It would be interesting to hear if the addition of Reddit data makes LLM stronger or weaker.
Outside of some (technical)subs most of the comment seems to be of a far lower quality than what you find in books or in websites.
HN user
It would be interesting to hear if the addition of Reddit data makes LLM stronger or weaker.
Outside of some (technical)subs most of the comment seems to be of a far lower quality than what you find in books or in websites.
For users, Reddit has been in a downwards spiral since the IPO.
It's pretty clear that their focus is no longer on providing the best communities (see the many examples of power hungry mods, see below) or the best UX (see the deprecation of the API or the recent changes to old.reddit.com described in the post).
The "product" is now selling data to LLM companies, everything else is clearly a secondary concern. It's a shame because at some point I really enjoyed participating on Reddit and it was a good source of valuable information.
to give some examples of mod abuse:
- /r/energy used to (or still does?) ban anyone in favour of nuclear power
- /r/ubisoft and /r/assasinscreed do not allow any posts criticizing the writing or characters, these are clearly run by the company
- /r/unitedkingdom is famous for shadow banning everyone who's critical of some government program (I got a ban for criticizing motability cars)
- /r/europe banned me 1 minute after submitting a chat control post that got 1.5k upvotes before it was removed
- some vaccination subs ban anyone who posted in anti-vaccination subs (irrespective of what they post)
Pricing often reflects what the vendors (expects) the customer is willing to pay. It seems that Google is still trying to find their niche in the market.
Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.
People overestimate what can happen in a year and underestimate what can happen in 5.
I'm betting that increased model efficiency and hardware optimisations will get us there a lot sooner. Biggest hurdle would be the memory prices though, if those do not drop back down it might take 15.
I feel that a better title for this article would be: "Some minor annoyances that, when fixed, would improve OpenCode"
# Prompt Cache Misses
It globs your filesystem and re-reads AGENTS.md (injected in turn-0 system prompt) on every SSE turn. If you put a quick note in AGENTS.md to be read in the next session, you immediately force a full re-evaluation.
Personal favourite: it puts the current date in the turn-0 system prompt and re-evaluates every SSE turn. If you’re using OpenCode at midnight you get a full prompt cache miss.
Okay, I can live with those.
# Compaction
Want to sit for 10 minutes while the LLM server prefills the entire session with a new prompt prefixed to it, just to turn it into 5 bullet points that go at the top of a new session? Me neither. I get what they are going for, but I’ve not seen it work well. Neither compaction nor pruning is implemented well, and they interact poorly.
Is this an OpenCode specific issue? I've seen the same with Codex and Claude
# System Prompts
The default system prompt is opinionated (fine) but it has shit opinions (not fine). It took me a while to figure out why my agent kept saying “Use ABSOLUTELY NO COMMENTS” when dispatching subagents.
Okay, so change it? Any LLM is opinionated, this system prompt enforces consistency across different models which seems reasonable.
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
That's a massive model!
The shift from "value" models to "intelligent, huge and slow" models coming from China is an interesting change in strategy.
My main issue with GLM 5.2 and Kimi 3 is that they're extremely token hungry and thus feel slow(er) to use.
As much as I like GLM 5.2 it's clearly a step below Opus (or even Fable) for more complicated tasks. I would place it at Opus 4.6/4.7 level.
Having said that, the safety system on Fable makes it an extremely unattractive model. It feels that half of the time you're paying double for Opus level performance.
Yes but Kubernetes takes these battle-tested scripts, and allow everyone to use them with a few lines of YAML ;)
I understand the dislike of YAML but a Kubernetes deployment is ~50 lines, if I had to build my own scripts with a similar feature set I don't think I would be able to get it down much more than that.
Kubernetes makes complex things (e.g blue/green deployment, auto-scaling, failover) possible irrespective of the underlying cloud/hardware with a good and standardized API.
It's absolutely overkill for small teams and homelabs (I run a cluster myself) but an absolute godsend if you do need the advanced functionality.
When we tried to do a pilot with their cloud we couldn't even sign-up. None of the corporate credit cards were accepted.
In addition to that the form basically only worked in Edge. We emailed support, they changed something on the backend. It still did not work. We gave up.
In retrospective that was a very clear warning sign that their priorities were misguided. I'm glad we did not waste any further time and effort on them.
This is an interesting point and the article should definitely take these factors into account.
It's indeed very worrying what we ask medical professionals to put themselves through for their jobs. I think we can all agree that having a well rested doctor or nurse would be preferable over a stressed/tired one. The amount of hours and night shifts that (young) doctors have to do and the extreme competitiveness of the field (partly) drives this.
I understand that it would drive wages down (somewhat) if we educated more doctors and obviously we shouldn't lower our standards substantially but it seems like everyone involved would benefit from this.
A friend of mine, whose a doctor, told me once that the best way to ask for medical advice is to ask the doctor what he/she would recommend for their own sister/brother. Siblings are close enough that he would not want them to suffer unnecessarily but it eliminates the personal factors. Obviously it differs per doctor but in my experience it usually leads to a good conversation about the trade-offs for medical care.
I've found these kind of models (Deepseek + Xiaomi) to be absolutely excellent when it comes to writing documentation for code. We have a bunch of internal tasks that need to be documented for a non-technical team.
I added 20 USD in credits for the Xiaomi models a while ago and they've been happily writing and updating hundreds if not thousand of pages and I still have 7 USD left!
It's really cool and interesting to see the kind of engineering that goes into Xiaomi (and Deepseeks) inference optimizations. Z.ai has also published some interesting papers although I haven't had a chance to go through them yet.
It does inspire hope that the Chinese labs seem to be so open although the sceptic in me does wonder what their end game is.
Surely, from a purely economic perspective it would be wiser to keep this proprietary and benefit from the increased API traffic?
While I mostly agree with your statement there's evidence that testosterone is linked to social status and mental well being.
A 50% drop most likely has a multifactorial explanation, being told that some traditional male traits are bad (and thus lower well being or social status) or medicated away (see e.g the rise in ADHD and Autism diagnosis) might have some effect.
I'm not nearly knowledgable enough to give any reasonable estimate but it would not surprise me if it was higher than 0%.
Short-term, follow the steps on the website and contact your political representative to explain to them why it's such a bad idea.
Long-term, switch to another messenger app that's opensource and truly E2E encrypted.
That also shows why this is such a foolish proposal.
The truly scary people are not on the "consumer" chat apps anyway and most certainly will be the first ones to switch to another communication channel if this passes. If this will have any effect it'll be that some, "dumb" criminals will be caught.
I remember looking into SDR when I was a student, really amazing what you can pull down for the price of a dinner these days.
I know that you can get the same pictures from the internet but building something like this and seeing how all the pieces fit together is extremely cool.
Also an extremely cool project to do with kids, the output is visual but it's such a cool combination of hardware, software, some light math (to design the attena) and crafting (to build it).
I'm not an expert but I think the SBB is already pretty good at handling this. I think they already run measuring wagons (Oberbaumesswagen) with grond penetrating reader and ultrasonic measurement and use flow sensors to monitor drainage.
I would expect that the solar panels impact the efficiency at least somewhat but apparently not enough to cause real and enough issues for the SBB or perhaps they see ways to improve this in the future.
I imagine that the cost to install is fairly low since train tracks require regular monitoring and maintenance so it's fairly cheap to add the installation and maintenance on top of the existing schedule.
The manufacturer claims that durability should not be an issue. Time will tell.
This is an extremely weird comment that doesn't add anything to the conversation.
Here on HN we discuss facts, jumping straight into racism has no place here.
They might be sending some user requests to Anthropic to gather trading data for their own models. If they do so, perhaps they need to add some tracer to request that they prefer to hide.
Split from: Left Party
So basically the left but with a stricter view on immigration?
I've gained and lost 10kg twice in my life. Maintaining the weight loss isn't that hard once you've a rhythm dialed in.
In my case I just weight myself daily, track the weight and scale my food consumption with the current trend. If I'm gaining weight I'll skip a meal.
It takes a while to figure out what works for you but I can tell you that making small lifestyle changes to maintain your weight is fairly easy compared to figuring out how to lose 10 kg.
GLM 5.2 is great but it heavily detoriates once the context window gets past 200k tokens.
I've had more success with creating a plan first and then implementing it in (short-lived) sub-agents.
Ironically good software architecture patterns (small functions, single responsibility) heavily impact the performance of these models as well. They do surprisingly well in well architectured codebases.
They do very poorly in anything that's a mess where Opus and GPT 5.5 still get reasonable performance.
6bn seems excessive but despite GPT 5.5 arguably being better than Claude I don't see a lot of adoption of Codex yet.
Some of my coworkers even use Sonnet (the default in Claude Code for the 20 USD subscription) and see no reason to change even though that model is definitely "outdated" compared to current SOTA.
I've been very pleased with it's performance over the last few days.
It's definitely not near Opus 4.8 level but it's very impressive nonetheless and it does do design extremely well.
This is too little, too late. Europe really need to start focussing.
All these tiny niche models are perhaps fun as an academic exercise or great for the researchers resume but I highly doubt that they'll add any value or will be used for anything serious.
Even if this becomes a somewhat decent model with a fantastic understanding of "gezellig", "kring verjaardag" or "pannenkoeken", how many people will interact with it before the limits of it will drive them back to a frontier model?
Even if the purpose of this is government & other regulated industries, do we really want our government to use a poor model? Either do it right or don't do it at all.
Thanks for pointing that out! I'm not in the US and I guess it's not illegal in China (given that Deepseek was more than happy to do it).
That does raise an interesting question, what kind of laws should LLMs (attempt to) follow? It's easy enough to spoof the country in the system prompt. I wonder how ChatGPT would respond if I told it I was located in a developing country without any piracy laws.
Same, I had Deepseek search for, download and transfer (to my Linux emulation machine) the best Dreamcast games yesterday.
GPT refused to do so (citing that it's illegal even though I own the games). Deepseek did a wonderful job for 7 cents.
At work I use Opus because, why not? But I could easily switch to a less capable model if needed.
And compared to other countries, I think Xenophobia is low
I would agree and also suggest that initiatives like this play a large role in doing so. While there's a lot of bullshit arguments coming from the "yes" camp they do make some reasonable points and it's important that we discuss them to show what the trade-offs are.
I cannot speak for all Swiss but knowing that it was a democratic decision to continue with some, high skilled, immigration makes it far easier to accept than if some government employee in Bern would've made that decision single handed.
Based on my first impressions it's about 6 months behind the frontier labs. So very similar to Opus in January.
That is, pretty damn impressive and very useable. When it comes to architecture or complex problems it does noticeable worse but I don't think anyone expected anything else.
One particular interesting strong point seems to be design and user interfaces. It does seem to punch above it's weight there but that might just be personal preference.