This seems like an arbitrary line to draw. Should your photo viewing software also not provide editing tools? If it's useful, why not?
HN user
RussianCow
I'm a full stack, jack-of-all-trades software engineer and leader. I use a combination of boring and cutting-edge tech to build useful products quickly.
My current venture is Semi-Decent, a software consultancy that helps businesses build and scale software products. We are intensely pragmatic and focus on combining lean MVPs with user research to guide our clients towards growth. We take the risk out of software projects by keeping scope short and focused and offering a money-back guarantee. https://www.semi-decent.com/
I also run a side business making bespoke, sustainable craft cocktails for events: https://www.theminimalmixologist.com/
Please get in touch! I'd love to chat.
sasha+hn@chedygov.com
It depends. For something high stakes or inherently complex, sure, you don't want to have to clean up the agent's mess afterwards. But for many tasks like building web UIs, the difference in output quality is going to be small enough that iteration speed will win over quality.
With a fast enough model, I can iterate on the UI of a given screen 4-5 times before Opus finishes its first attempt.
Once you've used a model that runs at hundreds of TPS, it's hard to go back. Everything completes so quickly that you can iterate without breaking out of flow state. My biggest gripe with slow (<50tps) LLMs is that I've lost all the mental context I built up by the time it's done, which makes it extremely difficult to explore or iterate on solutions.
The system prompt and available tools would likely only change for different agent types. So how they're launched probably doesn't matter.
I say this as I don't actually know how Clade Code does this, since it's not open source, but I fail to see why two agents doing the same thing launched at different times would have different tools and system prompts.
Yes, Codex has its own sandbox which you can disable: https://learn.chatgpt.com/docs/sandboxing?surface=app
Cursor has something similar. I don't know about Claude Code but I assume it does as well since Anthropic has open sourced their own sandboxing tool.
Compared to the primary agent, maybe. But it's highly unlikely that all the agents have different tools and system prompts than each other, and those account for the bulk of the context per the post.
You cannot; you must use either their Devin Desktop app or the Devin CLI.
But they don't appear to subsidize them to the same degree. I've only been using Devin for less than a month, but I've been hitting the limits of the $20/month plan way more quickly than I'd expect, and definitely more quickly than with Claude Code or Codex.
So far, Cursor provides the best value for their subscription, but I have to imagine they're basically lighting money on fire. There's no way their current pricing is sustainable.
But that 10% is the most important part! Getting the plumbing wrong means you might have bugs or your code is brittle. Getting the domain-specific business logic wrong means your product doesn't fundamentally solve the correct problem.
Ironically, Devin Desktop is one of those tools. It supports any harness that supports ACP (which is most of them)—you can use Claude Code, Codex, OpenCode, etc from the Devin Desktop UI.
I'm currently experimenting with OpenSpec[0] as the "framework" and using different subscriptions for different parts of the spec-driven process: Opus via Claude Code for exploration, Devin SWE for building, and GLM 5.2 via the Z.ai Coding Plan for verification. I don't love having to mix and match harnesses, but in practice it's barely more effort than switching models.
But some requirements you don't realize you have until you start building. With a fast model, you can surface those really quickly and have more time to iterate and explore different solutions. With a slower but smarter model, you just hope that what it produces after an hour is what you were imagining.
And yes, with Fable, the chance of that is higher than with SWE/Composer, but in my experience it's not so much higher that the extra time and cost is worth it. But it certainly depends on your goals and what you're building.
The regular one (not the fast variant) is free but slow. The "Lightning" variant (which uses Cerebras and gets supposedly 1000 TPS) costs $12.50/M output, $2.5/M input, $1/M cached input. So it's quite a bit more expensive than SWE 1.6.
The "Lightning" (Cerebras) variant isn't free, only the regular one, which runs closer to 50 TPS in my experience with SWE 1.6.
That's not why Elm is an ideal language for LLMs: it's because, if it compiles, it's most likely working software. Agentic workflows have gotten significantly better over the last year, so LLMs using languages like Elm, Haskell, or even Rust have an amazing feedback loop where even lower quality models can keep trying until things compile.
asking an LLM to reverse engineer and make your own plugin is trivial.
If you already have engineers on staff, a few tens (or even hundreds) of dollars per month per plugin is likely a rounding error budget-wise. If you don't have your own engineers, you're probably not going to be able to produce something as good (reliable, well thought out, etc) as a commercial offering.
I had the same gut reaction as you, but the reality is much more subtle. We work with several clients who are bought into at least one of these ecosystems, and there's no way the math ever works out in favor of building an in-house solution.
At least 3 times a month. I have a rental property and my tenant prefers to mail a check instead of paying extra to pay electronically. My spouse gets paid by check for dumb reasons I won't get into. I sometimes get dividends from my insurance company via check. And then several family members still prefer to use checks to pay each other back instead of Venmo or other electronic services.
I blame it on the fact that the US doesn't have a free electronic bank transfer system like the rest of the developed world.
Also why do you need bank app on your phone?
Many banks gate features like mobile check deposit behind the native app. The nearest ATM is 20 minutes away from my house, so unfortunately I consider this feature essential.
En, I think you’re just trying to justify your pre-existing position that this can’t work.
I never said it can't work. I just said that finding the correct medical digagnosis is different than finding a solution to a software problem.
Also, there are multiple "correct" ways to code something, so imperfect code that solves the problem is still useful. A medical diagnosis is either correct or incorrect.
At that point, cut out the LLM and just see the radiologist.
Probably because HDR on the vast majority of non-OLED monitors is useless. You really need a monitor with great contrast and a good HDR implementation for it to be of any benefit.
I'm not an expert, but I think those are the same thing. But for an LLM etched onto a whole wafer, it doesn't make sense to disable part of it since that would remove some weights entirely.
I don't know about you, but I generally don't write code in a vacuum. Other people may have touched it before me. Those other people may have made poor decisions.
Not that I'm immune from choosing the wrong abstraction sometimes. More than once the "other people" was me. We all make mistakes.
But even then a twice a week household cleaning hire is going to cost upwards of $1500/mo unless you're being particularly exploitative.
Sorry, what? Unless you're doing a deep clean of your house twice a week or you live in a particularly HCOL area, those numbers don't add up. You shouldn't be spending more than $1k/month on household chores, and even that seems high.
Source: A client of ours runs a "personal help" service (mostly focused on household tasks like laundry, tidying, organizing, etc as opposed to deep cleaning) so I have a lot of data on this. And they're a relatively premium service compared to some of the cheap labor you can actually buy. But they also don't operate in SF or NYC, so maybe prices are drastically different there.
The Chinese open weight models have been ahead of Sonnet (at least for coding) for a couple months now. I tend to take benchmarks with a huge grain of salt, but in my own experience, the latest versions of Kimi, MiMo, and GLM (pre-5.2) had already surpassed Sonnet in terms of output quality for a fraction of the price.
With that said, I'm excited to try GLM 5.2 because I still end up reaching for Opus and GPT 5.5 for many tasks because the open models tend to get stuck more often on complex problems.
Isn't that true of any provider? Anyone could be lying about what they're serving.
I've been doing the same, though admittedly out of curiosity more so than lack of funds. The open models are catching up quickly in their abilities, to the point where they're (mostly) not doing stupid stuff regularly, but you have to be very specific about what you want. I found that Opus, for example, is much better at asking me to clear up ambiguity in a request before starting, whereas the Chinese models tend to "fill in the blanks" and make their own assumptions.
My current workflow involves going from PRD -> execution plan -> build -> review, and this works nicely with open weight models like GLM 5.1, Kimi K2.6, and DeepSeek V4 Flash. With Opus I can generally skip the PRD entirely, and sometimes even skip the plan, and 80-90% of the time it does exactly what I want. But that can easily burn $5-15 for one feature, whereas it'll cost maybe $1-2 with the open weight models (at API pricing).
But those things won't be sped up by a faster LLM, so I feel like that's not what the OP is talking about.
Do you mean Flash and not Pro? I haven't tried it personally, but according to OpenRouter, the fastest DeekSeep V4 Pro providers are only ~50tps. That's slower than Claude Opus.
https://openrouter.ai/deepseek/deepseek-v4-pro?sort=throughp...
Same with cars. Half the battle is sometimes just unscrewing an old bolt that hasn't been touched in 10+ years without breaking it, or getting the rusted on rotors to come off.