HN user

nullbio

310 karma
Posts9
Comments246
View on HN

It gives you "assurances" that can be broken. They could still be clean-rooming your prompts like Anthropic does. The only way to be certain is if they don't have access to your data to begin with.

Qwen 3.8 4 days ago

I predict that no one will use this and everyone will use Kimi K3.

Qwen 3.8 4 days ago

If by Opus you mean Opus 4 and not Opus 4.8, then sure.

Even if it were slightly more expensive, it's still a better sales proposition for a company if they can run it from a hardware provider with their own locked down VPS and ensure that their IP is protected and that their data isn't being stolen or trained on. The fact that it's a little cheaper is icing on the cake.

Honestly, it's the only sane way for the market to move. The big labs are obviously stealing our data. Anthropic in particular clean-rooms everything you feed it, even if you opt out, so that it can train on your IP without getting sued. It's a copyright grey area they're abusing because the law has not kept up.

July 27th. But I agree with you that this is just normal competition. The only threat this poses is to Anthropic. OpenAI is more than capable enough to out-compete, their pricing is already reasonable. Greedy Anthropic will do their very best to try and stop this though, because they want to maintain the status quo of ripping everyone off.

Exactly. It's an incredibly stupid idea to restrict access to open source models if your want your nations economy to succeed.

Wild that we still haven't figured out how to make good benchmarks. What we really need is a way to properly quantify what makes a codebases architecture good, and then evaluate architecture of generated codebases, or evaluate refactors of existing ones.

Also, a way to evaluate a models ability to remove dead code, clean up slop, reorganize, etc.

None of the existing benchmarks test any of the things that truly matter. They were relevant when models struggled to one-shot functions, but we're so beyond that point right now, yet the industry has not kept up.

I can’t help but wonder whether constant use of “agent” harnesses will lead to an atrophy of the software engineering (or really any field) muscles.

This will and is 100% happening. I have a friend who hasn't written code by hand in around a year, but uses LLMs every day, and he tells me he can't remember how to write code by hand anymore. He has been a developer for 10 years. But he's not working for anyone at the moment, so I imagine if he was in a workplace the circumstances would be different and they probably wouldn't settle for this.

I think that as a result of this it likely also atrophies the problem solving and architecture building skills that writing the code manually gives you. It just ends up degrading into a loop of tell the agent to do X and assume it knows what it is doing.

I'm sure it's better than KimiK2.7 and GLM5.2. Benchmarks aren't the full picture. Despite GLM5.2 performing well on benches and supposedly near frontier, in reality it was nothing close to frontier in actual usage.

You did the easy part. Now do the managed database part, and at scale, whereby I don't have to worry about any chance of data loss. Otherwise this isn't "building PlanetScale" - it's building 1/100th of it.

It annoys me when people claim they've "easily and quickly" built something that took many developers many months or years worth of work and optimization to build a solid product.

It's like someone who generates a pretty looking HTML page with an LLM and claims they've built a customer-facing product. So much slop these days...

Agreed. There needs to be full transparency on the capabilities of the models. Being given access is one thing, but you can't have labs building powerful models that could be manipulating the entire planet without the public knowing. The public needs to have a good pulse on what the leading models are capable of so that we can remain in touch with reality.

This is bad. Where is the transparency?

So a small group of technocrats get together behind closed doors and secretly share their AI breakthroughs, and determine whether it's too powerful or not for the plebs in the public.

Who is watching the watchers?