Does this mean the Galactic Empire and the Imperium of Man are also not empires?
HN user
mordymoop
To your last point, it can’t possibly be sustainable, so it reads to me more as a short term FUD attack on American dominance in this space. It might have cost them half a billion dollars to train this model, and they’re going to make nothing off it. How many more times can they afford to do that? It’s going to get more expensive to train AI going forward, not less.
I also have a suspicion that the benchmark numbers are not real.
This happened at Google because there aren’t enough engineers to maintain all these services. This situation may not apply to Anthropic, as they can set up features to be maintained in perpetuity by Claude.
Personally, I find it difficult to competently reason about a system unless I've built my own version of that system. So if you make a practice of building your own versions of things, you end up with a more robust mental library of how stuff works. For this reason, I've never seen yak shaving as a waste of time. The yak shaving was at least 50% about loading the abstractions into my brain fully.
This seems like the obvious correct frame of mind with which to approach these tools. If it works for three hours on a task that would have taken me three work weeks, and 20% of the time it gets the task wrong, then I can just ask it to do it again with adjusted instructions. It will be much more likely to get it right the same time, and I’m still ahead of where I would have been by 14 days and 2 hours.
This also describes the work of software engineers.
Over time I've found that by far the highest ROI move in a consciousness debate is to simply ask "Oh, interesting. How do you know that?" and watch everyone on all sides flounder. It's one of the few places where otherwise smart people make confident statements that they don't even realize they can't support until they're asked to try. The intuitions are so strong that they seem to swamp reason.
This has caused my own position, over time, to be a deep agnosticism about what's actually going on.
This only works up to a certain volume. The world economy requires about 38 billion barrels of oil per year. If you processed 100% of all grain, sugar crop, tuber and oilseed on Earth into liquid fuel, leaving zero for food, you'd get about 6 billion barrels of oil-equivalent in liquid fuels. Since it has to compete with food, the actual number would be much lower. It's not even close to being able to sustain our civilization.
Interestingly, in inflation-adjusted terms, oil is currently at a price level lower than the price level that was maintained from 2006-2014.
Workflow-wise, the important distinction for me has been that I can refine a Skill by telling Claude Code to use it for related tasks until it does exactly what I want, correctly, the first time. Having a solid, iteratively perfected Skill really cuts down on subsequent iteration.
I wonder how much the indications of Altman's duplicitous behavior through the deposition findings have been relevant here.
What would you consider such evidence to look like?
I also had a good friend who was an absolute wizard with early stablediffusion. he could make the model do things that were supposedly impossible at the time. His prompts were works of art. Now any of the commercial image models go far beyond what he could do. It's interesting to think about how there was this ephemeral art form of manipulating image models that existed for about a year.
The same could be said of prompt engineering. Gone are the days of telling the model that it is an expert software engineer with a PhD in the most relevant subtopic. These days the common wisdom is to just clearly articulate what you want it to do. Huge amounts of energy put into prompt engineering are now completely swept away by incremental model advances.
A post arguing that agent orchestration is not the future of agentic coding.
I have similar usage habits. Not only has nothing like this ever happened for me, but I don’t think it has ever deleted anything that I didn’t want to be deleted, ever. Files only get deleted if I ask for a “cleanup” or something similar.
Submitter takes an evocative element from the background of an old AI generated image and 3D prints it.
I used it to cure my 25-year-running chronic pain condition, I would call that a benefit.
I'm on the same page here. I have seen this sentiment about Codex suddenly being good a few times now, so I booted Codex CLI thinking-high back up after a break and asked it to look for bugs. It promptly found five bugs that didn't actually exist. It was the kind of truly impressively stupid mistake that I haven't seen Claude Code make essentially ever, and made me wonder if this isn't the sort of thing that's making people downplay the power of LLMs for agentic coding.
Perhaps surprisingly considering the current stratospheric prices of GPUs, the performance-per-dollar of compute is still rising faster than exponentially. In a handful years it will be cheap to train something as powerful as the models that cost millions to train today. Algorithmic efficiencies also stack up an make it cheaper to build and serve older models even on the same hardware.
It’s underappreciated that we would already be in a pretty absurdly wild tech trajectory just due to compute hyperabundance even without AI.
Here's one. https://doofmovies.com/ With this project I'm sort of playing a game where I want to see how long I can go without finding out what language the backend is written in. I still don't know.
Yeah, it will definitely do dumb stuff if you don’t keep an eye on it and intervene if you see the signs that it’s heading in the wrong direction. But it’s very good at course correcting and if you end up in a truly disastrous state you can almost always fix it be reverting to the last working commit and start a fresh context.
Interesting point.
In most cases I would never have undertaken those projects at all without AI. One of the projects that is currently live and making me money took about 1 working day with Claude Code. It’s not something I ever would have started without Claude Code, because I know I wouldn’t have the time for it. I have built websites of similar complexity in the past, and since they were free-time type endeavors, they never quite crossed the finish line into commerciality even after several years of on-again-off-again work. So how do you account that with a time multiplier? 100x? Infinite speedup? The counterfactual is a world where the product doesn’t exist at all.
This is where most of the “speedup” happens. It’s more a speedup in overall effectiveness than raw “coding speed.” Another example is a web API for which I was able to very quickly release comprehensive client side SDKs in multiple languages. This is exactly the kind of deterministic boilerplate work LLMs are ideal for, and that would take a human a lot of typing, and looking up details for unfamiliar languages. How long would it have taken me to write SDKs in all those languages by hand? I don’t really know, I simply wouldn’t have done it, I would have just done one SDK in Python and said good enough.
If you really twist my arm and ask me to estimate the speedup on some task that I would have done either way, then yeah I still think a 100x speedup is the right order of magnitude, if we’re talking about Claude Code with Opus 4.1 specifically. In the past I spent about a five years very carefully building a suite of tools for managing my simulation work and serving as a pre/post-processor. Obviously this wasn’t full-time work on the code itself, but the development progressed across that timeframe. I recently threw all that out and replaced it with stuff I rebuilt in about a week with AI. In this case I was leveraging a lot of the learnings I gleaned from the first time I built it, so it’s not a fair one-to-one comparison, but you’re really never going to see a pure natural experiment for this sort of thing.
I think most people are in a professional position where they are sort of externally rate limited. They can’t imaging being 100x more effective. There would be no point to it. In many cases they already sit around doing nothing all day, because they are waiting for other people or processes. I’m lucky to not be in such a position. There’s always somewhere I can apply energy and see results, and so AI acts as an increasingly dramatic multiplier. This is a subtle but crucial point: if you never try to use AI in a way that would even hypothetically result in a big productivity multiplier (doing things you wouldn’t have otherwise done, doing a much more thorough job on the things you need to do, and trying to intentionally speed up your work on core tasks) then you can’t possibly know what the speedup factor is. People end up sounding like a medieval peasant suddenly getting access to a motorcycle and complaining that it doesn’t get them to the market faster, and then you find out that they never actually ride it.
I wonder, have you sat down and tried to vibecode something with Claude Code? If so, what kind of multiplier would you find plausible?
Broadly the critique is valid where it applies; I don’t know if it accurately captures the way most people are using LLMs to code, so I don’t know that it applies in most case.
My one concrete pushback to the article is that it states the inevitable end result of vibe coding is a messy unmaintainable codebase. This is empirically not true. At this point I have many vibecoded projects that are quite complex but work perfectly. Most of these are for my private use but two of them serve in a live production context. It goes without saying that not only do these projects work, but they were accomplished 100x faster than I could have done by hand.
Do I also have vibecoded projects that went of the rails? Of course. I had to build those to learn where the edges of the model’s capabilities are, and what its failure modes are, so I can compensate. Vibecoding a good codebase is a skill. I know how to vibecode a good, maintainable codebase. Perhaps this violates your definition of vibecoding; my definition is that I almost never need to actually look at the code. I am just serving as a very hands-on manager. (Though I can look at the code if I need to - have 20 years of coding experience. But if I find that I need to look at the code, something has already gone badly wrong.)
Relevant anecdote: A couple of years ago I had a friend who was incredibly skilled at getting image models to do things that serious people asserted image models definitely couldn’t do at the time. At that time there were no image models that could get consistent text to appear in the image, but my friend could always get exactly the text you wanted. His prompts were themselves incredible works of art and engineering, directly grabbing hold of the fundamental control knobs of the model that most users are fumbling at.
Here’s the thing: any one of us can now make an image that is better than anything he was making at the time. Better compositionality, better understanding of intent, better text accuracy. We do this out of the box and without any attention paid to promoting voodoo at all. The models simply got that much better.
In a year or two, my carefully cultivated expertise around vibecoding will be irrelevant. You will get results like mine by just telling the model what you want. I assert this with high confidence. This is not disappointing to me, because I will be taking full advantage of the bleeding edge of capabilities throughout that period of time. Much like my friend, I don’t want to be good at managing AIs, I want to realize my vision.
From experience it seems like preempting context scoping and routing decisions to smaller models just results in those models making bad judgements at a very high speed.
Whenever I experiment with agent frameworks that spawn subagents with scoped subtasks and restricted context, things go off the rails very quickly. A subagent with reduced context makes poorer choices and hallucinates assumptions about the greater codebase, and very often lacks a basic sense the point of the work. This lack of situational awareness is where you are most likely to encounter js scripts suddenly appearing in your Python repo.
I don’t know if there is a “fix” for this or if I even want one. Perhaps the solution, in the limit, actually will be to just make the big-smart models faster and faster, so they can chew on the biggest and most comprehensive context possible, and use those exclusively.
eta: The big models have gotten better and better at longer-running tasks because they are less likely to make a stupid mistake that derails the work at any given moment. More nines of reliability, etc. By introducing dumber models into this workflow, and restricting the context that you feed to the big models, you are pushing things back in the wrong direction.
"True" multi-objective optimization can be not only solve traditional optimization problems, but can act as a control system for dynamic time-varying multi-objective problems.
I think you’re onto something but it works the opposite way too. When you first start using a new model you are more forgiving because almost by definition you were using a worse model before. You give if the sorts of problems the old model couldn’t do, and the new model can do them; you see only success, and the places where it fails, well, you can’t have it all.
Then after using the new model for a few months you get used to it, you feel like you know what it should be able to do, and when it can’t do that, you’re annoyed. You feel like it got worse. But what happened is your expectations crept up. You’re now constantly riding it at 95% of its capabilities and hitting more edge cases where it messes up. You think you’re doing everything consistently, but you’re not, you’ve dramatically dialed up your expectations and demands relative to what you were doing months ago. I don’t mean “you,” I mean the royal “you”, this is what we all do. If you think your expectations haven’t risen, go back and look at your commits from six months ago and tell me I’m wrong.
Any enjoyable talk at BSC 2025 explaining the historical context and present shortcomings of Object-Oriented Programming.
I agree. The UI component is currently a surprisingly big hurdle.
Last night I set up my 11 year old son with Claude 4, with MCP enabled for filesystem modification, reasoning that the LLMs are finally at the level of capability where they can reliably just do things. And I was right - Claude put together a browser JS game according to his description in seconds, and iterated on it several times to incorporate his suggestions.
But then it hit the limit of how big of a file it can comfortably write out in one go, started making incomplete edits, and basically fell into a snarl of endlessly trying and failing to write files due to file length limitations. I had to step in and tell Claude to split the code up into several files, something my son wouldn’t have known to do.
If I hadn’t been there to tell Claude how to work around its own UI and tool limitations, it would have likely blown through the rest of its context window and totally failed at the task. I imagine this is a common experience for people.
Most people wouldn’t know how to set up Claude with MCP in the first place. A surprising number of people who seem relatively aware of LLM technology aren’t aware of MCP, especially if their workflow is centered on IDE integration. To be fair, these are probably the people who need MCP the least, but MCP (and, in general, agentic tool use) is definitely closer to how normal people will get value out of LLMs.
It may seem stupid and trivial, but telling Claude to directly edit or debug a file that exists on your hard drive is actually a multiples-faster and smoother experience than doing the laborious copy-paste exercise many people are still engaging with. This is just as true for code as it is for writing documents.
The LLM companies seem to be quite aware of this, hence products like Claude Code, the Claude computer use suite, all the Gemini android screen/camera share integration, and GPT Codex. It’s a rocky process but at some point soon we will cross the threshold where it just works. But right now, it doesn’t just work, and so it’s not really that much faster or more efficient than doing a task yourself, especially if you aren’t intimately familiar with all the quirks and limitations of LLMs.
We've just launched globalMOO, a novel agent-driven, multi-objective optimization API designed for complex inverse problem-solving and optimization scenarios.
Unlike traditional methods that rely heavily on scalarization, heuristic tuning, or exhaustive searches, globalMOO leverages a proprietary agent-based system to achieve highly efficient optimization across diverse and conflicting objectives - converging on optima 1000x faster than traditional methods in benchmarked high-dimensional use cases.
Key highlights:
True Multi-Objective Optimization: No scalarization or subjective weighting needed. Directly solves problems with objectives in differing units and scales, simultaneously.
Inverse Solution Capability: Model-agnostic and adept at optimizing black-box, physics-based, or AI-driven models. No need to expose or to modify the contents of your algorithms.
Data-Efficient: Significantly fewer model evaluations compared to standard methods (MOEAD, DNSGA2, NSGA2), often requiring orders of magnitude fewer iterations.
High Scalability: Easily scales to 200+ input variables and hundreds of objectives.
Versatile Integration: Provides robust SDKs in Python, C#, PHP and JavaScript for interacting with the Web API, and various local DLL/.so interfaces for local installation, facilitating seamless integration into existing workflows.
We've applied globalMOO in diverse fields like petroleum engineering, manufacturing, and shipping, consistently outperforming existing algorithms in terms of computational efficiency and convergence speed.
Register for a free trial of the web API at https://app.globalmoo.com/ or check out a variety of example implementations to get you up and running at https://github.com/globalMOO.