Actually I’ve found some flaws in the Torment Nexus and have a PR to improve it.
HN user
brookst
well-documented big catalogs are possible; they're just rare
They’re not that rare, it’s just that flat big catalogs don’t work well.
The most common failure mode for MCP server design is mapping MCP tools to system APIs. This results in a tool to list files, a tool to rename a file, a tool to copy a file, etc.
You can have 100 “tools”, you just need to define 10-12 areas (actual tools in the MCP model) and actions within each one. It creates a hierarchy that models can dive into when needed and avoid context clutter when that area is not needed.
So instead of 15 tools for file operations you have a single “file” tool, with actions like list, copy, rename, etc.
Single biggest arch requirement for any MCP server (well, any that has more than a handful of operations).
They don’t have goals any more than your phone had a goal.
Clearly one can get from Theorem 3 to Theorem 2 by composing with the isomorphism {X \cong {\bf C}^3} and using the previously mentioned fact that local injectivity implies non-zero constant Jacobian.
I mean, clearly, right?
You and me both, pal.
Hmm. When I have Opus or Fable write build plans, then invariably invariably read code extensively, but it’s true they don’t make changes and run tests.
I guess the way to formalize this would be to add a “make minimal change to confirm approach” instruction to the build planning prompt. Probably can even parallelize that so as the build planning prompt iterates subagents get launched to validate, similar to what research modes do.
I’m a little skeptical, but will give it a try.
The more honest ones are “how it looks on you” and don’t promise fit that depends on so many measurements that aren’t even visible in pics.
Sure, it’s idealized, but some people benefit from seeing color / neckline / etc on themselves as a visual reference.
Me, I’m a text-learner so I don’t get it at all. But I know people who get value.
How is that different than just having the frontier model write a build plan and store it, then have a cheaper model execute? That’s pretty normal practice.
The build plan should be better than the context that generated it since it strips out wrong turns and other noise. I think?
I said “if” for email, which I at least have recovery codes stored with someone trusted.
And that link is to recovery contacts, not codes. If you’re willing to trust one person you can get back into your Apple account after a disaster.
Phone number and email access, if you can regain access to them.
It is definitely a good idea to plan in advance and set up a recovery contact: https://support.apple.com/en-us/102641
If it’s 50 years and gradual, there will be many places very close. The only “winner takes all” scenario is a sudden breakthrough when nobody else is close.
Because it’s a gradient, not a binary.
If you use codex or Claude code or similar you can just have it ssh in.
I use a mikrotik at home, and I always get a review from Claude after making config changes.
I actually used to be a network engineer (many decades and career reinventions ago), and I appreciate that Mikrotik exposes all of the layers and connections between them.
But damn if Claude doesn’t always find something I got wrong in the config.
I use dictation to drive Claude code frequently, and it’s never had a problem with stream of consciousness and retroactive correction. Maybe try just direct voice and see if you notice any difference versus pre-cleaning?
Ant has quota resets two or three times a month, at least for me.
Yes, they’re following the Microsoft strategy of “if you name everything Copilot, you can report high Copilot usage”
Good sci fi premise, but not at all how AGI will happen.
It’s not going to be a singularity at one moment of time. It’s not going to be instant runaway self-improvement, no matter what doomers and fetishists say.
It’s going to be gradual. We’ll see glimmers of AGI, and the “G” part will be about gradual broadening of domains and deepening of capabilities.
All of the coding harnesses are already using their own tools to self-improve, and the HITL component is getting less frequent and at higher levels of abstraction.
That’s how AGI gets here: very gradually, no hard takeoff, and nobody will be able to pinpoint when exactly it happened.
So: also no single lab with a massive advantage, no government takeovers. It’ll be a lot less dramatic than the extremes believe. IMO, of course.
Great write up!
Why custom MIDI connectors of all things? What requirement couldn’t be met off the shelf?
Oh my goodness. I’ve used over 1B tokens on a single feature. I’m running at about 25B tokens/month right now.
My whole point was that by planning in advance you can shard the work into manageable sections with clear beginnings and outputs with acceptance croteria, and never compact, or even use more than a few hundfed thousand tokens in context.
It’s all hierarchical. Looking at an eval feature building right now, it’s 20ish build plans, each with zero to five or so /clear moments.
But maybe that’s the key thing… I don’t iteratively prompt ad hoc software writing. I do iterate on requirements, but if those are solid enough there is no “now write this function, now write that module”.
I’m not following the implication that these economics argue for jsing compaction instead of just clearing context?
Does it matter? As an end user I really only care about 1) how much I can do in a week, and 2) how long each task takes.
Subsidies would affect 1, but not 2. But if some VC wants to subsidize my Claude or Codex or whatever, awesome.
I love how AI means that all human-written code is now perfect, and any problems are onbviously a symptom of vibe coding. Another few years of hearing this and I may begin to believe I only wrote perfect code, when my current memory suggests otherwise.
Thanks, I wasn’t familiar with XTS.
Now I’m picturing the 3am end of quarter fire drill in finance when they discover the company has accounts receivable of fourteen billion $XTS and it’s appearing in the quarterlies.
How in the world does that absolve Dell/etc, OR reduce Microsoft’s culpability for letting their update service be abused?
This is the old “why buy a prebuilt PC when you can DIY for less”. Which is true, but completely misses the point that may people will happily pay to 1) not have to build it, and 2) have e2e warranty, and 3) have someone else do the software setup.
We can debate if it’s the most efficient use of money for a technical person, but it’s indisputable that many people get enough value to pay for the prebuilt.
I read the analysis and I don’t see anywhere that it claims to measure demand.
It seems pretty clear it’s about how many units Valve is selling (charging for, shipping) and explicitly not about reservations or demand.
Valve is being smart here. It is far better to be supply constrained at launch than to have enough capacity to meet initial demand (and then far too much capacity when demand slows over time).
Compacting at all is a mistake. With 1m context window there is no reason for a single task to require compaction.
Much better to spend tokens breaking the task into chunks, documenting and storing them durably, then executing each one in clean context and just /clear after.
It’s a similar concept to compaction, just planned in advance. Much much more effective, and doesn’t burn tokens and time (“wall-clock”, Claude) doing the compaction.
Ah, but who prompts the prompters?
$20m is really not much money to operate a company for 6 months. They must be close to break-even at least?
Or paint that turns black at low temp and white at higher temp?
Is GDP per capita really a good proxy for labor costs?