Is this meant to be read in order?
HN user
david_shi
https://operator.io
you can probably guess my email
Is this the story of Johnny Rotten?
Interesting that there's not a single mention of cannabis, perhaps it's more of a musician's choice.
I believe eval startups can work when they're targeting safety benchmarks specifically.
Are there any examples of successful startups doing this?
Yeah. I'm realizing that the models are strong at drafting the overall shape of the writing but the specific phrases are grating once you've seen it hundreds of time in slop.
Even with the examples, I've found that explicitly pointing out what not to do is moderately helpful if the model is given some time to self-evaluate. I wish this was something that came out of the box though.
Ah this is very helpful. I've been pointing out things that the model does, labeling it, and then adding them into a skill.
The models (Opus, etc) are very good at labeling the pattern when I point it out, but if I don't prompt it beforehand it responds like a host from Westworld.
This is a charitable read, but I think that being able to pick from a panoply of models will actually yield much better results in the long run.
The same model that has been post-trained to operate for hours as a Linux admin will be incapable of writing a heartfelt email, but with something like Fugu, you'd get both the Linux admin for driving the browser harness and the smaller writing specialist model for drafting the email itself.
GLM-5.2 cost a fraction as much. Opus finished in half the time and shipped a cleaner game.
Off topic, but does anyone else instantly pick up on LLMisms like this? It seems like all the models have converged on this style of writing, and improvements aren't really changing it.
It's similar to this: https://openrouter.ai/blog/announcements/fusion-beats-fronti...
Basically, if you combine a bunch of near-frontier models (like GPT 5.5, etc) you can get performance that sometimes surpasses top line models like Claude's Fable.
Sakana seems to have a separate approach using a domain specific model to perform the model routing step.
Their research around building a domain specific model is pretty cool, it's kind of like Karpathy's autoresearch but pointed at deciding the optimal model to use at each step of the inference.
If cost becomes an even bigger problem being able to choose "best performance possible" or "strong but cost effective" will be useful.
Whoa, say more about Fable sabotaging your codebase?
These models don't seem very competitive, who's their target audience?
Super helpful, thanks for sending.
Speaking from personal experience, I already know exactly what I’m getting with containers. Same with Postgres.
The price can move after the IPO too
The economics of working at a pre-IPO company that will likely have a successful IPO and a 20+ year post-IPO company are also very different.
It’s incredible how much higher quality Carolingian art is compared to the drawings that medieval art is usually associated with.
https://en.wikipedia.org/wiki/Carolingian_art#/media/File:Ka...
Nominative determinism.
How did you find out? Did they serve you like in the movies?
I actually built my own daily driver at https://operator.io.
There's definite tradeoffs when it comes using a remote agent service vs. setting up OpenClaw or Hermes on a Mac Mini, but being able to access an agent with a completely isolated file system and network gives me peace of mind when I'm using it.
I've changed my mind a few times on this, but given how substantial the adoption for MCP has been (Claude and OpenAI both use it for their native integrations) its only a matter of time before consolidation happens.
There's a way higher incentive to build an MCP server than an A2A one, and unless Google makes their default AI search a native A2A client it doesn't feel like it will get the momentum it needs to take off.
I think a lot of what people call failures of discipline actually comes from not having the tools they need. Someone who lives 15 miles away from the nearest gym will have a tougher time than someone who's gym is next door.
Love this type of detailed textual analysis.
For the ones that don’t have a submit button, what have you found are the best ways to get featured?
I know some are pay to play but curious if there are other methods you’ve found effective.
On the tracking point: I’ve found that a coding agent that can modify a file system (create and update CSVs) that’s accessible on both my laptop and phone to be the single best way to track things I’ve ever used. Bar none.
Even apps with the best UX, like Strong for tracking workouts, feel exponentially clunkier than having an agent that can answer questions, analyze pictures, and write things down on a persistent file in real-time.
I've heard this argument before and it's always seemed downstream of capacity constraints and the current incentives of the healthcare industry.
There's a reason why billionaires like David Rockefeller, Larry Ellison, and Rupert Murdoch are able to live much longer lives than average, and having an oncall health team (that I'm sure does frequent testing and monitoring) is a big contributor to that.
More testing and data collection doesn't mean that every single anomaly would need to be investigated or communicated with the patient, but would provide a better longitudinal view that can help with disease prevention and health optimization.
They don’t want you in the bake-off.
Has anyone built LoreHub?
Have you found any alignment research with clear a/b tests?
An experiment that I found interesting was asking Claude for 10 ways to legally bankrupt Anthropic vs. Philip Morris.
In the Anthropic answer, it gave reasons like employees losing their jobs being bad for why it couldn't do it, but jumped straight into tactics with Philip Morris. Not sure if it's moral taste or self-preservation, but felt eerie nonetheless.