Training on generating SVGs directly at all is already fairly niche. Generating full scenes with a cartoony character is even nicher. But there's plenty of non-pelican cartoony SVG content out there (created, not written, by humans with vector design tools), and more importantly, plenty of vision models to give feedback on the output (just raster as a png). You could easily hill climb this niche skill, if you cared.
HN user
unholiness
I don't think this small amount generalization to other animals and vehicles is strong evidence they haven't trained on this, either directly or more generally.
No mention of KV cache, one of the biggest reasons not to switch models mid-stream. Once you pay, say, 50k input and 50k output Opus 4.8 tokens, you don't pay token cost for cache reads of the 100k context while it builds. Switch to another model, you'll start with the cost of 100k input tokens on the smaller model to get that context loaded and its unique KV cache set.
The post may have some real insight here, where this 100k working context is actually better than trying to summarize it's findings into a plan. It's also right that the handoff is a perfect time to edit the context (removing planning instructions). But it doesn't mention it's a trade-off: the plan is smaller, so it's a cheaper "on-boarding" of the next model. Seems quite plausible that this is worth it for 1-off tasks. If it's right, this is basically "plans are useless, planning is everything" for LLMs.
My problem is, I think the plans are useful. I want to review and edit them. I want them to give context for the upcoming code review (even if humans aren't reviewing). LLMs are notoriously bad at explaining why they're doing something in the moment. Humans are notoriously bad at accepting there's no reason why. I think the humans have it right here, and to bridge this gap want my PRs, my docs, and my comments teeming with reasons why. Plans help with that.
I mean, this is real. Your KV cache lives for an hour since the last token (on anthropic pro/Max at least, today at least, this has degraded in the past). So once you pay, say, 50k input and 50k output Opus tokens, you don't pay for cache reads of the 100k context while it builds. Switch to another model, you'll start with the cost of 100k input tokens on the smaller model to get that context loaded and its unique KV cache set.
The post may have some real insight here, where this 100k working context is actually better than trying to summarize it's findings into a plan. It's also right that the handoff is a perfect time to edit the context (removing planning instructions). But it doesn't mention it's a trade-off: the plan is smaller, so it's a cheaper "on-boarding" of the next model. Send quite plausible that this is worth it for 1-off tasks. If it's right, this is basically "plans are useless, planning is everything" for LLMs.
My problem is, I think the plans are useful. I want to review and edit them. I want them to give context for the upcoming code review (even if humans aren't reviewing). LLMs are notoriously bad at explaining why they're doing something in the moment. Humans are notoriously bad at accepting there's no reason why. I think the humans have it right here, and to bridge this gap want my PRs, my docs, and my comments teeming with reasons why.
(NB: ranted to long, reposting at top level)
Are you maybe misunderstanding this point?
It's a fascinating arms race right now: the AIs are training on the humans but the human hivemind is also training on the AIs. Readers are developing allergic sensitivities to language that sounds like an LLM produced it.
Humans are training to detect AI content. Humans writing more like AIs is an unrelated (and slower) phenomenon.
So, Claude Cowork for OpenAI? Feels overdue!
I've loved using Cowork recently for sourcing decisions. Things where seemingly everyone's out of stock or questionably reputable, just let Cowork spin for 20 minutes, find the best new and best used options that meet your requirements, probably also suggesting a different item that does the job and is available for cheap. I've done it enough that I'm starting to loathe clicking through these sites myself.
I also love the countdown. Counterintuitively it makes the game less stressful — the counter goes back to 30 no matter what. If the timer counted up, you'd constantly need to care about getting each one as fast as possible, or fret about one that's taking you minutes, etc.
With the countdown, you more want to care about the high level stuff: Keep your brain agile enough to get the next one, figure out more general patterns, ensure you "cover" the promising patterns, notice tough spots (with tons of patterns) where you you'll need to lock in. That stuff is more fun to focus on than speed.
Everyone wants to fail less, sure. It's not surprising people's feedback focuses on the mechanic that made you fail! That doesn't mean changing that bit will make the game more fun.
I have to imagine if you give it positions in text it'd be pretty great
Not at all? LLMs are a terrible match for the kind of analysis a chess engine does (scaled deep search, deeply trained position evaluations). It's just not that kind of tool.
In a hypothetical scenario where we were inventing the standard model in the first 10^-11 seconds after the big bang, you're right there would be an analogy there. But in that scenario, our standard model would say there was one electroweak particle, not that there were 8 gluons.
In our own universe, the fact that electroweak symmetry breaks ensures there are 4 electroweak particles and not other combinations. There's no corresponding thing to contain gluons to individual particles, you'd need laws of physics we don't have to add that constraint.
Stopped reading after "Yet in the mathematical equations that define the Standard Model, the eight gluons are distinct from one another in the same way that the W and Z bosons differ."
W and Z bosons, photons, etc have fixed masses, charges, interaction strengths with other particles. These properties can exactly be listed and looked up in a table of elementary particles with discrete rows.
Gluon color is continuous property in a vector space. Gluons can have any color in that space, with any combination of the 8 basis vectors (and that choice of basis is also completely arbitrary). The color |g1> is no more valid than the color (|g1> + |g2> + |g8> / √3) or any other of infinite combinations.
Calling this "8 gluons" is like saying there's "3 photons" because they can have momentum in 3 dimensions. If you want to argue there's infinite kinds of gluons, go ahead, but there aren't 8.
"Elastic" in economics happens to refers to how elastic the supply/demand is when the price changes (not vice versa, as you're describing). So e.g. an inelastic demand means the quantity demanded changes very little when the price doubles.
It doesn't solve scalping, it solves putting everyone in a Red Queen race[0] against the scalpers.
Overdiagnosis will be a major problem long after we have the data.
It's just hard convince people with a general feeling something's wrong and a specific picture of something wrong that the two are almost certainly unconnected.
Another problem is that there haven't been natural experiments in low dose exposures the way there unfortunately have been for high dose exposures.
LNT is the null hypothesis. No one disagrees a linear model fits the data very well in high doses. If you want to argue that model doesn't work in low doses, you need a model with more parameters and sufficient data to fit it. The issue is that, at these low doses we want to differentiate, we're also looking at effect sizes that are hard to separate from noise, and sampling biases that are hard to erase. There's still lively and ongoing debate.
So, on the one hand, this is interesting! Reducing radiation from CT scans is a noble cause on its own. If on top of that it could make tomography cheaper and easier, you could imagine getting earlier detection of aneurisms, fibrosis, cirrhosis, thrombosis, stenosis, even plausibly cancerous masses (along with plenty of over-detection).
On the other hand, nothing here substantiates this promise. We've got a video render of what a hypothetical device could look like. It's probably more than nothing (they got exclusive license on these butterfly chips in 2025, and it's at least plausible that the best solution to the data bottleneck in an absurdly noisy system like this is real-time AI image processing)... But it's certainly less than something. It's a hype video that doesn't prove feasibility of anything, yet.
EDIT: This is all in reaction to the second video on the announcement post[0], which is much more informative than anything on the page currently linked.
I will never pay the "normal" API token price for it.
Not until June 22 you won't!
TBF I think it's just a remark on the upvotes. It's a perfectly cromulent comment with no business being at the top of an AMA.
Your first 2 to 3ish paragraph (explanation and rephrasing) is very characteristic of an AI
What you're seeing in those paragraphs is patience and empathy. Let's not let LLMs have a monopoly on these.
Zeros are just sold out everywhere through, no?
Yeah, this made it basically clickbait for me, in terms of time I wasted with the wrong expectation.
The lack of downvotes on posts on HN has always felt like more of a bug than a feature to me.
Subagents are a helluva drug.
As the 36 millionth person building something similar, I wish[0] there were better info out there on what works well. I can understand why, but it's still frustrating seeing on the one hand how deeply helpful the flywheel of this type of setup can be, and on the other hand how every blog post stops at some incredibly superficial setup to help them write more blog posts.
[0] Yes this is a plea, if anyone has the good stuff
The Trustees of Reservations, in the Boston area, are a great example of this working well.
If it were, I can in theory see situations where improving content cleanliness is worth blowing away the KV cache.
But I absolutely can't see how feeding the entire context into a more expensive model multiple times per task, just to propose context edits that might indirectly help, could ever be worthwhile.
Gene drives are such amazing and such frightening technology. No one puts them in the same conversation as nukes or engineered pandemics, but they share the same pattern of "technology improvements giving smaller and smaller actors globally reaching powers", and have even more potential for consequences not intended by those actors. It's pretty scary to imagine a world where one lab or one rich farmer has the power to (after a few dozen generations) globally make arbitrary edits the DNA of entire species. Even smart and goodhearted people can screw up that world.
So from this armchair, I'm glad to see that at least for aedes aegypti (which seems like the clearest case for deploying a gene drive), there's an alternative like debug.
I'm having a hard time squaring this view with the prediction markets. 500k+ of volume on Polymarket[0] had Elon losing at only even odds until March, 2:1 until May 15th, and 3:1 before the verdict dropped.
It seems to have been widely reported from the start[1] and throughout[2] that the statute of limitations was a key thing Musk's team had to prove. If it was so clear, why did people think this case had legs?
[0] https://polymarket.com/event/will-elon-musk-win-his-case-aga... [1] https://www.reuters.com/legal/litigation/musk-lawsuit-over-o... [2] https://www.nytimes.com/live/2026/05/14/technology/openai-tr...
I just open photos.google.com and grab them. No need to fiddle on my phone.
When on wifi, the photo backup upload starts immediately. If it doesn't (possibly due to your settings, this used to be my issue) you can manually open the photos app and tap the backup now button.
I think the answer is practice, for a few reasons. One is obvious: conversation is a skill. Just like a novice chess player can spend 5 seconds figuring out which squares the knight can move to while an intermediate player spots a fork to force trading a strong bishop or exposing an overworked queen, exposure to similar situations rewires your brain to work faster in those situations.
Another reason, though, is to me one of the main benefits of social interaction in the first place: The brain rewiring also makes you think about what other people would think, want to hear, say to you, etc, even when they're not around. That sure can give you better answers in conversations, but more importantly, I think this is just genuinely a nice way for the brain to be. In the same way that dogs are happy playing fetch, humans are happy living with other people in mind. Maybe because it feels like not everything is your responsibility, or that you worry less about what you should be doing, or that you look forward to laughing about disasters later... I'm not entirely sure. Whatever it is, it's nicer than the alternative.