Refilling the context of the cheap executor == just switching the model mid-conversation instead of /clear’ing and passing some plan doc?
HN user
swingboy
On the first syntax example: there’s something funny to me about using three pipe operator and four different functions to turn “hello world” into “Hello World”.
This doesn’t look as fun as a Tamagotchi.
What are some of the things you’re doing with the Codex app-server?
The `repo_path` field.
Well, it looks like he was running the agent in his home directory to begin with considering the `repo_path` field is exactly that.
Is it because Pi’s default system prompt is so simple?
Haven’t heard this one in a while XD
Which 2 or 3 skills?
There was a big thread about it here the other day. https://news.ycombinator.com/item?id=48734373
Serious question in good faith: what was the deal with the “calamari” (clots?) the anti-vax crowd kept talking about being found in the veins/arteries of folks who took the Covid vaccine?
“Anthropic wants to MAKE A DEAL!”
These guesstimate release dates seem…soon. Usually the Polymarket markets are pretty accurate and really accurate when an insider at one of the companies puts a bet.
How is this enforceable?
And it’s a paper from Alibaba researchers, the company/lab that Anthropic called out by name.
I’ve always assumed any LLM output that was some type of rating or score was bullshit. Unless the LLM writes a Python script to calculate the score (and even then…) then the score it outputs is just the next most likely token, taking into account temperature and what not.
You see a lot of frameworks for things like spec-driven development make use of scoring how good the spec/design/plan is and it’s like, uhhh…
Did he not listen to music made by others when he was isolated in the cabin?
All members of the Epstein class.
Nice try, Barbra.
How would export controls apply if OpenAI or Anthropic released a model as open weights? Not that they would, but asking out of curiosity.
Because the American people weren’t outraged enough to push them out.
Here’s to hoping that Alibaba (and other Chinese labs) have collected some really good distilled data.
My work involves asking LLMs about both Tianenmen Square and what’s going on in Gaza, so I can’t use Chinese or American models!
I get a 500 when clicking “Explore the Models”
Anthropic’s best practices still include the use of XML: https://platform.claude.com/docs/en/build-with-claude/prompt...
*Advice only applies to neighborhoods without an HOA.
Does “the best machine for AI use” apply here considering these models are still server-side?
GPT5.5 xhigh seems to benchmark about on par with Mythos for cybersecurity.
The Trump administration would never do anything to manipulate the markets. /s
I realize these models are locked up pretty tight and terabytes in size, but in a future like that, I don’t see them not being leaked via an insider. The weights have to be loaded into VRAM at some point.