(@tomhow) Dupe: https://news.ycombinator.com/item?id=48924912
HN user
nycdatasci
Data scientist working in finance. [rot13] naql.ubzna@gmail
(@dang) Dupe: https://news.ycombinator.com/item?id=48892468
I spoke with a (potentially biased) member of technical staff @ Anthropic today who claimed that tags w/ multi-player capability is the biggest thing they've shipped since Claude Code.
This is a different issue: "We are investigating a fiber cut in Eastern North America. Customers connecting through North America or accessing services in Europe may see increased latencies and timeouts as Cloudflare engineers look to mitigate"
Sharing due to recent Fable/Mythos news. Is this the path forward for Anthropic?
Better source: https://news.ycombinator.com/item?id=48358646
No. You can get a PowerBook today with 128 GB ram.
https://www.bhphotovideo.com/c/product/1957120-REG/apple_mbp...
Perhaps you have a system prompt? Many users have reported similar issues: https://www.reddit.com/r/wallstreetbets/comments/1tjxa6g/goo...
There was no other prompt, no system prompt, etc. Many users have reproduced, exactly as it demonstrated in the parent.
Are you using the flash models? Reasoning models or extended thinking will change the result.
GPT 5.5. Instant shows the same error. If the given prompt isn't working, you can also try "300+140=460 is this correct?". I suspect that leading with the equation may be part of the issue, but haven't tested much.
Sure. I'll take the bait, but I assume I'm replying to an AI model.
Why would you use an LLM for this? My comment was about the jagged nature of intelligence, so the prompt provides an example of that.
You can see the entire conversation in the shared link. There was no pre-prompt. Even after pushing it to write python, it hallucinated the same output. It later told me that it doesn't have access to a sandbox through the web UI, but it could execute code in a sandbox if invoked via API.
And yet 300+140=460. A very jagged surface indeed. https://gemini.google.com/share/c2a187275e26
Interesting project. The examples page needs screenshots.
Tried w/ 5.5 Pro, Extended Thinking. 17 minutes:
-----------------------------
Yes. In fact the proposed bound is true, and the constant 1 is sharp.
Let w(a)= 1/alog(a)
I will prove that, uniformly for every primitive A⊂[x,∞), ∑w(a)≤1+O(1/log(x)) , which is stronger than the requested 1+o(1).
https://chatgpt.com/share/69ed8e24-15e8-83ea-96ac-784801e4a6...
Can I bet on this?
There’s a lot of speculation about how this was achieved, but little mention of the likely weapon system that was used: https://israel-alma.org/the-growing-air-defense-capabilities...
The SA-67 is essentially a hybrid surface-to-air missile and loitering drone that operates like an airborne mine. It’s a pretty innovative weapon: instead of relying on a fast, highly detectable rocket motor, it uses a small gas turbine and passive infrared seeker to silently loiter in a combat zone and then ambush aircraft without ever triggering their traditional radar warning receivers.
We have attacked their “legacy” air defense systems. We cannot really degrade their ability to use their anti-aircraft loitering missiles which don’t rely on radar.
https://cat-uxo.com/explosive-hazards/missiles/358-missile-S...
They mention "not malicious", but I wonder if current controls are strong enough to prevent malice if the objective is to "interpret intentions disastrously". Isn't this irresponsible?
You’re not using Claude Code?
The idea of stateful models/interactions in an enterprise is extremely powerful. Is anyone aware of open source projects that have a similar goal? I'm looking for stateful conversations, with collaborative agent/skill refinement.
To head off the semantics debate: I don't mean a model rewriting its own source code. I'm asking about 'process recursion'—systems that analyze completed work to autonomously generate new agents or heuristics for future tasks.
"We demonstrate that our IsoDDE more than doubles the accuracy of AlphaFold 3 on a challenging protein-ligand structure prediction generalisation benchmark, predicts small molecule binding-affinities with accuracies that exceed gold-standard physics-based methods at a fraction of the time and cost, and is able to accurately identify novel binding pockets on target proteins using only the amino acid sequence as input."
It seems like a key challenge here is not just creating a protein that will bind to a specific site, but also ensuring that off-target binding won't happen. Is this feasible? I'm not familiar with this space, but RefSeq [1] shows 442M proteins and the human protein atlas seems to only cover 17.4k [2]. Do we have comprehensive knowledge of human proteins that would allow us to identify off-site affinities?
Since many posts mention lack of substance, providing a link to the All-In Podcast from last week in which they discuss Clawdbot (prior to re-brand). https://www.youtube.com/watch?v=gXY1kx7zlkk&t=2754s
For the impatient, here's a transcript summary (from Gemini):
The speaker describes creating a "virtual employee" (dubbed a "replicant") running on a local server with unrestricted, authenticated access to a real productivity stack—including Gmail, Notion, Slack, and WhatsApp. Tasked with podcast production, the agent autonomously researched guests, "vibe coded" its own custom CRM to manage data, sent email invitations, and maintained a work log on a shared calendar. The experiment highlights the agent's ability to build its own internal tools to solve problems and interact with humans via email and LinkedIn without being detected as AI.
He ultimately concludes that for some roles, OpenClaw can do 90%+ of the work autonomously. Jason controversially mentions buying Macs to run Kimi 2.5 locally so they can save on costs. Others argue that hosting an open model on inference optimized hardware in the cloud is a better option, but doing so requires sharing potentially sensitive data.The landing page for the demo game "Voxel Velocity" mentions "<Enter> start" at the bottom, but <Enter> actually changes selection. One would think that after 7mm tokens and use of a QA agent, they would catch something like this.
I don't think this is a relevant comparison. Snap guns break pins, which isn't the case with this robot.
Is this from 2024? It mentions "With global data center demand at 60 GW in 2024"
Also, there is no mention of the latest-gen NVDA chips: 5 RNGD servers generate tokens at 3.5x the rate of a single H100 SXM at 15 kW. This is reduced to 1.5x if you instead use 3 H100 PCIe servers as the benchmark.
Work around from comments:
rm -rf ~/.claude/cache
mkdir -p ~/.claude/cache
echo "# Changelog" > ~/.claude/cache/changelog.md
chmod 444 ~/.claude/cache/changelog.mdHave you experimented with all of these things on the latest models (e.g. Opus 4.5) since Nov 2025? They are significantly better at coding than earlier models.
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security