HN user

Philpax

8,329 karma
Posts198
Comments1,084
View on HN
luau.org 20h ago

Luau Type Functions

Philpax
1pts0
github.com 20d ago

crustc: entirety of `rustc`, translated to C

Philpax
386pts92
www.baseten.co 28d ago

How we built the fastest API for GLM-5.2

Philpax
4pts0
store.steampowered.com 1mo ago

Steam Machine 512GB

Philpax
34pts1
www.anthropic.com 1mo ago

Claude Fable 5

Philpax
2626pts2160
openai.com 1mo ago

Building the Infrastructure for the Intelligence Age in Michigan

Philpax
2pts0
www.anthropic.com 1mo ago

Anthropic Cofounder Chris Olah's Remarks on Pope Leo XIV's "Magnifica Humanitas"

Philpax
87pts99
github.com 2mo ago

talkie-coder: From 1930 to SWE-bench

Philpax
2pts0
www.blender.org 2mo ago

Anthropic Joins the Blender Development Fund as Corporate Patron

Philpax
256pts218
huggingface.co 2mo ago

XiaomiMiMo/MiMo-v2.5

Philpax
3pts0
store.steampowered.com 2mo ago

Steam Controller: It's almost here

Philpax
32pts5
github.com 3mo ago

LLM on a 1998 iMac G3 (32 MB RAM)

Philpax
3pts0
www.youtube.com 3mo ago

Mirror's Edge Early Prototype (Feb 7, 2008) [video]

Philpax
2pts0
www.theguardian.com 3mo ago

Artemis II lifts off: four astronauts begin 10-day lunar mission

Philpax
280pts1
www.theregister.com 3mo ago

Linux kernel czar says AI bug reports aren't slop anymore

Philpax
14pts1
www.percepta.ai 3mo ago

Constructing an LLM-Computer

Philpax
1pts0
epochai.substack.com 4mo ago

First AI Solution on FrontierMath: Open Problems

Philpax
4pts0
www.githubstatus.com 4mo ago

GitHub: Degraded Performance for Various Services

Philpax
3pts0
techcrunch.com 4mo ago

Anthropic's Claude rises to No. 2 in the App Store following Pentagon dispute

Philpax
42pts2
www.vectorware.com 5mo ago

Async/Await on the GPU

Philpax
228pts55
lpalmieri.com 5mo ago

Can agentic coding raise the quality bar?

Philpax
1pts0
simonwillison.net 5mo ago

Deep Blue

Philpax
35pts13
news.ycombinator.com 5mo ago

Tell HN: Yet Another Round of Zendesk Spam

Philpax
11pts3
store.steampowered.com 5mo ago

Steam Hardware: Launch timing and other FAQs

Philpax
99pts51
news.ycombinator.com 5mo ago

Tell HN: Another round of Zendesk email spam

Philpax
105pts54
www.anthropic.com 5mo ago

Claude on Mars

Philpax
28pts4
zulip.readthedocs.io 6mo ago

Zulip AI use policy and guidelines

Philpax
2pts0
www.interconnects.ai 6mo ago

Get Good at Agents

Philpax
1pts0
over.world 6mo ago

The Path to Real-Time Worlds and Why It Matters

Philpax
3pts0
codeberg.org 6mo ago

Open Slopware

Philpax
6pts2

Over the last two years, this weight class has doubled its scores and/or saturated several benchmarks in the Qwen lineup alone without loss of generality: https://claude.ai/public/artifacts/9f249169-3623-417e-86cd-7...

There is undoubtedly a limit somewhere (there is only so much you can pack into a given size) but it's really not particularly clear where that limit is. I don't think it's superintelligence - that much I agree with you - but I think "We already have a 1gb model that is as capable as it will ever be" is strictly false.

This is science fiction, these models don't have access to their own weights.

The models are being used to train, and improve the infrastructure for training, other models [0][1]. Several RL techniques rely on using the currently-being-trained weights as part of their process. I really would not take "don't have access" as a given, especially during the training phase.

What would be a lot more scary is a model as capable as sol that's able to run on consumer hardware without taking up several terabytes of storage, but of course that is simply not possible as we need 4t parameters to even begin emulating a small fraction of what a human brain can do.

The Poolside Laguna S 2.1 model [2] purports to compete with models several times its size, and inference compute is becoming increasingly plentiful. Again, would not hold anything here as a given.

[0] https://openai.com/index/gpt-5-6/ ("GPT-5.6 accelerates OpenAI")

[1] https://www.kimi.com/blog/kimi-k3#coding

[2] https://poolside.ai/blog/introducing-laguna-s-2-1

Annoyances around the general workflow (i.e. having to dedicate an entire output to it and/or preventing you from doing anything else with that output), as well as general stability/connectivity issues (i.e. not streaming one of [video|audio|input] on connection). Interacting with its UI also makes me believe that it is in dire need of some TLC.

Comparatively, Moonshine "just worked". Claude set it up, porting across relevant settings from Sunshine, I connected to it, punched the client pairing code into the server, and then I was able to play Clair Obscur on my TV without being logged into my workstation. Pretty cool!

Oh, I have no doubt that they could have extracted those gains from Zig! My point is more that, from a relatively naive line-to-line port, they were able to claim these benefits without much effort.

It's not great for Zig if you have to put in more work to end up at the same place efficiency-wise, especially for a language marketed at people who like to get the most out of their metal.

Without commenting on Bun itself as a project, or the nature of the rewrite, it can't be good for Zig that a naive rewrite away from it fixed memory leaks, improved stability, shrunk binary size by 20%, and improved performance by 5%.

The Steam Deck is the closest thing the PC world has to a console (barring the Steam Machine, of course), and features near-console levels of hardware/software integration.

I am _very_ familiar with Claudish, and to some extent, the other AIs' writing styles. This article is human-written and features human writing quirks.

The very first sentence

Back in 2022 and 2023 there were two big branches of machine learning happening at Meta.

is unmistakably human. That's not how a LLM would phrase this sentence, and if it did, it would have put a comma after 2023.

In looking at the code that the LLMs have produced for the project, especially given the pretty massive and widespread architectural changes needed to make the implementation libified and memory safe, we decided that the codebase is not a derivative work that would require carrying forward the GPL license and have decided to release the code under the MIT instead.

Hmm. That's going to be interesting.

Claude Fable 5 1 month ago

"Releasing a model this capable comes with risks. Without safeguards, Fable 5’s capabilities in areas like cybersecurity could be misused to cause serious damage. We’ve therefore launched the model with safeguards that mean queries on some topics will instead receive a response from our next-most-capable model, Claude Opus 4.8. To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessions. With more capable models arriving in the coming months, we’re working to improve our safeguards and reduce false positives as quickly as we can.

For a small group of cyberdefenders and infrastructure providers, we’re also launching Claude Mythos 5. It’s the same underlying model as Fable 5, but with the safeguards lifted in some areas.2 Mythos 5 will initially be deployed through Project Glasswing, in collaboration with the US Government, as an upgrade to Claude Mythos Preview. It has the strongest cybersecurity capabilities of any model in the world. Soon, we intend to expand access to Mythos 5 through a broader trusted access program."

The Chinchilla scaling laws give you a minimum for the number of tokens you should be using for a given size: if you can't meet what they suggest for that size, you should shrink the size, as, otherwise, the capacity of the model is going to waste.

I do agree that it is a datapoint, but GP's point is that this model was undertrained, so it's hard to draw the same conclusions from it that we would from other research.

I agree with the others - I'm sure that you've provided your own input, but Claude's writing and design style is so overwhelmingly dominant that those who have spent time with it can immediately recognise it, and it makes it hard to take at face value that you were the primary author, even if you were.

For your workflow, I'd suggest drafting with a LLM to help you find the right balance of content, and then throwing all of that out and writing it yourself. Otherwise, it won't sound like you.