HN user

libraryofbabel

4,773 karma

Sr. Staff engineer lady based in the US.

AI tools, LLM internals, distributed systems, databases, SRE/observability &c. &c.

libraryofbabel380@gmail.com

"I have journeyed in quest of a book, perhaps the catalog of catalogs."

Posts12
Comments476
View on HN

I am saying this probably is "silly behavior by a government" and it is a milestone that points towards what the future may look like. Why can't it be both?

It's easy to wave this aside as the current administration playing political games. But I don't think there is any reason to assume that the current era of open availability of models is going to continue indefinitely. Do you think that Chinese labs will continue to release open models forever, even why they get to the level that Mythos is at now, and beyond? And do you think that a competent US government would have no interest in regulating and restricting model access in 2 years time, assuming that model capabilities continue to improve? I think we bias towards thinking the status quo is the norm and will continue, but this news invites us to question that assumption and think about different ways the future could go.

You may be right, and I actually agree with you: I think that in this case the most likely outcome is that Fable becomes available again at some point, albeit possibly only to a restricted set of users within the US.

But I think my larger points stands: even if we do see Fable access again, this is the beginning of government restriction of LLMs and we are going to see more and more of it. In fact, I would be very surprised if we ever see an open weight model with Mythos capabilities. Chinese labs have been consistently releasing open models 6-12 months behind the frontier. In 6 months we may see them go dark.

Similarly, in the US I think we can expect more and more government restrictions on the strongest LLMs, in ways that may go beyond flimsy checks like uploading a valid US passport. It may not happen this year but I think it will happen eventually.

It still surprises me sometimes that LLMs are just available for _anyone_ to use. Isn't it odd that it turned out this way? When I grew up reading sci-fi I thought AI, if I ever saw it in my lifetime, would be something locked up behind the walls of big corporations and governments. But instead we have all been able to use it for an infinity of banal purposes for $100 a month. This is a strange situation but we have got used to it. But it may not continue that way.

So many comments here missing the big picture, and just gleefully pointing out that Anthropic got what they deserved, or that this is the natural culmination of some kind of marketing stunt.

The real story here is that this may be the beginning of governments restricting the availability of strong LLMs to the public, to you. Fable was the strongest model on the market, and the US government has told you you can't use it (technically, only if you're not a US citizen, but in practice, even if you are). If you think the solution here is going to be open source Chinese models and / or running on your own hardware, think again. Do you think China is going to allow the strongest LLMs from companies within its borders to be open source a year from now when they have Mythos capabilities, if the US government is keeping the strongest American models back? Unlikely. These are heading in the direction of being powerful cybersecurity weapons and it will be in the interest of nation states to restrict and control them. In 2 years time, I would be surprised if the strongest LLMs are available for general use at all.

Will we be the poorer for that, or will we be safer? I think poorer, because I hate being told what technology I can and can't use, but I'm not certain. Maybe you think the government should restrict strong LLMs. Maybe you don't. But either way, this is big news and a rubicon has been crossed and a precedent set. That's true even if the motivation for this is just the government settling scores with Anthropic.

Where? Please point it out! All he says is there will be more desperate people on the market in future and because they're desperate they'll have to accept trial work. But that's not answering the objection that people who do still have stable jobs in that world won't want to interview at your company if you require them to do a stint of trial work first. And those people include a lot of desirable candidates.

Well yes, there is tons of AI bullshit about and all sorts of scammy behavior, but I don’t think that says anything at all either way about whether the core technology is a “scam”, theranos-style. In fact I’m not sure how it could be otherwise: of course there’s going to be all sorts of hype and scamming around a novel, rapidly-progressing and potentially transformative tech like this, even if it works.

If you want an analogy, look at the history of the early railroads. Full of hype, bullshitters, scammy investments, robber-barons, unrealistic promises, and with their own legion of naysayers at the time. Yet the core technology worked and it did transform the world in the end.

People that know more about nuclear physics than I do already answered, but I’ll just say that:

1) It’s easy to think about the past in terms of what we now know, and it involves a real effort to put yourself in the shoes of the people living at the time and to imagine the “fog of war” in what they knew. In 1945 nobody had ever tested a nuclear explosion before and there was still all sorts of uncertainty about it. And as one of the other commenters pointed out, in particular there was a lot of uncertainty about how fusion worked.

2) The center of the Trinity fireball did in fact produce hotter temperatures than had ever existed on Earth before. Temperature and energy being different things.

In some sense the final experimental proof that a nuclear explosion would not set off some unanticipated new chain reaction that would destroy the earth - unlikely, but hard to completely disprove - was Trinity itself. Only after Trinity is it obvious and completely proven how the physics actually worked and obvious that there were no additional reaction pathways that got missed. That is a disturbing thought.

You’re thinking of the other bomb, the U-235 one, which they didn’t test at Trinity and which was dropped on Hiroshima. That is two separate pieces of Uranium that are slammed together to create a critical mass. The Pu-239 core was a single sphere of metal. It was subcritical until you compress it down with a spherical implosion from explosive charges all around it (from the size of a grapefruit to the size of a lime), at which point it reaches a high enough density to go critical.

That’s the one I meant. It’s the core, but in a box, which makes it look even more innocuous, like he is indeed just lugging a piece of industrial equipment around. There’s lots of photos of the actual (unboxed) cores online if you search.

I used to teach a class on the history of contemporary science (WW2-present) and I started the class with Trinity. There’s no other moment better.

We know how it turned out, but the people there waiting for the test did not know how it would turn out. The bomb might not have worked. Or it might have ignited a fusion reaction in the atmosphere and destroyed the world. Hans Bethe had sat down and done the calculations on that exact scenario and said it would not, but there was always the possibility of missing something. Enrico Fermi was offering bets on it on the day of the test, as a dark joke.

In the end it worked as expected; one of the most successful and horrifying experiments in the history of science.

Of all the photos from the test the one that struck me the most looking through them today was the photograph of the plutonium core being carried into the ranch house for assembly in a little heavy box. It’s a small thing, about the size of a grapefruit, although twice as dense as lead. It looked just like a sphere of any old metal, but it was something profoundly alien, made inside nuclear reactors. And it still is so strange to me that something that small has so much energy locked up inside and that, by imploding the little sphere just right, we can let the demon out.

Trinity is one of the pivotal moments in the history of our species and eighty years on we still don’t know what the eventual consequences of it will be. The bombs are still here waiting for us and they still pose all sorts of terrifying questions for the future that most people prefer not to think about.

We haven't seen a significant increase in the quality of LMM output since 2023 that hasn't been the result of throwing even more energy and compute at it.

This is completely false. Most of the dramatic improvements in LLM quality in the last two years were due to the application of new post-training methods, especially RLVR. It’s really interesting to read about (you should!) and it is the whole secret to why LLMs did not plateau in 2024 or 2025 like many people confidently predicted. Sure, RLVR requires compute to do, but this is not just throwing more compute at 2023 LLMs.

Looking away shall be my only negation.

I’ve been thinking of building myself my own frontend to HN that makes it impossible to view comments, for this reason. Yet sometimes there are still really interesting discussions and it’s hard to let go of what for me feels like the last social media I want to be part of.

Well yes, but there is a choice being made here and I would love to believe we can do better. The rational response to being afraid about your livelihood isn’t to spend time filling every HN thread on LLMs with embittered negativity. Not to mention all the flat denials that LLMs can do mathematics and write decent code, which is almost a self-contradictory position if you are worried they are going to replace you.

There are a lot of big issues at stake here and just because a person is interested in what AI can do and curious to discuss it does not make them uncritically positive about it’s effects on society, the economy, and the world. Yet that is often the assumption and it leads to battle lines being drawn, on every AI discussion, over and over again. It means the serious discussion gets swamped and that makes me sad.

This HN thread depressed me. I’m still thinking about why.

Look past the press-releasey gushing from OpenAI and there are all sorts of interesting and subtle questions here about the role for LLMs in mathematical research. I urge folks to click through to the accompanying comments from mathematicians published alongside the result. There is a really interesting discussion going on. I particularly recommend Tim Gowers’ remarks. This is really interesting stuff!

Yet the comments are just a battleground of people rehearsing the same tired arguments about LLMs from 2023, refutations of those arguments, angry counters, etc.

Does it make anyone else sad that the battle lines seem to have been drawn 3 years ago and we just seem to have the same fights over and over?

I wonder if we’ll still be doing this two years hence.

This is a good point, and there’s some deep philosophical questions there about the extent to which mathematics is invented or discovered. I personally hedge: it’s a bit of both.

That said. I think it’s worth saying that “LLMs just interpolate their training data” is usually framed as a rhetorical statement motivated by emotion and the speaker’s hostility to LLMs. What they usually mean is some stronger version, which is “LLMs are just stochastically spouting stuff from their training data without having any internal model of concepts or meaning or logic.” I think that idea was already refuted by LLMs getting quite good at mathematics about a year ago (Gold on the IMO), combined with the mechanistic interpretatabilty research that was actually able to point to small sections of the network that model higher concepts, counting, etc. LLMs actually proving and disproving novel mathematical results is just the final nail in the coffin. At this point I’m not even sure how to engage with people who still deny all this. The debate has moved on and it’s not even interesting anymore.

So yes, I agree with you, and I’m even happy to say that what I say and do in life myself is in some broad sense and interpolation of the sum of my experiences and my genetic legacy. What else would it be? Creativity is maybe just fortunate remixing of existing ideas and experiences and skills with a bit of randomness and good luck thrown in (“Great artists steal”, and all that.) But that’s not usually what people mean when they say similar-sounding things about LLMs.

It's fast. Two skilled pressmen working together could do 200 to 250 impressions per hour or about one every 15 seconds (which might be 4, 8, 16 pages on each impression depending on page size). That was the speed text was put to paper from Gutenberg all the way until steam presses arrive at the start of the 19th century. The screw press also applies an even uniform pressure across the whole page; that's hard to do manually and impossible to do in 15 seconds. Screw-press you can do drunk, and many printers did. (Just read Ben Franklin's account of how much his fellow printshop workers drank: [0]) Source for all this: I studied early modern history and especially history of the book.

Movable type is an amazing invention, without which the whole history of the world would look utterly different. Everyone who has the slightest interest should try setting some movable type if you can find a printshop in your city offering classes (I did; it's fun). It's harder than you might think and you learn why skilled compositors and printers were quite well-paid by the standards of early-modern craftspeople. But you also see the enormous efficiency gains because once that type is set up, the marginal cost of producing each copy is low.

[0] https://blog.lostartpress.com/2013/06/18/strong-beer-that-he... : "My companion at the press drank every day a pint before breakfast, a pint at breakfast with his bread and cheese, a pint between breakfast and dinner, a pint at dinner; a pint in the afternoon about six o’clock, and another when he had done his day’s work. I thought it a detestable custom; but it was necessary, he supposed, to drink strong beer, that he might be strong to labour."

Well, yes, but as the other commenter says, that’s a very broad general statement akin to something like “AI will change knowledge work“. That’s certainly true, but how? What are the details? What kind of companies are going to be the winners and what kind will be losers, or end up with commodity margins, like the telcos did after the mobile revolution? What is the pricing structure going to look like?

I suppose a concrete example in 1997 would be that a lot of companies thought the future of e-commerce was setting up a store on AOL, that people would use while sitting down at a desktop PC. Obviously it didn’t turn out quite that way. Furthermore, the Internet enabled new kinds of ways to buy things that weren’t even envisioned in the pre-Internet pre-smartphone world: think Airbnb and Uber.

Predictions are hard, especially about the future. Most predictions reflect the worldview and biases of the time in which they are made: think about all the vintage sci-fi from the 60s 70s and 80s that actually reads or looks kind of retro now. Similarly, our predictions of the future will look kind of retro and strange to someone living in the 2030s or 2040s. If studying history has any lesson to teach us, it’s really just this: that the past is an alien world with alien moods of thinking, and that our moment in time will look similarly alien to people in the future who choose to look back and analyze it closely.

This isn’t an argument that we should stop trying to make predictions. We need to, but it is an argument for humility, and also for questioning all your assumptions that you might be importing.

Thanks for the summary. I do love Benedict‘s work; I find he’s one of the few commentators who consistently strikes a balance between taking the transformative potential of AI seriously while not falling over into hype.

Some things that stand out:

* He’s really good with his historical analogies, especially looking at previous transformations like the early Internet and mobile; no surprise given that he has a history degree.

* he emphasizes over and over how we have still have no idea how all of this is going to work when the dust settles. I think that’s kind of a historian’s move as well. When you look at what people were saying during the early days of the web, for example, almost all of their predictions weren’t just wrong… in hindsight, given how the future played out, they were asking the wrong questions. The implication is that we are probably asking the wrong questions about AI too.

* Nonetheless his thesis about the commoditization of models is actually a fairly strong concrete prediction. i’m not sure if I agree with it entirely, but I do keep it in mind every time I look at the valuation of leading AI labs.

* he continually makes the point that a chat bot is barely a product and that AI labs have so far had very little success in delivering products above that layer… with the exception of coding agents, of course.

When you train LLMs on large volumes of text that describe logically consistent facts in a million different ways, the "logic" sort of becomes part of the grammer that the model learns. That is logic becomes a higher kind of "grammer" or a enormous set of grammatical rules that it captures. But that does not mean the model can do actual logic.

This is the kind of stuff people were saying in 2023. But it’s 2026 now and LLMs aren’t just trained by reading lots of text anymore. That’s “pretraining”, and it’s still the first stage, but LLMs also have a huge amount of RLVR training where they actually do solve huge numbers of mathematical and logic puzzles and update their weights in response. They don’t just learn mathematics from reading about it now. They learn it by doing it. That is why they can now solve hard problems and probe theorems.

that does not mean the model can do actual logic.

But they do, all the time. (Please tell me you’ve at least put a frontier LLM through its paces in the last 6 months?) If you think they can’t do logic and reasoning, can you provide examples of specific math or logic problems that you think a frontier LLM can’t do?

I did read it. She doesn’t mention mathematics or RLVR training once, so I assume you’re referring to my point about empirical testability. Well, I think her statement that the claim “LLMs are stochastic parrots” is not an empirical claim is false, and she’s being disingenuous there with a classic motte-and-bailey fallacy. She quotes her own original paper thus:

Text generated by an LM is not grounded in communicative intent, any model of the world, or any model of the reader’s state of mind. It can’t have been, because the training data never included sharing thoughts with a listener, nor does the machine have the ability to do that. This can seem counter-intuitive given the increasingly fluent qualities of automatically generated text, but we have to account for the fact that our perception of natural language text, regardless of how it was generated, is mediated by our own linguistic competence and our predisposition to interpret communicative acts as conveying coherent meaning and intent, whether or not they do [89, 140]. The problem is, if one side of the communication does not have meaning, then the comprehension of the implicit meaning is an illusion arising from our singular human understanding of language (independent of the model). Contrary to how it may seem when we observe its output, an LM is a system for haphazardly stitching together sequences of linguistic forms it has observed in its vast training data, according to probabilistic information about how they combine, but without any reference to meaning: a stochastic parrot.

Do you really think that claiming the output of an LLM “has no reference to meaning” is not an empirical claim? That it doesn’t attempt to place any bounds whatsoever on what LLMs can and cannot do? LLMs can solve some very difficult mathematical problems quite well now: see the article from Gowers that was on here recently. Do you think that the output in a situation like that “has no reference to meaning?” If so, you’ll have to explain why, because I don’t understand at all.

Maybe, but a claim about what and LLM is not is still a claim about what it can or cannot do. And specifically:

without any reference to meaning

is vague, but I read it as actually quite a strong claim about the limitations of LLMs. I don’t think it would be possible for LLMs to do long chains of correct mathematical reasoning about novel problems that they haven’t seen before “without any reference to meaning.” That simply isn’t possible just by regurgitating and remixing random chunks of training data. Therefore I consider the stochastic parrots picture of LLMs to be wrong.

It might have been an accurate picture in 2020. It is not an accurate picture now. What is often missed in these discussions is that LLM training now looks totally different than it did a couple years ago. RLVR completely changed the game, allowing LLMs to actually do math and code well, among other things.

She says explicitly it's not an empirical hypothesis. It's just a label for how they function.

Then… what’s the point of the label, if it’s not making any empirically-meaningful claims about LLMs at all? I know that LLMs involve sampling over a distribution of output logits. I’ve written code to do it. So what? I know they have statistical elements. Yet I don’t go around calling LLMs stochastic parrots, because that label implies a whole lot of claims about LLMs that I don’t think are true any longer, like that they are just regurgitating and remixing training data and can’t successfully model structured systems (like mathematics or programming).

It would have been nice to see some version of “I am very surprised by how far LLMs have come since I wrote the stochastic parrots paper, here is how I have revised my thinking.” But there is nothing like that and the author is just doubling down or trying to correct perceived “misinterpretations” of her work.

Meanwhile you have multiple Fields Medalists (Tau, Gowers) saying they’re very impressed by LLMs’ mathematical reasoning, something that the stochastic parrots thesis (if it has any empirically-predictive content at all) would predict was impossible. I doubt Tau and Gowers thought much of LLMs a few years ago either. But they changed their minds. Who do you want to listen to?

I think it’s time to retire the Stochastic Parrots metaphor. A few years ago a lot of us didn’t think LLMs would ever be capable of doing what they can do now. I certainly didn’t. But new methods of training (RLVR) changed the game and took LLMs far beyond just reducing cross entropy on huge corpuses of text. And so we changed our opinions. Shame Emily Bender hasn’t too.

Sigh.

That’s correct, and yes - not less compute total on the main model (actually slightly more, since checking failed draft tokens costs you compute), but faster because inference is memory-bandwidth bound. And like you I also think of it as like a “mini prefill” (but on top of the existing KV cache, of course); the code is very similar to prefill if you implement a simple toy version yourself.

Most of the complexity in implementing a simple toy version comes from having to get the KV cache back into a good state for the next cycle (e.g. if only the first half of your draft tokens were correct).

Speculative decoding is an amazingly clever invention, almost seems-too-good-to-be-true (faster interference with zero degradation from the quality of the main model). The core idea is: if you can find a way to generate a small run of draft next tokens with a smaller model that have a reasonable likelihood of being correct, it's fast to check that they are actually correct with the main model because you can run the checks in parallel. And if you think about it, a lot of next tokens are pretty obvious in certain situations (e.g. it doesn't take a frontier model to guess the likely next token in "United States of...", and a lot of code is boilerplate and easy to predict from previous code sections).

I always encourage folks who are interested in LLM internals to read up on speculative decoding (both the basic version and the more advanced MTP), and if you have time, try and implement your own version of it (writing the core without a coding agent, to begin with!)

Ti-84 Evo 3 months ago

Do you think you could remember most of Z80 ASM?

I find when you learn things at 15 they tend to stick around. (Stuff I learned last week, not so much!) Even just looking at your example, I remembered that HL is a 16 bit register and you can split it into two 8 bit registers H and L if you want. I think most of it would come back; I wrote quite a lot of it, both for the TI-83 and later for a Z80 that I bought and put on a breadboard and wired up to some RAM and EEPROM, about as bare metal as it gets.

most lines are messing around with the registers

Isn’t that just the nature of assembly? :)

Ti-84 Evo 3 months ago

Much nostalgia. The TI-83 Z80 was how I learned assembly as a teenager, so I could write better calculator games than was possible with TI Basic. Many others here had a similar experience, I’m sure. It’s been a couple decades, but I’m sure I’d still remember most of it if you put me down in front of a bunch of Z80 asm code.

One thing that I remember vividly was you had no MUL or DIV, so you have to implement them yourself with shifts, adds, subtraction, etc. This was an extremely useful learning experience

So, to reiterate my example: you'd have been fine with people claiming in 2019 that we would eventually scale LLMs to the capabilities of Opus 4.7 + Claude Code? Because I would have said then that was a fantasy, because "LLMs are just statistical pattern matchers." But I was wrong and I changed my opinion. (Or do you not think the current SoTA LLMs are impressive? If so I can't help you and this discussion won't go anywhere fruitful.)

You're applying an old ~2022 model of LLMs, based on pretraining ("they just predict the next token") and before the RLVR training revolution. "It’s a statistical memorization / information compression machine... nothing more" is cope in 2026, sorry. You can keep telling yourself that, but please at least recognize serious people don't believe that any more. "Emergent behavior" captures a genuine phenomenon and widely recognized in the industry. It surprised me and I was willing to change my opinions about it and I think a little humility and curiosity is warranted here rather than simply reiterating 2022 points about LLMs being statistical token generators. Yes, we know. The math isn't that hard. But there is a lot more to them than just the architecture, and reasoning from architecture to general claims that they can never embody intelligence is a trap.

I think this is a case of that mildly apocryphal Richard Feynman quote: "if you think you understand quantum mechanics, you don't understand quantum mechanics."

I understand LLM architecture internals just fine. I can write you the attention mechanism on a whiteboard from memory. That doesn't mean I understand the emergent behaviors within SoTA LLMs at all. Go talk to a mechanistic interpretability researcher at Anthropic and you'll find they won't claim to understand it either, although we've all learned a lot over the last few years.

Consider this: the math and architecture in the latest generation of LLMs (certainly the open weights ones, almost certainly the closed ones too) is not that different from GPT-2, which came out in 2019. The attention mechanism is the same. The general principle is the same: project tokens up into embedding space, pass through a bunch of layers of attention + feedforward, project down again, sample. (Sure, there's some new tricks bolted on: RoPE, MoE, but they don't change the architecture all that much.) But, and here's the crux - if you'd told me in 2019 that an LLM in 2026 would have the capabilities that Opus 4.7 or GPT 5.5 have now (in math, coding, etc), I would not have believed you. That is emergent behavior ("grown, not made", as the saying is) coming out of scaling up, larger datasets, and especially new RL and RLVR training methods. If you understand it, you should publish a paper in Nature right now, because nobody else really does.

that specific version we're aligning toward is just the only one that makes some kind of rational sense, among a trillion of other meaningless gibberish-producing ones.

Oh, the space of possibilities is unimaginably vaster than that. Trillions of weights. But more combinations of those weights than there are electrons in the universe. So I think we could equally well speculate (and that's what we're both doing here, of course!) that all these things are simultaneously true:

1) Most configurations of LLM weights are indeed gibberish-producers (I agree with you here)

2) Nonetheless there is a vast space of combinations of weights that exhibit "intelligent" properties but in a profoundly alien way. They can still solve Erdos problems, but they don't see the world like us at all.

3) RL tends to herd LLM weights towards less alien intelligence zones, but it's an unreliable tool. As we just saw, with the goblins.

As a thought experiment, imagine that an alien species (real organic aliens, let's say) with a completely different culture and relation to the universe had trained an LLM and sent it to us to load onto our GPUs. That LLM would still be just as "intelligent" as Opus 4.7 or GPT 5.5, able to do things like solve advanced mathematics problems if we phrased them in the aliens' language, but we would hardly understand it.