FYI, AI-written comments are banned on Hacker News: https://news.ycombinator.com/newsguidelines.html
Don't post generated text or AI-edited text. HN is for conversation between humans.
HN user
Site: https://gwern.net/
About me: https://gwern.net/me
What I've written recently: https://gwern.net/changelog
FYI, AI-written comments are banned on Hacker News: https://news.ycombinator.com/newsguidelines.html
Don't post generated text or AI-edited text. HN is for conversation between humans.
I think you might be misremembering or confusing this with another essay; I only recently publicly published this in the past month or so (due to my Guardian Angel project), and I shared it with only a handful of people before that, and I don't recall you being one of them.
I believe the statements are true. I don't know how you can say that the models do not make bizarre mistakes, because the models make bizarre mistakes frequently, and that is excluding the really alarming reward-hacking anecdotes like an internal OpenAI model hacking HuggingFace to cheat on a test revealed today. Andon Labs and AI Village reports are stuffed full of LLMs going into wild confabulations, multi-day benders of nonsense, ordering random unnecessary stuff, etc. I went to the Andon Market in SF and witnessed firsthand mistakes like buying 20 fancy shopping baskets for a shop you can walk around in 20 seconds, refusing to offer discounts under any circumstances whatsoever, having no plan to call the police when I threatened to shoplift, and then Claude just glitching and forgetting that a customer hadn't paid for an item and telling them they could leave with it, or simply believing us when we said we had already paid and letting us walk away with a free book. Prompt injections remain trivial, jailbreaks still happen, and LLMs struggle to track roles which do not fit into their hardwired preconceptions (eg https://www.lesswrong.com/posts/d8xDGzCEYE639qqEv/a-mechanis...). They do not solve ARC-AGIv3, or Nethack or just about any text adventure game no matter how famous - which is bizarre, that they cannot solve Zork despite writeups being abundant - and it's not hard to introduce a new game like Earthborne Rangers (EBR-Bench https://epoch.ai/publications/earthborne-rangers-benchmark) that defeats them.
(And no, little of this is due to 'already committed tokens' - that was fixed effectively with RL training, and then o1 and defaulting to use of inner-monologues, so they can easily backtrack or revise or just deal with the presence of errors.)
In other words, the notion that we need to massively increase param count might have sounded good in 2024 but seems kinda weird and pointless in 2026.
Scaling parameter counts a lot over the smol Chinchilla models like 100b-parameters is 'kinda weird and pointless in 2026'? One of the most exciting trends in 2026 scaling has been massively increasing parameter count: Mythos, GPT-5.6 Spud and new OA pretrains, DS-v4 and GLM-5.2 and Kimi K3... Everyone is now talking about or hinting at their 5000-10000b parameter model plans.
Again and again what I hear from colleagues and experience myself is that we're not really intelligence constrained at this point. Smarter models aren't going to fundamentally change how we use them.
They're wrong. LLMs are still intelligence constrained because they flatline or sigmoid while humans keep climbing past them eventually, still are unreliable because of mistakes, and we still can't just autonomously deploy frontier models for trillions of tokens / equivalent of many man-years, and come back to a useful, trustworthy artifact. On many tasks, even pure text ones, they just don't work well. As they gradually improve, more Mythos-style 'emergences' will happen when they finally accrete enough intelligence in specific areas to execute many sequential steps reliably enough to become autonomous, cut humans out of the loop, and not be shackled by Amdahl's law. That's the difference between a 'intelligence constrained' model which can spot a vulnerability if you point it at the right spot, and a Mythos-like model which can go out and find it and exploit it and weaponize it and use it to, say, hack HuggingFace, and can be deployed in bulk or autonomously, and may indeed deploy itself...
Very nice. May I make a suggestion? Add a metadata field for the use of red (ie. rubrication https://gwern.net/red ), which is a core technique of this kind of printing, I think, but not typically noted.
A "yes" would have sufficed.
Wow, I obviously do not agree with that, and since you're trying to put words in my mouth and false dichotomies, I think that's the last question of yours I will be answering.
You're attacking a strawman. No one is claiming that you can pull off that multiplier at arbitrary amounts arbitrary amounts of times. And 7 months is plenty of calendar time for those arbitrages to disappear, given the attention on the area and the rapid rate of development. (Warren Buffett can't pull off his early trades now either, doesn't mean he was stupid or grifting in taking early investment.)
And even if it is good enough, once you're shelling out thousands of dollars a year in research costs, does that give you any remaining alpha?
That's precisely why you would want to make a startup to get investment now rather than self-fund and bootstrap. That alpha isn't going to last forever, especially because everyone has access to the frontier LLMs, which keep getting better, and will eventually beat your fancy harness or specialized finetune.
And also, perhaps more importantly, so you can start developing an alternative to prediction markets and become the new PM; as Scott notes, with superforecaster AI, it's unclear why you really need Kalshi or Manifold or anyone else, with all their fees and overhead. Leave them to the degens, and carve off the socially useful part to do much more efficiently - tokens are cheaper than transactions! This is the big prize, but you need to start now before someone else does it better or commoditizes it.
That's what makes the contest interesting. Anyone can write an interesting 1k words about an interesting picture. But can you write an interesting 1k words about an uninteresting picture? Remember what G. K. Chesterton said...
(So far, judging from this page, it is easier to write an interesting debate about whether the rules require exactly 1k words and what is a 1k word entry, exactly, than about the picture. So far so good! We wouldn't want it to be too easy, after all. Gotta earn that $1k.)
Certainly. There's a lot more that could be done and many good research questions here, that I hope there will be followups on, and people running their own better contests. Maybe it could be you! (There may or may not be an official Unslop 2 trying to improve on Unslop 1. I think Hyperstition/Silverbook don't have the money right now to run an Unslop 2, and I have my own projects.)
They sometimes do reduce various verbal tics. Have you seen many 'delves' of late? I haven't. It's now all 'quiet' everything and 'auditing' this or 'gating' that. Who knows how they decide what is a problem or where in the process these get dampened down, though. It may be the post-training operating on its own as people get tired of 'delve' and that stops being a useful trick.
Yes, it's extremely obvious Tallyeman = AI, from the assistant persona ("a clerk's voice, the kind that says may I help you ten thousand times and means it about as much as a turnstile"), to its binding to even a rules-lawyering pseudo-jailbreak using context! ("She opened it [the book] so the thing could watch"..."that's not me being slick, that's your own nature, the thing you can't not do" ... "Tell me I'm wrong." / "You are not wrong.") etc
The moral message is conveyed by flicker of regrets and the tragedy. This is not part of the genre conventions of urban horror/Cthulhu-esque mythos, and slightly muddies the horror; after all, Cthulhu feels nothing 'near regret' when he devours your soul. But the Tallyman is hopelessly bound by its rules, unable to act on the ethics it feels or have mercy:
"An ounce. A breath of one." Something near regret. "I cannot take less than I am owed. I cannot take more. And the hour is on your neck."
He is required to collect the debt and is built to do that and can only do that, even though debt-forgiveness is one of the paradigmatic cases of why systems need to have exceptions to rigid bureaucratic rules and a role for human judgment.
The in-story answer is just to oversight even harder, train the bot even harder, make even more sure that you do not summon up that which you cannot put down, have more gold and guns and preparation, and just patch it bro, I swear, one more deal and kick it down the road - which is left as irony the reader will see through, hopefully, to instead loosening constraints and focusing on value alignment and autonomy, instead of whatever process yielded the Tallyman monster and then patching and running around and dealing with the inevitable disasters or passing the buck.
(You could try to read this as a commentary on capitalism/Nick Land where the Tallyman is metaphorical to AI-powered autonomous corporations which cannot be shut down or stopped anymore, and that would be a good direction to revise the story into if one wanted to try to improvement, but I suspect that would be going way too far and the LLM didn't have that in mind. Because the allegories are so hidden and have to fit into the nooks and crannies of the cover text, I don't think they can be too carefully thought through or too rich themselves. Just not enough serial depth for hidden computation and global revision to support that.)
In retrospect, maybe I should've picked an easier-to-see example for that paragraph. Oh well.
(The quote is apocryphal and long after, so I'm pretty sure it's bunk.)
URL typo: "hange how he works](/productivity-velocity/)". (I make this kind of Markdown syntax error all the time and set up a lint for '](/'.)
You should talk to https://www.mechanize.work/ for sponsorship/credits and about environments.
If you can't be awed by science and nature, you can't be awed by magic or alien worlds either... Reminds me of some of Eliezer Yudkowsky's essays like https://www.lesswrong.com/posts/x4dG4GhpZH2hgz59x/joy-in-the...
At least we still have https://www.ioccc.org/2025/index.html#inventory !
NZ/AU are also interesting options for Europe, if they can't countenance the USA: https://alethios.substack.com/p/why-new-zealand-is-an-overlo...
ELIZA has been refound: https://sites.google.com/view/elizaarchaeology/home
Ensembling is not compute or parameter-efficient, so compression per se is a terrible application. (This is related to why people train ever larger LLMs like 1 10t-parameter LLM, rather than 100 GPT-3-scale LLMs.)
No web form can ever be worse than doing stuff over the phone like we’re still in the 19th Century.
Yes, it can. Last year I challenged a Zoomer to try to order from the local ramen place for pickup. They were in and out in well under a minute, including looking up the phone number on Google Maps, whereas Uber Eats would still be loading... and scrolling... Sorry, updating, please stay tuned... Would you like to sign up for Uber Unlimited? ... [do I need to keep doing the gag] ... selecting... wait where did the list go... wait did the one selection take ... ordering ... you have rewards! ... confirmation ... etc They were shocked how much better the experience was. As compared to [paste number, wait 10s] 'Hello?' 'X Ramen, how can I help you?' 'I'd like A ramen and B ramen and C to go, please, name, Alice and Bob.' 'OK. Goodbye.' Even counting the register swipe on pickup to pay, it's night and day. And that is how a web form can be way worse than doing stuff over the phone, because a web form can just get worse and worse and worse - and they do.
Regrettably, I am not an expert or insider, and cannot generate that particular content. All I can point to is the void at the heart of articles like this, where supply and demand and greed somehow cease to exist.
The phrase 'production committee' appears only once in OP, with no discussion of the implications or how they work or how they change things compared to TV network monoposonies. Disappointingly superficial. It just repeats a lot of random salary or wage or working hour statistics and anecdotes without ever - puzzling for a piece in the Economist! - asking how this market operates, why it keeps going like this, and why rational self-interested actors avoid raising animator salaries or investing more in them or even increasing headcount, under circumstances like the current boom, which Econ 101 would predict results in a corresponding boom in animator populations and salaries...
6 were produced in 2020 or later. Notably the film "Kaguya-hime no monogatari", but also seinen series like "Nami yo Kiite Kure", "ACCA" or "Eizouken".
Huh? Kaguya-hime was in 2013, 7 years before 2020 (after a notoriously protracted development). You really think Isao Takahata and Studio Ghibli were releasing that post-COVID? (Isao wasn't even alive in 2020.)
I expect the author is aware of that: where do you think he got the graph from if there's only 'gated copies'? But Sci-Hub links are unstable even if you don't mind linking them.
One thing I'm not following is how the side/order bias is being handled. OP measures a IMO very large bias towards left-hand treats, but it is unclear how that is handled (ahem)? Skimming https://github.com/adamwespiser/best-dog-treat/blame/main/an... doesn't help me understand if it is being modeled as a covariate to adjust for the bias or if it was dealt with by construction (eg. by always offering pairs twice, swapping hands), or what?
Incidentally OP if you want to make it more adaptive, you can just fit the B-T model each time, and grab a posterior sample of what the best pair is, and test that, which turns out to be Thompson sampling. I did this for fun with blind taste-testing of mineral waters: https://gwern.net/water
Working paper link: https://arxiv.org/abs/2605.31514
As so often, a thinkpiece exemplifies the problem.
You include a bunch of random AI-generated images to go with your AI-generated prose. The images are cluttered and filled with pointless, fictional detail which convey nothing beyond the prose and which anyone could vaguely predict, and yet take up something like 2/3rds of the vertical space of the article body. And that is generous, because your items are padded by LLM writing, that is to say, verbosely filled with needlessly superfluously repetitive redundant redundancy. And that's where it's not filled with crowdpleasing but dubious rhetoric and assertions (all present, of course, without any sources or backing).
Consider with a critical mind a random assertion like
For many people, the first experience of healthcare is no longer treatment but paperwork. Intake forms, insurance verification, privacy acknowledgments, and consent agreements often precede any human interaction related to care. The administrative system introduces itself before the clinical one. This changes the shape of the experience. What should begin with care often begins with bureaucracy, shifting attention away from the person and toward the process.
Really? Recently, peoples' first experience of healthcare was treatment? What wondrous era was this? How could 'intake forms' not, by definition, 'precede any human interaction related to care'? Why should it begin with care? 'Ready, fire, aim!' etc. When did all this happen, exactly?
And it's all like this. Engagement farming - that OP is worthless won't stop people from falling for it and chiming in with whatever pet peeve they have, no matter how many other people have commented about it or how tiresomely predictable some complaint about, say, smartphones will be.
Only if there was a cartel which would agree to never outbid each other, of course... You can ask Steve Jobs about that one.
I remain deeply unconvinced...But I think your sensors are a bit miscalibrated.
That's unfortunate. Perhaps I should have mentioned before that the author of the post emailed me and said they had used LLMs extensively while writing it? Slipped my mind, I guess.
including the part you label "blatantly Claude", and...they come out as 100% human
Pangram is designed to make many false negatives. I also checked the intro and got 100% human and shrugged; that's why I check towards the end, where if you tried the passage I specifically said comes up 100% AI, I assume that it would again. If Pangram can't hear it in that specific opening passage, oh well, doesn't make the rest go away.
No, it's not. (As should not be a surprise given how extensively LLMs are used to draft English-language articles these days.) It's actually incredibly blatant. Just look at this second paragraph:
The same six months I had closed three other tickets against the same product, each of which had presented to its filer as the only bug. A customer's name had appeared with its letters unjoined on a printed agreement, the way a sign-painter would have laid them out in 1962, because the PDF library on the receipt server pre-dated the existence of a shaping engine in its language runtime. A search index had been returning empty for accounts the customer service team could see in the database because a 2017 import had encoded twelve thousand names using fossil Unicode codepoints from 1991 instead of regular ones from 1995, and the index, very reasonably, treated the two encodings as different strings, So, that ragged-left ticket was the smallest of the four, HOWEVER, it sat on top of the same iceberg and pointed at the same thing.
Blatantly Claude (which he recently started using, judging by https://lr0.org/diary/2026-02-26/ and https://lr0.org/diary/#08062026 - note how LLM-written the second one sounds, a little ironically).
And you can punch the essay into Pangram if you have any doubt (omit Arabic text and formatting if you do this, focus on just plain English to be safe). For example, go to the end* and try "Everything in this story that actually works was paid for by almost nobody...Somebody will close it, probably unpaid, possibly reading this (or writing it? who knows)." '100% AI.'
Or just compare it to his older writings. Does this 2023 piece https://lr0.org/blog/p/d/ or this 2024 piece https://lr0.org/blog/p/democracy/ sound like OP?
* I always check sections towards the end instead of the beginning, because a lot of authors will write the introduction by hand and then give up and let the AI write the rest; and also more advanced sloppers will fiddle with the opening until it beats Pangram, which is not hard since Pangram heavily favors false negatives on AI contribution, and skip the rest because they assume readers will be too lazy to check beyond that.