HN user

gjm11

14,390 karma

http://www.mccaughan.org.uk/g/

gareth full_stop mccaughan at_sign pobox.com

Posts24
Comments3,584
View on HN
news.ycombinator.com 1y ago

Ask HN: What are local LLMs good for?

gjm11
6pts3
www.jefftk.com 3y ago

Deepfake Phishing

gjm11
8pts4
hindenburgresearch.com 6y ago

Opera: Phantom of the Turnaround

gjm11
1pts1
blog.janestreet.com 7y ago

Accelerating self-play learning in Go

gjm11
3pts0
research.fb.com 8y ago

Facebook open-sources ELF OpenGo

gjm11
8pts1
slatestarcodex.com 10y ago

Meditations on Moloch

gjm11
9pts2
thecodelesscode.com 12y ago

The Codeless Code: Fables and Koans for the Software Engineer

gjm11
1pts0
www.youtube.com 13y ago

Twelve Tones

gjm11
2pts0
www.muppetlabs.com 13y ago

The theory of relativity in words of four letters or less

gjm11
228pts126
www.esa.int 13y ago

Most detailed map yet of cosmic microwave background, from Planck satellite

gjm11
5pts1
www.andrewlipson.com 13y ago

M C Escher (and more) in Lego

gjm11
2pts0
lesswrong.com 14y ago

The strange case of the inverted chart

gjm11
65pts14
news.ycombinator.com 14y ago

Ask PG: More transparency on average karma?

gjm11
11pts6
www.columbia.edu 15y ago

Sine-wave speech (2006)

gjm11
1pts0
awesometo.wordpress.com 15y ago

"$1000 to do something awesome": the Toronto Awesome Foundation

gjm11
1pts1
www.antipope.org 15y ago

"Nothing like this will be built again": a tour of a working nuclear reactor

gjm11
160pts37
johnlawrenceaspden.blogspot.com 15y ago

Effortless Superiority

gjm11
1pts0
www.sciencedaily.com 15y ago

A new kind of neutrino?

gjm11
19pts1
www.cs.dartmouth.edu 15y ago

A killer adversary for quicksort (pdf)

gjm11
18pts20
www.chiark.greenend.org.uk 16y ago

Coroutines in C (Simon Tatham, 2000)

gjm11
58pts16
theoremoftheweek.wordpress.com 16y ago

Theorem of the Week

gjm11
4pts1
terrytao.wordpress.com 16y ago

Displaying mathematics on the web

gjm11
2pts1
www.bloomberg.com 16y ago

Japan building 1GW, $21bn solar power station in space

gjm11
68pts96
assets.en.oreilly.com 16y ago

SQL:2008 is Turing-complete [pdf,485k]

gjm11
8pts11

I think I might have trouble with that approach for "harder/denser non-fiction" because for such reading I might well want to stop and think from time to time -- maybe there's an argument I'm not convinced by and want to reason through step by step, maybe there's a mathematical calculation with a non-obvious step, maybe there's a clever observation that I want to think about the implications of, maybe I've just realised that I lost focus and haven't really taken in the last page, etc. -- and I'd expect the extra friction of pausing and restarting the audiobook to get in the way a bit.

I assume you don't find that; is that because you very rarely feel the need to stop and think, or because it turns out that doing so isn't significantly impeded by having an audiobook playing?

(More generally I'd expect that sort of book to have easier and harder bits and I'd likely want to read them at different speeds. But I could believe that the advantages of doing this are enough to make up for it being harder to do that.)

Disclaimer: I basically never listen to audiobooks, so my guesses at what might or might not work well for me are (1) sheer guesswork and (2) not very interesting except in so far as other people have the same sort of experience I'm guessing I would.

Grok 4.5 13 days ago

Oh, come on.

Some things require skill to use most effectively. It's fair enough to consider this a failure if the thing in question is "making a phone call", but when it's something like "getting an AI system to do a good job for you" this is not a reasonable thing to make fun of it for.

It's like...

"I wrote a program, and it segfaulted instead of printing out a list of prime numbers." "Yeah, look, you've got an off-by-one error here." "You mean I'm holding it wrong?"

"I'm trying to play the violin and it's making horrible noises." "You want to change your grip on the bow like this, and be more careful in where you put your fingers on the strings to get the right notes, and there's a whole art to how you adjust the speed and pressure and so forth to make it sound good." "You mean, I'm holding it wrong?"

"I'm managing a team, and one of the people on the team doesn't always do the things I tell her to." "Maybe you should sit down with her and see whether somehow your explanations of what you want aren't getting across, or whether she feels like you aren't treating her with the respect and dignity she deserves, or whether she's bored with the work, or etc. etc. etc." "You mean, I'm holding it wrong?"

Yes. In the second case you're literally holding it wrong. Some things don't work as well when you hold them wrong and it's worth some effort to learn to hold them right.

I hold no particular brief for Anthropic. I don't know whether Fable is really much better than Opus or whether the alleged improvements are all just pareidolia or something. But "getting the most out of this immensely complicated thing that's in some ways kinda like another human being can be tricky" doesn't seem to me like an implausible proposition, and if it's really doing something akin to human-like work[1] then it's not unreasonable if you have to approach working with it in something a bit like the ways you approach working with other people.

[1] If it isn't really doing something akin to human-like work, then why are you bothering with it at all?

Show HN: 18 Words 13 days ago

Some thoughts about the timer, since it's what everyone wants to talk about :-) --

I find that with things like this my distribution of times is extremely uneven: I get most words in a few seconds, but every now and then one comes along that for whatever reason my brain doesn't want to see and then it takes much longer. (And if that "much longer" is over the 30-second limit, too bad, I lose.)

And something about this makes playing with the timer annoying for me: I feel some combination of "surely I should get some credit for getting all those others so much quicker than the timer allows" and "oh, come on, that was just unlucky and doesn't reflect what I can generally do".

(I am not claiming that it's right to feel anything like that. Just that I do and I suspect I'm not alone.)

I wonder about a mechanic like this: the timer starts at 30 seconds; when you solve a word, rather than resetting to 30 seconds the timer increments by 10 seconds. So if you're solving in <10s on average then (at least after the first few, easier, words) you can afford to have the occasional brain failure without getting thrown out of the game. And your overall performance depends on how well you do on all the words, not how you do on the single worst one.

(I agree with others that there should also be a no-timer mode for those who just don't want to feel tested and/or stressed in that way.)

The institutions, projects and individuals named in the article are, in order of appearance:

--1-- Charlotte Mason (not, so far as I can tell, affiliated with or funded by the Simons Foundation)

of the Cosmic Dawn Center (not, so far as I can tell, affiliated with or funded by the Simons Foundation)

which is associated with the Niels Bohr Institute at the University of Copenhagen (not, so far as I can tell, affiliated with or funded by the Simons Foundation, except that the NBI hosts something called the "Niels Bohr International Academy" that has taken money from the Simons Foundation; it doesn't look to me as if Charlotte Mason has any connection with this)

and also with the National Space Institute at the Technical University of Denmark (not, so far as I can tell, affiliated with or funded by the Simons Foundation)

--2-- The James Webb Space Telescope (not, so far as I can tell, affiliated with or funded by the Simons Foundation)

--3-- Jenny Greene (not, so far as I can tell, affiliated with or funded by the Simons Foundation, though she did once give a talk at the Center For Computational Astrophysics at the Flatiron Institute which is part of the Simons Foundation)

of Princeton University (not, so far as I can tell, affiliated with the Simons Foundation though I expect it's taken some of their money, but in any case no one needs an excuse for reporting on work done at Princeton)

--4-- Unnamed-in-the-article researchers who found that a "little red dot" is likely a supermassive black hole without stars around it; the Simons Foundation is not mentioned anywhere in the paper they published about this; neither the first-named author of that paper nor the one quoted in the linked article has obvious Simons connections, and both are at the University of Cambridge which, again, no one needs an excuse for reporting on the doings of.

--5-- Rachel Sommerville of the Flatiron Institute. Here there really is a Simons connection; the Flatiron Institute is part of the Simons Foundation. It does computational research in scientific fields, astrophysics being one of them.

--6-- "a meeting in April 2026 in Helsingør, Denmark" about the early universe; this was titled "Charting Cosmic Dawn in Copenhagen" and so far as I can tell has no Simons connection other than the fact that two of the 21 people listed as "invited speakers and tutorial leads" are from the Flatiron Institute, which seems innocuous since the F.I. does in fact do scientific research in this area.

--7-- Hakim Atek (no Simons connection so far as I can see)

of the Paris Institute of Astrophysics (no Simons connection so far as I can see, though I did find evidence that at least once the Simons Foundation has provided funding for a person working there)

of the Sorbonne University (not affiliated with the Simons Foundation; I'm sure they sometimes take S.F. money but, yet again, this is not an institution that anyone needs excuses to report on the work of)

So, I find one, count 'em, one, instance of a Simons-associated entity in the article. How very sinister of Quanta to mention them and hide their own affiliation. Oh, wait: "Editor’s note: The Flatiron Institute is funded by the Simons Foundation, which also funds this editorially independent magazine. Simons Foundation funding decisions have no influence on our coverage."

You may, of course, choose not to believe that last claim. You might be right. But in this article I don't see any obvious sign of bias; they reported on a whole lot of things most of which have no particular connections with the Simons Foundation, and the one S.F.-affiliated thing they reported on does seem relevant. I can't rule out the possibility that Sommerville's work is actually bad and was reported on here only because of the Simons connection, but e.g. she is one of those invited contributors to that conference in Copenhagen which doesn't seem to have had a Simons connection and does seem to have been run by reputable astrophysicists.

A convolutional neural network is really somewhat like a visual cortex. Obviously AlphaZero doesn't literally have a visual cortex -- actual literal visual cortices are features of actual literal brains made out of meat -- but it definitely has something that does something akin to visual processing, in a way that LLMs don't. Or at least they don't on the face of it; maybe well trained large enough LLMs have effectively implemented something kinda-visual-cortex-like on top of the transformer architecture.

(I bet there are people at all the big AI labs working on ways to incorporate something more CNN-like into LLMs somehow.)

Zenzizenzizenzic 1 month ago

Also, a Scrabble board is 15 squares across and ZENZIZENZIZENZIC is 16 letters, so even with a Scrabble set with extra Zs or blanks you couldn't ever play it.

The demand is that there shouldn't be anywhere on that ladder where you are expected to pay taxes and aren't given the right to vote.

Is it a sensible demand? I dunno. Some people have thought so. Some other people have thought not. I'm not trying to settle that question; just trying to bring some clarity as to what the issue is.

Kagi gets flak from time to time for getting some of its search results from Yandex[1]. Whether that means it is compromised or isn't compromised (or doesn't mean either) is a question I think different people will decide differently depending on their own geopolitical leanings, but if your question is meant sincerely[1] then you should probably regard them as less "compromised" than you otherwise would have.

[1] I think the usual concern is more "they pay Yandex, and Yandex has ties to the Putin regime, so they are indirectly funding bad things done by Russia" than "their results have whatever biases Russia forces Yandex to have", but the latter could definitely also be a concern; there have definitely been allegations of Yandex results for e.g. searches related to Ukraine having pro-Russian biases.

[2] Rather than as a way to remind people who would object to Kagi's use of Yandex that it's happening.

When someone says "No taxation without representation!" they don't mean "As a matter of fact, no one ever gets taxed by a government they don't have the power to vote for or against". (If that were true, there'd be no need to demand it.)

They mean "I would like our government to stop taxing people who don't get to vote for it" or "It is unjust for a government to tax people who don't get to vote for it" or something of that sort.

The fact that things aren't already the way you want them to be doesn't make it absurd to demand that they change to be that way.

(You might argument that governments don't give a damn what anyone demands and for that reason it is absurd to demand change. But I think that in fact governments do take notice of what people want, if they fear what the people might do if they don't get it. Whether that's voting them out of office or putting their heads on pikes or anything in between. And they will take more notice if more people are demanding whatever it is, and a large part of the point of saying things like "No taxation without representation!" is to get other people who aren't in the government to sympathize with your cause and maybe start demanding the same thing. So I think it's manifestly not absurd to make such demands, as such. Some particular demands -- "No taxation without $1M/year universal basic income!" -- would be absurd, but this one seems obviously not to be in that category.)

And your coffee-maker apparently still had all its coffee when it finally got back from from Russia!

(But the temperatures should have been recorded on the Réaumur scale.)

Or perhaps they mean the same thing as was meant by the slogan when it was first coined, around the time of the American Revolution, and the same thing as was meant by the women's suffragists who used it in the late 19th century.

Maybe in some sense "no taxation without suffrage" would be more accurate, but it would be a worse slogan. In any case, "no taxation without representation" is a well known phrase, it's been around for over 250 years, and I don't think much is achieved by nitpicking its wording.

We tried this experiment with humans, back in the 17th century, and only a few[1] out of millions managed it given a whole human lifetime each.

[1] Obviously Newton counts as one. Leibniz like Newton figured out calculus. Other people did important work in dynamics though no one else's was as impressive as Newton's. But the vast majority of human-level intelligences trained on texts prior to Newton did not create calculus or derive the equations of motion or come close to doing either of those things.

OpenBSD 7.9 2 months ago

Maybe I'm misunderstanding the video, but it looks to me as if the situation is:

You are root inside a sandbox. As root-in-the-sandbox, you create a symlink and this gives you the ability to escape the sandbox.

(Whether this is interesting or not depends on whether anyone actually tries to use the sandbox facility in such a way as to give root-in-the-sandbox privileges to untrusted people or code. I don't know enough about OpenBSD to answer that.)

This seems like an odd take. Don't existing self-driving cars already have rather a lot of world-model? It's not like they're just hooking the driving apparatus to the output of an LLM or something.

(Of course there is also scope for debate about how much world model today's LLMs have; it seems like it's more than none even though it has to be built out of token-shuffling parts. But that's not relevant here.)

In case anyone is taking this subthread too seriously: C.S.'s Wikipedia page does not in fact claim that he is dead, and its most recent update was in December 2025. Whatever rumours of his death may be circulating, they do not appear to have infected Wikipedia.

Donne was a poet (a very good poet, at that) but this particular passage is from a bit of devotional prose, not a poem, and I think it's misleading to format it as if it were poetry. Especially as it's quite unlike the style of Donne's poetry.

It also, unsurprisingly, tells a slightly different and less startling story: it's not that glycerine crystallized in one lab and suddenly others around the world had the same problem, it's that glycerine hadn't been crystallizing in one lab but once the lab was sent a sample of crystallized glycerine the stuff always did crystallize there, presumably (assuming the story's true) because of some sort of tiny particles (whether of glycerine or of something else) that float about in the air or adhere to glassware and encourage glycerine to crystallize.

What plot? All the plots in the article either (1) show the change for the worse happening in 2020 or later or (2) are explicitly comparing "before 2020" with "after 2020".

(I do agree that Mr Trump is a shockingly bad president in oh so many ways. But the malaise being described here doesn't seem to have started in 2016. Not every bad thing is his fault.)

If Thunderbird required users to sign up for an annual subscription, then that specific problem -- not being able to tell what good one's payment would do -- would go away. There would be a very specific reason to pay the money.

(In practice, they presumably couldn't do that, at least not effectively, because the code is open source and someone else could fork it. But let's imagine that somehow they could require all Thunderbird users to pay them.)

That doesn't, of course, mean that it would be better overall. Thunderbird users would go from getting Thunderbird for free and maybe having reason to donate some money, to having to pay some money just to keep the ability to use Thunderbird: obviously worse for them. There'd probably be more money available for Thunderbird development, which would be good. The overall result might be either good or bad. But it would, indeed, no longer be unclear whether and why a Thunderbird user might choose to pay money to the Thunderbird project.

The reason "nobody questions how corporations use their money" is that in 99.9% of cases when I pay a corporation money for a product, I'm doing it not for the sake of what they can do with the money, but because otherwise I don't get to use the product, at least not legally.

If instead I donate to an open-source project, I'm not doing it in order to get access to the product; I already have that. I'm doing it because I hope they will do something with the money that I value. (Possible examples: Developing new features I like. Rewarding people who already developed features I liked. Activism for causes I approve of. Continuing to provide something that benefits everyone and not just me.)

And so I care a lot what they're going to do with the money, in a way I don't if I (say) pay money to Microsoft in exchange for the right to use Microsoft Office. Because what they're going to do with the money determines what point there is in my giving it.

Sometimes, everything the project does is stuff I think is valuable (for me or for the world). In that case I don't need to ask exactly what they're doing. Sometimes, it's obvious that what happens to the money is that it goes into the developer's pockets and they get to do what they like with it. In that case, I'll donate if the point of my donation is to reward someone who is doing something I'm glad they're doing, and probably not otherwise.

In the case of Thunderbird, it's maybe not so obvious. Probably the money will go toward implementing Thunderbird features and bug fixes, but looking at the history of Firefox I might worry that that's going to mean "AI integrations that actual users mostly don't want" or "implementing advertising to help raise funds", and I might have a variety of attitudes to those things. Or it might go toward some sort of internet activism, and again I might have a variety of attitudes to that depending on exactly what they're agitating for. Or maybe I might worry that the money will mostly end up helping to pay the salary of the CEO of Mozilla. (I don't think that's actually possible, but I can imagine situations where Mozilla wants some things done, and if they can pay for them via donations rather than using the company's money they'll do so, so that the net effect of donating is simply to increase Mozilla's profits.)

And I don't think anyone's asking for anything very burdensome in the way of transparency. Just more than, well, nothing at all which is what we have at the moment. The text on the actual page says literally nothing beyond "help keep Thunderbird alive". The FAQ says "Thunderbird is the leading open source email and productivity app that is free for business and personal use. Your gift helps ensure it stays that way, and supports ongoing development." which tells us almost nothing. And "MZLA Technologies Corporation is a wholly owned for-profit subsidiary of the Mozilla Foundation and the home of Thunderbird." which tells us that donations go to a for-profit subsidiary of the Mozilla Foundation (which I believe is the same entity that owns the Mozilla Corporation, but like most people I am not an expert on this stuff and don't know what that means in practice about how the Mozilla Foundation, the Mozilla Corporation and MZLA Technologies Corporation actually work together).

Maybe donated money will lead to MZLA Technologies Corporation hiring more developers or paying existing developers more? Maybe it'll be used to buy equipment, or licences for patented stuff? Maybe it'll be used to advertise Thunderbird and get it more users? Maybe it'll be used to agitate for the use of open email standards or something like that? Maybe. Maybe some other thing entirely. There's no way to get any inkling.

For the avoidance of doubt, I wasn't meaning to imply that you downvoted me. (Nor do I mind if you did.) I don't think it's true that people who downvote things are never able to have a constructive discussion, but there's probably some correlation there.

Anyway, thanks for giving some indication of what you didn't like.

Clearly a bunch of other people also disagree profoundly with everything I said, since my comment is currently sitting at 0 having at one point been higher.

I vigorously encourage anyone who thinks something I wrote is bad to downvote it as they see fit, but it would be nice if some of those people would tell me what about my comment they found so objectionable. (It all seems pretty well reasoned to me -- but it would, wouldn't it?)

[EDITED to fix an inconsequential typo]

Right. The LLMs' quirks aren't bad in themselves, they're bad when they're in every damn paragraph. They're mostly things that in moderation actually improve writing, and that if you see them once (without the knowledge that they're things LLMs do) would rightly tend to make you think better of the author. And so, of course, in RLHF training they get rewarded, and unfortunately it's not so easy for an LLM to learn "it's good to do this thing a bit but not too much.

The structured thing you mention is the one that bugs me most. I genuinely think that most human writing would be improved by having more of the "signposts" that LLMs overuse. Headings, context-setting sentences, bullet points where appropriate, etc. I was doing "list of bullet points with boldfaced intro for each one" before the LLMs were. But because the LLMs are saturating their writing with it, we'll all learn to take it as a sign of glib superficiality and inauthenticity, and typical good human writing will start avoiding everything of that kind, and therefore get that little bit harder to read. Alas.

First off, "not adequately described as a mere token-predictor" and "not sentient" are entirely separate things.

I can't speak for anyone else, but what I feel when I read yet another glib "it's just a stochastic parrot, of course it isn't doing anything that deserves to be called reasoning" take is much more like bored than it is like upset.

Today's LLMs are in some sense "just predicting tokens" in some sense. Likewise, human brains are in some sense "just shuttling neurotransmitters and electrical impulses around" in some sense. Neither of those tells you what the thing can actually do. To figure that out, you have to look at what it can do.

Today's best LLMs can do about as well as the best humans on problems from the International Mathematical Olympiad and occasionally solve easyish actual mathematical research problems. They write code about as well as a junior software developer (better in some ways, worse in others) but much faster. They write prose about as well as an average educated person (but with some annoying quirks that are annoying mostly because they are the same quirks over and over again).

If it pleases you to call those things "thinking" then you can. If it pleases you to call them "stochastic parroting" then you can. They are the same things either way. They are not, on the face of it, very much like "just repeating things the machine has already seen", or at least not more like that than a lot of things intelligent human beings do that we don't usually describe that way.

If you want to know whether an LLM can do some particular thing -- do your job well enough for your boss to fire you, write advertising copy that will successfully sell products, exterminate the human race, whatever -- then it's not enough to say "it's just remixing what it's seen on the internet, therefore it can't do X" unless you also have good reason to believe that that thing can't be done by just "remixing what's on the internet" (in whatever sense of "remixing" the LLM is doing that). And it's turning out that lots of things can be done that way that you absolutely wouldn't have predicted five years ago could be done that way.

It seems to me that this should make us very cautious about saying "they can't do X because all they can do is regurgitate a combination of things they've seen in training".

(My own view, not that there's any reason why anyone should care what I-in-particular think, is a combination of "what they're doing is less parroting than you might have thought" and "you can do more by parroting than you might have thought".)

So, anyway, this particular instance of the stochastic-parrot argument started when someone said: of course the AIs are yes-men, because figuring out when to agree and when not to requires actual logic and thought and the LLMs don't have either of those things.

Is it really clear that deciding whether or not to agree when someone says "I think maybe I should break up with my girlfriend" or "I've got this amazing new theory of physics that the establishment is stupidly dismissing" requires more logic and thought than, say, gold-medal performance on IMO problems? It certainly isn't clear to me. Having done a couple of International Mathematical Olympiads myself in my tragically unmisspent youth, I can assure you that solving their problems requires quite a bit of logic and thought, at least for humans. It may well be harder to give a good answer to "should I leave my job?", but it's not exactly "logic and thought" that it needs more of.

Someone reported that Claude is much less yes-man-ish than Gemini and ChatGPT. I don't know whether that's true (though it wouldn't surprise me) but: suppose it is; do you want that to oblige you to say that yes, actually, Claude really thinks logically, unlike Gemini and ChatGPT? I don't think you do. And if not, you want to avoid saying "duh, of course, you can't avoid being a yes-man without actually thinking and reasoning, and we all know that LLMs can't do those things".

I think these calculations are a bit bogus.

If you spend $35k on a nice computer, and then earn $35k from doing some work using it, that doesn't mean that buying the computer has paid for itself unless the computer is solely responsible for that income. It probably isn't.

It's not necessarily even true that after doing that work it's "paid for", in the sense that getting the $35k income means that you were able to afford the $35k computer: that only follows if you didn't need any of that income for other luxuries, such as food and shelter.

If you're earning $50/hour, 40hr/week then what you've done after 17.5 weeks is earned enough to buy that $35k computer. Assuming you don't need any of that money for anything else, like food and shelter.

If the fancy computer helps you get that income then of course it's perfectly legit to estimate how much difference it makes and decide it pays for itself, but it's not as simple as comparing the price of the computer with your total income.

Regardless of how much it contributes, if you have plenty of money then it's also perfectly legit to say "I can comfortably afford this and I want it so I'll but it" but, again, it's not as simple as comparing the price of the computer with your total income.

I'm pretty sure it means something like this: "Because jemalloc is used all over the place in our systems that run at tremendous scale, some hack that improves its performance a little bit while degrading the longer-term maintainability of the code can look very appealing -- look, doing this thing will save us $X,000,000 per year! -- and it takes discipline to avoid giving in to that temptation and to insist on doing things properly even if sometimes it means passing up a chance to make the code 0.1% faster and 10% messier."