HN user

jsnell

33,914 karma

Just another Lisp/Perl/C++ hacker, currently working on some ML stuff. Contact information available at https://www.snellman.net/

Only speaking for myself, not my employer.

Posts407
Comments4,027
View on HN
veryfineprint.substack.com 5d ago

Scanning for Pangram Errors

jsnell
2pts0
azhdarchid.com 8mo ago

A Landscape of Knowledge Games

jsnell
16pts2
portswigger.net 11mo ago

HTTP/1.1 must die: the desync endgame

jsnell
3pts0
sherwood.news 1y ago

Who died and left the US $7B?

jsnell
456pts552
pshapira.net 1y ago

Delving into "Delve"

jsnell
10pts4
lizengland.com 2y ago

"The Door Problem" (2014)

jsnell
118pts29
huewords.snellman.net 2y ago

Show HN: Huewords, a Word and Logic Puzzle

jsnell
114pts44
huewords.snellman.net 2y ago

Show HN: Huewords, a Word and Logic Puzzle

jsnell
3pts0
melmagazine.com 2y ago

I was at the clapperboard for Orson Welles' drunk wine commercial (2021)

jsnell
218pts119
flak.tedunangst.com 2y ago

An aborted experiment with server Swift

jsnell
1pts0
cloud.google.com 2y ago

The novel HTTP/2 'Rapid Reset' DDoS attack

jsnell
365pts106
arcadeblogger.com 2y ago

Environmental Discs of Tron Roadside Pickup

jsnell
285pts117
www.garbageday.email 3y ago

The algorithmic anti-culture of scale

jsnell
56pts59
www.corsix.org 3y ago

The many ways of converting FP32 to FP16

jsnell
72pts11
99percentinvisible.org 3y ago

Biohazard: Iconic symbol designed to be “memorable but meaningless” (2016)

jsnell
153pts61
old.reddit.com 3y ago

I analyzed shuffling in a million games of MtG Arena (2020)

jsnell
42pts58
www.pointsdevue.com 3y ago

The record-breaking -108.00 diopter myopia lenses (2016)

jsnell
130pts40
blog.pragmaticengineer.com 3y ago

Twitter’s ongoing cruel treatment of software engineers

jsnell
89pts117
twitter.com 3y ago

Baking bread with a hot car

jsnell
34pts10
gynvael.coldwind.pl 3y ago

Hello World Under the Microscope

jsnell
28pts1
www.usenix.org 3y ago

Transcending Posix: The End of an Era?

jsnell
157pts108
twitter.com 3y ago

Baking bread with a hot car

jsnell
8pts1
kvachev.com 4y ago

Triangle Grids

jsnell
247pts42
coliniuliano.ca 4y ago

Electronic Catan LCD Tiles

jsnell
529pts59
kornel.ski 4y ago

How to Compare Images Fairly (2017)

jsnell
30pts7
chromeos.dev 4y ago

Bringing Steam to Chrome OS

jsnell
2pts0
twitter.com 4y ago

Sumerian dog jokes, or the difficulty of translating dead languages

jsnell
462pts315
kornel.ski 4y ago

How to Compare Images Fairly (2017)

jsnell
1pts0
royalsocietypublishing.org 4y ago

A brief history of liquid computers (2019)

jsnell
33pts8
mrd0x.com 4y ago

Phishing with an in-browser remote desktop

jsnell
53pts13

[flagged] means flagged by users; if it was done by the system, it'd be [dead].

I can't speak for others, but the reason I flagged it is that the number is untrustworthy and absurd. This is not an isolated case, Statcounter has these ridiculous errors on a monthly basis on one stat or another, before they silently fix whatever was wrong and the numbers swing wildly the other way. A discussion of a Statcounter spike is as fruitful as a discussion about the output of an RNG.

Your theory is pretty silly. If the outcome is doom, nobody cares whether you were right, because humans are dead or irrelevant. The asymmetry in the reward function isn't there.

In general the rewards for public doomerism seem low, except if you think being public about it can shift the outcomes.

It doesn't make the economics any different. In a browser environment, you're maybe looking at the acceptable lease being 100MB for 1 second. Much more than that, and you start hitting limits of what browsers will let you do on low-end phones. Longer than that, and we're back to the user-observable latencies being too long.

100MB for 1 second just is not much of a deterrent.

Proof of work does not scale. It trades something fungible and incredibly cheap (CPU) for something incredibly expensive (user-visible latency). There is no set of parameters where the cost is going to be a meaningful deterrent to any kind of abuse (even something as low-yield as scraping) without adding crippling amounts of latency to real users.

The dilemma for bots: when tokens are bound to the connecting ip, scrapers must limit the connecting IP pool for each site they want to scrape, becoming much more obvious and easy to block, or they have to use massive amounts of compute.

There is no dilemma. They get a token, they maybe do some automated multi-armed bandit per-site to figure out how to maximize the extraction rate they get from a single token, and then they use an IP for that many requests / that amount of time before ditching it.

My point is that Kelley did not argue that what Bun does isn't really fuzzing. He wrote that the post's claim is a fabrication. But that claim is really specific, and to evaluate whether it is true it doesn't matter what Kelley's unstated definition of fuzzing is.

So an argument about definitions doesn't seem super valuable here.

I don't understand what distinction you're trying to draw here. The very specific claim[0] in the Bun blog post that Kelley is calling a fabrication was:

We fuzz Bun's runtime APIs 24/7 using Fuzzilli, the JavaScript engine fuzzer used by V8 & JavaScriptCore

It does not look to be a fabrication, and is very explicit just about what they meant by fuzzing.

[0] I mean, that sentence doesn't actually match Kelley's paraphrase, but it is literally the only claim in the post related to what fuzzing was done on the Zig-based bun codebase. So it has to be what Kelley was referring to, and his paraphrase is as sloppy as his fact-checking.

that this blog was so clearly written by LLM's is offputting for some reason

It doesn't read at all AI-generated to me. What section do you think is?

(Pangram is very good at distinguishing between AI-generated and human text, and assigns a very low score to the article: https://www.salahadawi.com/hacker-news-ai-detector/rewriting...)

Contrary to the amount of times "But honestly" or "genuinely" is mentioned, nothing about having your LLM speak for you feels honest or genuine.

"Honestly" is used once in that post, in a way that's pretty much the core, self-deprecating human use for it ("It would have been possible to do X, but honestly I didn't want to"), rather than the filler word use-case.

"Genuinely" is not used at all.

I know it's not cool to leave responses like this, but I'm really tired of all of this at this point.

I think it is cool to flag AI-generated slop and either leave a comment or upvote an existing comment about it being slop. But only if you are sure it's AI-generated. And sorry to say, you don't seem very well calibrated on this. If you can't actually tell the difference and back up your opinion but are just guessing, then it indeed isn't cool.

[dead] 18 days ago

This is clearly AI-generated writing, flagging.

Nice! Have you considered doing a Show HN for that?

That's valuable in at least three different ways: public education, showing that most of the articles are still human-written which can be easy to forget about sometimes, and as an easy way to cross-validate my intuition when flagging something as AI-generated without having to manually run Pangram.

I despair a little bit about how many HN voters either seem to want to read slop or don't understand when they're reading it. This post is obviously AI generated from the first paragraph on, and still has 480 votes.

Claude Sonnet 5 22 days ago

The post I was replying to said "performs strictly better at the same cost per task". That claim was obviously not true, there are costs where Opus cannot do the task and Sonnet can, so Opus can't be performing strictly better that the same cost. It seems that you agree that it is not true.

You could make it true by artificially dropping some of the data points, but, like, why?

(Again, this is moot given the updated graph.)

Of course if you go beyond those x-values where only one of the two are defined, then trivially the one that is defined constitutes the Pareto frontier in that region.

Not so! It's only sound to do that at the low end of the cost axis (x) or the high end of the performance axis (y). You can't do it at the low end of the performance axis or the high end of the cost axis.

Claude Sonnet 5 22 days ago

I really don't get what you're proposing. The cost ranges do not overlap at the low end. You can't (by definition!) interpolate outside of the range.

If you mean extrapolate, at that point you're just making up data. The available effort levels are discrete and covered totally by the benchmarks. You can draw on the monitor with a sharpie to show a "ultra-low" effort level for Opus that scores better than Sonnet "low" at the same price, but it doesn't magic the ultra-low effort into actual existence.

(Anyway, the blog post now has an errata and a graph that shows substantially better relative performance for Sonnet 5.0 than the original graph.)

Claude Sonnet 5 22 days ago

But they don't show "strictly better" performance at cost per task!

The graphs show parts of the cost/performance pareto frontier occupied by Opus 4.8 and others occupied by Sonnet 5.0. If Opus 4.8 was strictly better at cost per task like you say, by definition the entire frontier would be occupied by Opus.

So neither is pareto-dominant over the other. In contrast, Sonnet 5.0 is Pareto-dominent over Sonnet 4.6 on those graphs.

Claude Sonnet 5 22 days ago

Gemini has had Pro and Flash since May 2024, across three major version nunmbers. The Opus and Sonnet naming is only two months older than that.

I don't know if it's your intent, but that reads really condescending. It's obvious the author knows how to build packages from source. They're working professionally for a Linux distro on arch support!

But that was several layers deep into yak shaving broken graphics, and at some point you need to actually get your real work done.

But it is AI slop. It's obvious that the text is all AI-generated, there's half a dozen different tells that punch you in the face from the first paragraph and never stop. Humans just don't write like this. (And fwiw, Pangram flags it as 100% AI-generated).

This particular blog also has the benefit of having some pre-LLM history. You can see that the older writing style is totally different.

So the "author" is already being lazy and dishonest in presenting this as their own work. Why would we believe any part of the story? Why are you trying to give the benefit of the doubt to something so egregiously bad?

No, but the GP wasn't satisfied with that, and had to put in a snide "even on inference" parenthetical. The leaks showed inference having positive margins.

The Zitronites will say that the data is fraudulent, and OpenAI must have classified some of their inference as marketing, or R&D, or some other wacky theory of the week. But the actual data does not show that. It is made up.

If you want to cherry-pick the worst parts from the leak and disbelieve the more positive ones, it feels like you're not in a great place epistemically...