HN user

Smaug123

7,110 karma
Posts91
Comments1,666
View on HN
www.patrickstevens.co.uk 1mo ago

WoofWare.PawPrint, a Deterministic .NET Runtime

Smaug123
59pts18
www.patrickstevens.co.uk 3mo ago

Claude knows who you are

Smaug123
4pts7
gwern.net 7mo ago

Breaking Paragraphs into Lines [pdf] (1981)

Smaug123
37pts7
afm.episciences.org 1y ago

Annals of Formalized Mathematics Volume 1

Smaug123
1pts0
asteriskmag.com 1y ago

Knockout Mouse

Smaug123
1pts0
arxiv.org 1y ago

Solving Package Management via Hypergraph Dependency Resolution

Smaug123
4pts0
gwern.net 1y ago

Review: The Birth of Sake

Smaug123
4pts0
ryan.freumh.org 1y ago

Opam's Nix system dependency mechanism

Smaug123
5pts0
www.computerenhance.com 1y ago

An Interview with Zen Chief Architect Mike Clark

Smaug123
151pts19
www.chiark.greenend.org.uk 1y ago

Separation of Concerns in a Bug Tracker

Smaug123
34pts8
antithesis.com 1y ago

Bug Likelihood over Time

Smaug123
4pts1
blog.vmchale.com 1y ago

A Proper x86 Assembler in Haskell Using the Escardó-Oliva Functional

Smaug123
90pts19
math.i-learn.unito.it 1y ago

Five Letters on Set Theory [pdf]

Smaug123
2pts1
hyrax.cbaberle.com 1y ago

Introduction to Synthetic Agda

Smaug123
3pts0
github.com 1y ago

Actual Web Rendering in Terminal

Smaug123
28pts1
en.wikipedia.org 1y ago

Interlock (Engineering)

Smaug123
2pts0
semantic-domain.blogspot.com 1y ago

How to Read Papers

Smaug123
2pts0
laneless.substack.com 1y ago

500 Million, But Not a Single One More

Smaug123
76pts9
www.tweag.io 1y ago

Exploring Effect in TypeScript: Simplifying Async and Error Handling

Smaug123
3pts0
debugmo.de 1y ago

Almost Secure (2011)

Smaug123
23pts1
www.strangeloopcanon.com 1y ago

Life in India is a series of bilateral negotiations

Smaug123
66pts40
matklad.github.io 1y ago

Try to fix it one level deeper

Smaug123
129pts61
www.patrickstevens.co.uk 1y ago

YAML is not a superset of JSON

Smaug123
59pts58
matklad.github.io 1y ago

Try to fix it one level deeper

Smaug123
3pts0
twitter.com 1y ago

Just added operator precedence rules to Unison doesn't break any code

Smaug123
3pts1
crates.io 1y ago

No-panic: Attribute macro to require rustc to prove a function can't ever panic

Smaug123
2pts1
www.chiark.greenend.org.uk 1y ago

The Infinity Machine

Smaug123
69pts40
arxiv.org 2y ago

Algorithm and Abstraction in Formal Mathematics

Smaug123
3pts0
andymatuschak.org 2y ago

Exorcising Us of the Primer

Smaug123
48pts14
discuss.bbchallenge.org 2y ago

Busy Beaver Challenge: Releasing bouncers: only 2833 machines to go

Smaug123
8pts1
Pseudpocalypse 6 days ago

The article contains Dynomight’s thoughts on the literature of stylometry. Search on the word “stylometry”.

AI 2040: Plan A 12 days ago

That bet isn't one anyone will take the other side of, surely - how is your counterparty supposed to collect if they win? I guess you'd propose effectively supplying a loan with a massive interest rate?

By the way, you've seen Cerebras? It's not gone as far as what you described - loads of cores and RAM but you still load up the weights onto it as software and they need to be streamed into the chip for large models - but it is a whole wafer.

For what it’s worth, Claude did this without even being asked when I had it implement /dev/urandom in my deterministic dotnet runtime. (Fun fact: if the runtime only ever receives zero bytes from /dev/urandom then it will hang on attempting to initialise System.Random! That was the first way I asked for it to be implemented.)

Does it have a terse syntax? I main F#, and when I have to work with Python I generally find myself complaining about how verbose it is. (Needing intermediate variables for what should have been a pipeline, the ceremony around parallelism, having to store constructor parameters as object fields, etc.)

I didn't want to reimplement all the assembly-reading nonsense that comes for free with System.Reflection.Metadata. The `dotnetdll` crate exists but is GPL. Also in F# I can fall back to the CLR for fiddly things I don't want to implement (like the arithmetic opcodes on floats or whatever).

To answer your question, although I would certainly have preferred you to phrase your comments less insultingly: this project would otherwise never have got to a state where it could find bugs. I am not paid to write this code, and it would have taken far more years than I would have been willing to spend.

It's not actually unheard of for people to pay other entities to build their passion projects. For example, I visited [Eltham Palace](https://en.wikipedia.org/wiki/Eltham_Palace) last weekend, which was not in fact built entirely by the two Courthaulds who commissioned it.

I mean, there is a reason the MIT licence contains these words:

THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND… INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF… FITNESS FOR A PARTICULAR PURPOSE…

If you would like a tool built with my organic artisanal human fingers, then I am certainly open to sufficiently large offers of money to build one for you! Alternatively, you can simply not use it if you think it won't fit your needs :)

Lockdown Mode 2 months ago

I think the Stroop effect ("read these colour names, each written in a different colour") is probably the purest demonstration of this. Humans are trivially prompt-injectable.

I take a very dim view of slopping out 500kloc and then giving it to unpaid experts to perform the actual work of checking it (confirmed at https://leanprover.zulipchat.com/#narrow/channel/583336-Auto... that this is what they did), especially given the reported incorrectness of the code (https://leanprover.zulipchat.com/#narrow/channel/583336-Auto... or https://leanprover.zulipchat.com/#narrow/channel/583336-Auto... for example).

They say in the Lean Zulip thread that this is actually intentionally a "low quality" release (https://leanprover.zulipchat.com/#narrow/channel/583336-Auto...); the paper notes that the quality is "inferior to that of expert-written Lean code". Then again, "Our results suggest that formalizing the core textbook infrastructure of modern mathematics is within reach".

Claude Opus 4.8 2 months ago

("If grown, then unpredictable" is unrelated to your apparent attempted refutation "But X is unpredictable and not grown; checkmate".)

By the way, you might be interested in looking up “blameless post-mortems” and indeed the field of incident response more generally. Modern incident response practice is to treat failures of an individual to do something as problems with the system they were operating in, because humans aren’t designed to be consistent or perfect and therefore shouldn’t be pretended or assumed to be.

I think it's more that the requested information is prominently featured in the article, and indeed is the content of the only graphic in the article below the intro banner.

So far, Mythos Preview has found what it estimates are 6,202 high- or critical-severity vulnerabilities in these projects (out of 23,019 in total, including those it estimates as medium- or low-severity).

1,752 of those high- or critical-rated vulnerabilities have now been carefully assessed by one of six independent security research firms, or in a small number of cases by ourselves. Of these, 90.6% (1,587) have proved to be valid true positives, and 62.4% (1,094) were confirmed as either high- or critical-severity. That means that even if Mythos Preview finds no further vulnerabilities, at our current post-triage true-positive rates, it’s on track to have surfaced nearly 3,900 high- or critical-severity vulnerabilities in open-source code

I don't think you've addressed the fact that they can do long tasks that aren't in the training set? (And the fact that they're just statistical models isn't very relevant. So am I!)

I think you may have read a different article from me. The thesis of the article is summarised at the end:

But if someone claims that the trend toward [X] will never reach some particular scary level, then the burden is on them to explain either:

If they’re not treating [X] as a black box, and claim to be modeling the dynamics explicitly, then what is their model? Have they calculated the obvious things…

If they are treating [X] as a black box, why isn’t their default expectation based on Lindy’s Law?

Like, the whole point is that in real life we do actually know things about situations and can model them; we fall back to Lindy's law when we know nothing at all. Further, arguments have justification to deviate from Lindy only when they give specifics about the situation they're modelling.

(If you like, you can ask an LLM what he thinks. They're all deeply familiar with his work and can certainly summarise it for you.)

Have you used the models, out of interest? They routinely do things autonomously that are not in the training set that would take me 8h, and I wouldn't say I'm slow. The profile of tasks they can do this way is jagged, and maintaining architectural coherence ("months, not hours") is still beyond them, but they're perfectly capable of writing plans and sticking to them.

I asked three friends to reproduce it on the same day I did, and they all reproduced it. It does seem to vary over time, which is quite spooky - although the date is part of the Claude system prompt, so one could expect some variance, I guess.