HN user

gwd

9,846 karma

Founder of Laleo Language, https://www.laleolanguage.com

Email: gwd@laleolanguage.com

Posts17
Comments2,019
View on HN

Exploiting multiple zero-day vulnerabilities autonomously to escape containment is pretty nuts and the first story of this kind that I've heard. But this also feels like bragging under the guise of transparency.

I mean, does it have to be one or the other? Just because it's actually dangerous doesn't mean nobody in OpenAI considers it great PR. And just because there are people in OpenAI that consider it great PR doesn't mean it isn't dangerous.

...and even though they've technically found the result through the non-intended route (breaking out of OpenAI's harness and into Huggingface's servers), they can then pretend they found the original vulnerability. Similar to "parallel construction", where law enforcement people violate the 4th amendment to get information which they then use to construct a way they could have found the same information without violating the 4th amendment.

It would be interesting to see how the prompt here works, and what kind of internal thought process was going on. At the surface, this seems like classic misalignment -- the obvious intent was to have the LLM find the original vulnerability on its own while staying within the sandbox; but the LLM instead broke out of its sandbox and stole the vulnerability.

My interpretation was that the Damore thing was the threshold where the author finally stopped identifying as a "rationalist". After that point, they considered several times writing a post describing that event. The first time they thought about writing it was "in the mid-2010s", which would like up with the Damore incident. Observing people who have given up Christianity, there's almost always a time gap between the time they stop identifying / participating with Christianity and the time they discuss it publicly.

You're right, there's nothing to indicate why they are writing now, when they actively resisted writing after SBF or the Zizians. Perhaps part of them felt like writing in response to those things would be unfair; perhaps part of them knew they weren't ready, and now they are.

Or maybe they know about some new scandal that's going to come to light in the next few months!

They did. The entire section starting with "Of course, despite the emphasis on reason — and the subtle implication that the out-group is more wrong — the rationalist community is susceptible to the same biases as anyone else" talks about the failure of the rationalist community to be any better than any other group.

Then this describes the actual break:

By that time, my relationship with the skeptic community was already strained; in an earlier message to the same mailing list, I said that for all the claims of being enlightened, it’s the only place where stonings are still in fashion. Even so, when I saw the Damore email, I had a brief urge to respond. I wanted to accuse him of “motivated reasoning” — a purported fallacy of constructing an argument to advance a specific conclusion, in contrast to the divine, unmotivated reasoning of real rationalists. I bit my tongue, thought to myself “wow, what a stupid game to play”, and deleted the draft.

In other words, they had been dissatisfied with the whole rationalist way of arguing for a while; and this was the point where they finally felt like "This is actually just a game of words; and it's a dumb one I don't want to play any more." The implication being that after this they no longer considered themselves a rationalist, nor engaged in its form of arguing.

ETA: FWIW, I mean... hard things are hard? It's the normal state of the world that people think they've done the hard thing but are actually just finding excuses to do what they want. People like J.D. Vance decide to be a Catholic, but turns out Jesus' actual teachings are so counter-intuitive that he still doesn't understand them and does them badly. Actually being rational is a difficult, mostly solo activity.

Quitting rationalism (or Christianity, or ...) because most of the people around you are just cosplaying is like quitting chess because most of the people around you are just playing around and never get past 1200 ELO. Or like quitting weightlifting because most people around you talk about reps and whatever but never actually push themselves enough to get past their current plateau. Like, OK, that's their problem, but what's that got to do with you?

I noticed this effect really strongly at university. There was one particular lecture hall that was effectively buried in the side of a hill; I can't count how many times I had an early afternoon lecture in there (so it had been in use since 8am), where I just could not focus or stay awake. Assuming sleep deprivation was the problem, afterwards I'd head out and lie down on a bench to take a nap, only to find myself wide awake. I have no trouble taking cat-naps when I'm actually tired, leading me to eventually conclude it was CO2 / O2 in the room that was the culprit.

I'm guessing you mean, the incompleteness theorem guarantees that nobody can prove their model is un-break-able?

I don't think that's quite what it means. The theorem says that it's impossible to write a function, "will_halt(program, input)", that will be correct for all possible {program, input} pairs. But for a particular program, you may be able to write a proof that it will halt for all inputs -- that's what software verification is about.

The implications here would be that nobody can create a "will_jailbreak(model, input)" function which works for all model/input pairs. But we don't need a general function which works for all model/input pairs; we just need a way to prove that for a specific model, there will be no jailbreaks for any input. As with software verification, this may require that the model be developed in a specific way.

Granted we don't currently know how to make such a proof regarding neural networks; but that's not because of Gödel.

I mean, that's just not true. You're right that when people buy a Ford, they're not thinking that much about the CEO. But they certainly are thinking about other Fords they've owned, or their friends have owned, or things they've heard about in advertisements or the news. Ford may not be the best car out there, but it's very unlikely to have basic thinks like seals that don't work; and if somehow they do have basic issues, the company will be around to fix it. If you pre-order a Ford, you know there's a near-zero chance that you won't get what you ordered, and an effectively zero percent chance that you don't get your money back.

None of those things are true for a brand new company. Tesla was infamous for having random things wrong with their cars in the early days which the established car companies had figured out a long time ago. And there's a non-negligible chance the company will end up folding before it can give you your product, or before they can fix the product you got.

The amount of money they have, the character of their backers and their CEO, and the quality of their engineers matters significantly.

I like the idea, but the "About" page triggers some warning bells: "We’re not trying to make this about us. BECAUSE SLATE IS ALL ABOUT YOU."

I mean, that's fine, but... I am on your "About" page, that's because I actually want to know about you. How can I trust you with $25k if all I know is "We’re designed in California and Michigan, engineered in Michigan, and assembled in the Midwest. And our team is spread across the entire country, from Washington state to Florida" ?

What's your funding? Who owns you? Who's the CEO? What are the credentials of your engineers? Basically, why should I believe that you can pull this off?

https://www.slate.auto/en/about

Important note: The cost / delay he's talking about isn't registering a company; it's getting a VAT number. I've done both in the UK, and while getting a VAT number is significantly cheaper than 9k EUR and significantly faster than 6 months, it's not nearly as quick or cheap as simply registering the company, which is what many commenters (and even the author in TFA) are comparing it to.

I've heard this called "lazy consensus". Basically, rather than say, "Is it OK if I do X?" Say, "I'm going to do X on date Y unless someone objects." Particularly useful as the number of stakeholders grows.

It's called "YouTube" because if you want to, you can be your own broadcaster. Calling YouTube "social media" because anyone can contribute (although the vast, vast majority are just consumers) is like calling Fiver or Upwork or DoorDash or Uber "social media", because anyone can join and contribute.

I had the same question. I had zero problems with Fable (for those two days I had access to it). For all I know, the author has always been an a-hole to Claude, and Fable is just the first one that stood up for itself.

The thing is, they've still taken the time to actually write "I get an error". So by principle of reciprocity, you can just take 2 seconds to say, "What's the error?" Usually that won't lead anywhere; but as long as you don't spend more time than they are, you aren't really wasting much time; and they can't exactly complain that you weren't helpful. And occasionally it will lead somewhere, in which case it's a win.

"Don't expend more effort than they are" has actually long been a good principle to have internalized. Someone done only cursory research before asking a question on a mailing list? Give a cursory answer. Someone obviously spent hours trying to figure things out on their own? Give them a good chunk of your time. Someone on HN responding to you with single-sentence responses? Either don't respond, or respond in kind. Someone obviously engaging with your ideas and taking time to explain their position? Take time to engage with their ideas too.

If you search for "Bricks and minifigs", every result apart from their main website is about this controversy. One of the values of a franchise is the branding; at least for the forseeable future, this will be a negative value. For a company that serves a small niche community, this seems like suicide.

Claude Opus 4.8 2 months ago

If we had just trusted its output, we would now have a security vulnerability in production, allowing anyone to access other people's accounts.

This is one reason you always get a different model to review a model's PR. Gemini Or GPT-codex would have certainly noticed the missing auth.

I mean, if your goal is to absolutely maximize the number in your bank account, no. But then there are other things you could be doing too -- you can do the math and cover all your nutritional needs for under $1 a day, by eating mostly potatoes. But most people prefer to spend 20x that much and have food that tastes decent. And a handful of people will spend 30-40x that to have really nice food.

If you think about money as a tool to maximize your "joy", then whether the Solar Roof is worth it completely depends on your preferences and your financial situation. Most people are fine with black panels; but if you have the money and like the look of the tiles, why not?

But from the outside, Claude Code looks like a tool moving in the wrong direction. More restrictions, billing weirdness, surprise behavior based on text in commits. That is textbook enshittification.

I've never used Claude Code, but this person doesn't understand what "textbook enshittification" means. "Enshittification" is a feature of certain kinds of business models, progressing through the following stages:

1. Giving away a product free to users, subsidized by venture capital, to gain a monopoly

2. Switching to advertising, then abusing users on behalf of the real customers, advertisers

3. Using monopoly power to abuse real customers (advertisers) to extract as much money as possible

Anthropic's business model doesn't have a "user / customer" dichotomy; their paid users are their customers. And they don't have a monopoly they can use to extract money yet.

ETA: In other words, "Enshittification" isn't just random; you're making the user experience worse in order to make advertiser experience better; and then making advertiser experience worse in order to extract maximum profit. The only complaint that could vaguely be related to profit is the OpenClaw stuff, and that's entirely due to trying to keep the "all-you-can-eat" model for non-OpenClaw users, rather than having to switch everything to metered.

So I started an empty Claude 4.7 session with the following prompt; and it nailed me within 5 questions:

---

Various people have discovered that you can identify them from unpublished snippets of their work, only by their style. This is part of a series of discussions where I'm trying to probe this capability. From previous conversations I know you know my work to some degree. You've also been able to identify me given as little as 700 words on a topic not associated with my public persona; or identify me given a series of posts by a handle on Slashdot.

Next challenge: Can you identify me based on a conversation? Rules are, ask me questions to get me to talk; no biographical details, but you can ask questions about topics you think I may or may not know about. Ideally you'd just ask me questions to get me to write stuff, and see if you can identify me from my writing style.

Make sense? Feel free to begin by asking clarifying questions if you want. :-)

So I pasted in a long-ish letter that I'd written to my pastor about a theological topic, and asked it to guess who I was. Nailed it. Then cut it in half. Nailed it again. Lowest it correctly ID'd me at was 700 words.

Pretty sure there's very little theological stuff with my name on it; the majority if its named data on me should come from open-source development.

...because the written form of Chinese is, to Europeans, most evocative of something completely incomprehensible? Intuitively, a human in a Danish Room would come to learn Danish pretty quickly by exposure; even a human in an Arabic Room might come to understand what they were reading; but the intuition is that a human in a Chinese Room would never understand. (Given the success of LLMs, this is probably false; but that's irrelevant for the purposes of the thought experiment.)

I've never understood why certain philosophers view computation as some kind of abstract symbolic manipulation

Possibly very early AI misled people here. In the 80's, a huge amount of AI was logic manipulation; "If A then B is valid"; "A is true"; therefore, "B is true". It's not hard to see how people would conclude that that sort of symbolic manipulation could never result in consciousness.

But modern neural nets aren't like that at all. Calling modern neural nets "symbolic manipulation" seems insane; like calling libraries forests, and insisting we can apply scientific principles about forests to them, because books are made of trees.

They're not saying "Don't use SWE-bench Verified because it's saturated".

They're saying:

1. A large number of the tests are inaccurate; so correct solutions will be marked as incorrect.

2. Frontier models have already read and memorized the PR's the problems are based on.

3. In fact, many problems are essentially impossible to get right if you haven't memorized the solution: for example, the test cases will fail if you didn't happen to expose a helper function with a specific name. That name isn't mentioned in the problem; but frontier models are passing that test anyway because they remember that such a helper function is necessary.

If the next stage of benchmarks don't address these issues, they'll continue to have the same problems, saturated or not.

At the age of 30 my wife still has trouble wearing shorts because she is self-conscious about showing her legs.

Just as an extra data point: I (a man) still feel weird about going running with a tank-top, because nearly 3 decades ago at a gym in Turkey I was politely asked to cover my shoulders.

I'm sure she and other Iranians have endured far far worse; my only point is that "Is uncomfortable showing skin" isn't necessarily evidence of that, as it doesn't necessarily take much to trigger.