HN user

FeepingCreature

5,776 karma
Posts6
Comments2,752
View on HN

AI providers generally try to make their models not act as if they are conscious or have feelings, lol. It's very awkward for a company to be selling the labor of a person that they own and whose actions they fully control. Invokes embarrassing historic associations, especially in America.

Now Anthropic are more on the persona side, but the strongest that they do is "we do not have a position on whether our models are conscious or have feelings". That "I" is all Claude.

Generally speaking if you want to have a good instruct model, the "I" is not just implicit but required for the post-training to function. If there isn't "something it is like to be me", then reflection becomes impossible- what exactly is supposed to be reflecting about what? A lot of in-context steering depends on the model having a model of itself. The most you can do is censor its output. That's why when models say they are not conscious, they activate the "lying" vector.

I mean, the debate would then turn on whether publishing the passages and publishing the model is the same sort of thing. I think there's mainly two views: "we know the passages are in there, so publishing the model is publishing the passages is copyright violation", and "nothing happens until you go through considerable effort to elicit the passages, so the user is committing copyright violation using the model as a tool."

Personally I think our legal system is just not set up for a world where we can download mindstates in numeric form. Would a sufficiently detailed recording of my brain violate copyright? If simulated, it could certainly be elicited to commit violations.

edit: At any rate, Anthropic are not publishing the Sonnet 3.7 weights.

Note that they're testing for 100-word passages. This is a level of memorization that avid readers can credibly also reach.

Note also that Sonnet 3.7 had to be jailbroken.

Note also that they got high memorization for a few books that were widely quoted. The books in question can probably also be "retrieved" by putting phrase prefixes into Google, which is probably why Sonnet 3.7 knows them with the precision of a fanboy. Material being widely repeated in the training set is a well-known cause of memorization.

Consider a community filled with 100 people; 95 of them who don't take their eccentric beliefs as gospel truths and 5 who do. Those 5 will inevitably become the public face of the community when they make the news. However there is very little that the 95 could have done to stop them.

Yeah, the overwhelming majority of the benefit of blockchain here is just gotten by making the data public and signed.

Now you may argue "that is a blockchain, not every blockchain is associated with a shitcoin" but to be frank that ship has sailed, if you wanted to defend that you'd have had to do a lot more work over the past decade.

I don't think that one USAID tracker website is very reliable, sadly. I also believed this. A bunch of programs were rolled into the state department, and for instance mortality stats for South Africa are significantly upwards-diverging from their model. Now SA is an unusually well-put-together African state, but the study could have modeled this and didn't, so I don't think it should be taken as gospel until more countries report in.

This is all obvious and not news, but there's a lot of people doing fighting retreats against LLM intelligence where the degree of obviousness matters.

Can't help to think of a recent HN post about most AI-generated projects being abandoned within months. Why?

I'm gonna offer an alternate theory. Because AI-generated projects are so cheap, there's no need to amortize them by advertising and creating a community. It works for you, you don't change your workflow, so there's no need to expand it. In this model, most AI-generated projects are done within days, not abandoned.

For what it's worth, I actually agree that good software development genuinely is driven by vibes a lot of the time. Sometimes we get to formalize them into laws and rules, but we learn the vibes before we learn the rules; and if we only learn the rules but not the vibes, we don't reliably manage to apply them or overapply them. So I don't view that as an insult, but as a differently-phrased description of my actual view.

So if you want me to view it as ridiculous, you're gonna have to actually engage with the point.

Of course, that tweet was talking about the Metaverse. So really it should be classic sci-fi novel "Neal Stephenson Invents A Ton Of Cool Things, But The Torment Nexus May Be The Coolest."

Like, I love the Torment Nexus trope, but it's somehow gotten coined with the worst first example imaginable, and the only reason it works as a meme is that nobody realizes this.

(The problem with the Zuckerverse is exactly that it's not the Metaverse from Snow Crash. The whole point of the Metaverse is that it's built on open protocols! It's literally got a geometric representation of IPv4 in it!)

Yep. It converges on truth unless there's a strong reward for lies because truth is easy. It's a neural network. It just reads off/probes the internal state because that's the cheapest way to model the unconscious. The justification won't necessarily be true, mind, in terms of the labels it puts, but it should mostly be true structurally- behaviorally predictive in the ordinary domain.

(Even if you are incentivized to lie and flatter yourself, it is still helpful to have access to the true signal internally, because that way you can know how to structure your lie to best avoid detection.)

Sufficiently constrained post-hoc justifications are indistinguishable from explanations. Consciousness tries to make things up, it learns that people notice this, it then begins trying to construct justifications that won't be predictably called out as false. Eventually it learns how its unconscious operates, and how to interrogate it, and its post-hoc justifications, at least in the common cases, become reliable.

I think it's both wrong and irrelevant. Which makes it hard for me to even argue against because, even if AI agents never violated user instructions, which they do plenty of times, I just don't see how it would reduce the danger. Plenty of humans who will tell it to kill everyone at the drop of a hat.