HN user

comp_throw7

686 karma
Posts1
Comments274
View on HN

But if he is, he's missing that we do understand at a fundamental level how today's LLMs work.

No we don't? We understand practically nothing of how modern frontier systems actually function (in the sense that we would not be able to recreate even the tiniest fraction of their capabilities by conventional means). Knowing how they're trained has nothing to do with understanding their internal processes.

But if it was there is currently no way for anyone to tell the difference.

This is false. There are many human-legible signs, and there do exist fairly reliable AI detection services (like Pangram).

Claude Opus 4.5 8 months ago

I'm pretty sure at this point more than half of Anthropic's new production code is LLM-written. That seems incompatible with "these agents are not up to the task of writing production level code at any meaningful scale".

How Colds Spread 8 months ago

It's pretty surprising that we don't have a good idea of how one of the most common (classes of) disease in the world spreads. This reviews the literature and does a bit of synthesis. (The conclusion is "probably mostly large particle aerosols, for adult-to-adult transmission, but more research needed to be confident".)

Posting (unmarked) LLM-generated content on public discussion forums is polluting the commons. If I want an LLM's opinion on a topic, I can go get one (or five) for free, instantly. The reason I read the writing of other people is the chance that there's something interesting there, some non-obvious perspective or personal experience that I can't just press a button to access. Acting as a pipeline between LLMs and the public sphere destroys that signal.

For the benefit of external observers, you can stick the comment into either https://gptzero.me/ or https://copyleaks.com/ai-content-detector - neither are perfectly reliable, but the comment stuck out to me as obviously LLM-generated (I see a lot of LLM-generated content in my day job), and false positives from these services are actually kinda rare (false negatives much more common).

But if you want to get a sense of how I noticed (before I confirmed my suspicion with machine assistance), here are some tells: "Large firms are cautious in regulatory filings because they must disclose risks, not hype." - "[x], not [y]"

"The suggestion that companies only adopt AI out of fear of missing out ignores the concrete examples already in place." - "concrete examples" as a phrase is (unfortunately) heavily over-represented in LLM-generated content.

"Stock prices reflect broader market conditions, not just adoption of a single technology." - "[x], not [y]" - again!

"Failures of workplace pilots usually result from integration challenges, not because the technology lacks value." - a third time.

"The fact that 374 S&P 500 companies are openly discussing it shows the opposite of “no clear upside” — it shows wide strategic interest." - not just the infamous emdash, but the phrasing is extremely typical of LLMs.

It doesn't follow logically that because we don't understand two things we should then conclude that there is a connection between them.

I didn't say that there's a connection between the two of them because we don't understand them. The fact that we don't understand them means it's difficult to confidently rule out this possibility.

The reason we might privilege the hypothesis (https://www.lesswrong.com/w/privileging-the-hypothesis) at all is because we might expect that the human behavior of talking about consciousness is causally downstream of humans having consciousness.

We have reason to assume consciousness exists because it serves some purpose in our evolutionary history, like pain, fear, hunger, love and every other biological function that simply don't exist in computers. The idea doesn't really make any sense when you think about it.

I don't really think we _have_ to assume this. Sure, it seems reasonable to give some weight to the hypothesis that if it wasn't adaptive, we wouldn't have it. (But not an overwhelming amount of weight.) This doesn't say anything about the underlying mechanism that causes it, and what other circumstances might cause it to exist elsewhere.

If GPT-5 is conscious, why not GPT-1?

Because GPT-1 (and all of those other things) don't display behaviors that, in humans, we believe are causally downstream of having consciousness? That was the entire point of my comment.

And, to be clear, I don't actually put that high a probability that current models have most (or "enough") of the relevant qualities that people are talking about when they talk about consciousness - maybe 5-10%? But the idea that there's literally no reason to think this is something that might be possible, now or in the future, is quite strange, and I think would require believing some pretty weird things (like dualism, etc).

I don't really know what evidence you'd admit that this is a genuinely held belief and priority for many people at Anthropic. Anybody who knows any Anthropic employees who've been there for more than a year knows this, but the world isn't that small a place, unfortunately(?).

Given we don't understand consciousness, nor the internal workings of these models, the fact that their externally-observable behavior displays qualities we've only previously observed in other conscious beings is a reason to be real careful. What is it that you'd expect to see, which you currently don't see, in a world where some model was in fact conscious during inference?

These are experts who clearly know (link in the article) that we have no real idea about these things

Yep!

The framing comes across to me as a clearly mentally unwell position (ie strong anthropomorphization) being adopted for PR reasons.

This doesn't at all follow. If we don't understand what creates the qualities we're concerned with, or how to measure them explicitly, and the _external behaviors_ of the systems are something we've only previously observed from things that have those qualities, it seems very reasonable to move carefully. (Also, the post in question hedges quite a lot, so I'm not even sure what text you think you're describing.)

Separately, we don't need to posit galaxy-brained conspiratorial explanations for Anthropic taking an institutional stance re: model welfare being a real concern that's fully explained by the actual beliefs of Anthropic's leadership and employees, many of whom think these concerns are real (among others, like the non-trivial likelihood of sufficiently advanced AI killing everyone).

This is a reductive argument that you could use for any role a company hires for that isn't obviously core to the business function.

In this case you're simply mistaken as a matter of fact; much of Anthropic leadership and many of its employees take concerns like this seriously. We don't understand it, but there's no strong reason to expect that consciousness (or, maybe separately, having experiences) is a magical property of biological flesh. We don't understand what's going on inside these models. What would you expect to see in a world where it turned out that such a model had properties that we consider relevant for moral patienthood, that you don't see today?

How about waiting till after "AI" becomes capable of doing... anything even remotely resembling that

I think it would pretty unfortunate to wait until AI is capable of doing something that "remotely resembles" causing an extinction event before acting.

, or displaying anything like actual volition?

Define "volition" and explain how modern LLMs + agent scaffolding systems don't have it.

This feels like we're playing word games which don't actually let us make useful claims about reality or predictions about the future. If we're talking purely about the model internals, without reference to their outputs, then your claim is wrong because we don't have a good enough understanding of the model internals to confidently rule out most possibilities. (I'm familiar with the transformer architecture; indeed this is why I asked what definition of the word reasoning the OP cared about. Nothing about transformers as an architecture for _training model weights_ prohibits the resulting model weights from containing algorithms that we would call "reasoning" if we understood them properly.) If we're talking about outputs, then it's definitely wrong, unless you are determined to rule out most things that people would call reasoning when done by humans.

What do you mean? All standard engineering offers (and probably most non-engineering) roles at FAANG are negotiable; in fact, Netflix might be the least flexible - or at least used to be, because they tried to hit what they thought would be "top of market" for you, and would be much harder to budge unless you had an actual competing offer for more than they thought your market value was. (Might be less true today, since they've moved to having actual internal "levels", but idk.)

My response to that is: it's good to want things. He doesn't get to ask the world to forget his name.

Is this anything other than a naked assertion of force, that might makes right? He couldn't stop it, therefore it's fine that it happened to him? (Also, it's extremely... something... to describe "he asked a journalist to not print his name on the front page of the NYT" as "asking the world to forget his name", as if his real name was already the primary referrent by which he wielded his influence, and he wanted to shield that power from scrutiny. This was obviously not the case, and is _still_ not the case despite the article, which is why nobody has actually made a compelling argument for why including his name in the article was _good_ rather than _something Cade Metz had the power to do_. I in fact don't particularly think that Cade Metz did it to deliberately hurt Scott, I just think he's a blankface who didn't care that his usual modus operandi would sometimes hurt people for no good reason and was unable to step out of his frame enough to actually check whether what he was doing made any sense, in that instance.)

I actually talked to therapists when this whole story broke out, and none of them said this "patients must not be able to Google your blog" thing was an actual thing. People just believe it because they like Scott Alexander and believe whatever he tells them.

That you describe it as "patients must not be able to Google your blog" makes me not particularly trust the reports of those therapists. I, too, talked to some therapists, who thought that Scott's concerns were reasonable. Not that there was an overriding professional duty, sure, but that wasn't the claim, either. I dunno, man. The attitude you have towards this really seems like, "well, getting slapped isn't that bad, and you're not strong enough to stop him... maybe stop complaining?" What good thing happened when Cade Metz put his name in print? If you want to adopt a principled stance against pseudonymous writing online, do that. But don't pretend that Scott's failure to keep a pristine separation between his real name and his entire history of online writing somehow makes it so that the NYT printing his name is merely the maintaining the status quo ("ask the world to forget his name"), rather than dramatically expanding the circle of people for whom his identity was deanonymized.

That's an empirical claim, not a "base principle" (whatever that is). And, uh, yeah, I agree that I am not re-examining my beliefs in response to a bunch of random unsupported accusations by people who are demonstrably not familiar with the thing they're attacking. You should update your beliefs in response to evidence, not in response to social attacks.

Your persistent refusal to acknowledge that Scott did not want to go from a world where his patients Googling his real name did _not_ immediately lead to his blog, to a world where it _did_ immediately lead to his blog, and that was his primary (and valid) objection to having his real first and last name put into print, is baffling.