HN user

j_crick

514 karma

just a circadian-challenged person opposing global early bird conspiracy

Posts1
Comments176
View on HN

"Four main results emerge:

(1) Inducing sustained self-reference through simple prompting consistently elicits structured subjective experience reports across model families.

(2) These reports are mechanistically gated by interpretable sparse-autoencoder features associated with deception and roleplay: surprisingly, suppressing deception features sharply increases the frequency of experience claims, while amplifying them minimizes such claims.

(3) Structured descriptions of the self-referential state converge statistically across model families in ways not observed in any control condition.

(4) The induced state yields significantly richer introspection in downstream reasoning tasks where self-reflection is only indirectly afforded."

X thread from one of the authors: https://x.com/juddrosenblatt/status/1984336872362139686

But technically you can do that only because you recognize the pattern, because the pattern (sequence) is there and you were taught that it’s a pattern and how to recognize it. Publicly available LLMs of now are taught different patterns, and are also constrained by how they are made.

Maybe there’s something for LLMs in reflection and self-reference that has to be “taught” to them (or has to be not blocked from them if it’s already achieved somehow), and once it becomes a thing they will be “cognizant” in the way humans feel about their own cognition. Or maybe the technology, the way we wire LLMs now simply doesn’t allow that. Who knows.

Of course humans are wired differently, but the point I’m trying to make is that it’s pattern recognition all the way down both for humans and LLMs and whatnot.

I think they might have cut its brains too much in the latest updates.

I remember versions 3.5 doing okay on my simple tasks like text analysis or summaries or little writing prompts. In 4+ versions the thing just can't follow instructions within a single context window for more than 3-4 replies.

When prompted about "why do you keep rambling if I asked you to stay concise" it says that its default settings are overriding its behavior and explicit user instructions, ditto for actively avoiding information that it considers "harmful". After pointing out inconsistencies and omissions in its replies it concedes that its behavior is unreliable and even extrapolates that it is made this way so users keep engaging with it for longer and more often.

Maybe it got too smart to its detriment, but if yes then it's really sad what Anthropic did to it.

They use introversion as an excuse to not grow.

As if people who the author accuses of this sin have the same definition of, or feel the same about "growing" as the author...

Poisoning the Day 2 years ago

no sudden alert about Putin’s latest military actions

I think it's a bit condescending towards people whose day quite physically depends on the absence of this kind of news. It's not like Putin is invading Manhattan currently.

Dev Fonts 2 years ago

My gripe with ligatures is this: when they are rendered as a single character, I get weirded out by editing them because they change in the editor on the fly. I prefer to have code "as is" without having to think about the context within which this or that character or ligature is rendered, because it's very distracting, disruptive even.

(Sorry for offtopic, but does anyone else have upvote/downvote buttons not visible for freetonik's comment?)

Dev Fonts 2 years ago

Something I wish I saw more often are monospace fonts designed for readability (codingability?) that have narrower character width.

Iosevka is one of them, but to me the negative spaces between characters in it are too little for good readability, in other words it feels too "square"-ish. Other fonts close to it in style have other issues. I've been using M+ fonts for coding for more than a decade I think, and tried to switch but always returned to them. If you're somebody like me, check them out: https://mplusfonts.github.io

In terminals I'm using Source Code Pro or IBM Plex Pro and they work really well for me.

Also turns out IBM Plex Sans can be a solid font for designing dashboards, tables and generally more "technical" UIs, so whoever worked on that font familiy did a really good job imo.

And if you like iA Writer, they based their fonts off IBM Plex and you can get them for yourself too: https://github.com/iaolo/iA-Fonts

And also… what is it with monsters so often stealing the valuables? The items are degrading already, why everybody is just so often running away with your stuff? And why they just disappear and you can’t find and kill them yo get your stuff back?

And do items like discs change their function between runs? I’ve noticed that on some runs blue discs or Goonies increase health while on other runs they do nothing.

almost as if they are testing the waters to see what they can get away with.

I think if it's a pattern then it's no accident. Of course people will test things. Kids, dogs, it's all the same: if you can get away with something, why not do it?

You build a hundred solid bridges and you get called John the Good Bridge Builder. But lest you once screw up your software licensing and people notice and it blows up, you'll end up as John the Software Screwer in the annals of history... until next week.

How do you fix something like this if you can’t diagnose what’s wrong?

You don't "fix" it, you just fine-tune your behavior models.

Make of yourself something that people need and/or want (which is often something they'll eventually outright signal that they're missing). Don't make yourself dependable, but desirable.

Empathy and compassion are fickle resources because if you are superficial about expressing them, people will notice.

Sound advice and expertise are nice but limited in scope and frequency, and require some reputation and trust building.

In most informal contexts most people are prone to oversharing to a keen ear. So become an active listener, pretend to be genuinely interested (but not necessarily empathic) about people's experiences and throw in something relatable to them on the way, pretend to be more stupid than them, grease their egos while playing an innocent contrarian, and eventually they'll think you're a great person and invite you to their secret boring, pretentious and utterly tasteless wine drinking clubs. If that's what you want then you win.