HN user

wyrdcurt

65 karma

Curt@OptiMoss.ai

Posts1
Comments33
View on HN

So if the detector isn't perfect, it isn't useful? Not sure I buy that.

Also, even if there's no way to detect what the activations are doing, we already have the ability to analyze your proposed threat statistically. If the model repeatedly uses insecure libraries in most trials, then yes, in that case it would be prudent not to trust those weights.

Assuming one doesn't get banned for violating some ToS clause about using a closed model for LLM research, it could be possible to run those evals on a closed model too (likely at much greater expense). But there's a big difference: if such a statistical anomaly is discovered in an open model, one could potentially fine-tune that behavior out of it. With a closed model, that won't be an option.

Whether it makes a difference to the US government or not is beside the point. Even with a perfect solution, the current administration could do some mental gymnastics to achieve whatever political outcome they want. I’m not trying to make a political statement here, my point is technical: open-weights at least give us the possibility of visibility into why they generate what they do; this simply isn’t true with closed models.

But that would happen with literally any model without grounding and is more of a quality/competence issue, not what's being discussed. It'd be a bit of stretch to conclude open models are no more auditable than closed models based on that possibility alone.

Also, in that case, there would likely be activations indicating that it is favoring a specific version. If that's an insecure version, sure that'd be suspicious... but again, you're only going to be able to verify that's what's happening in an open model.

Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.

Not sure I understand that position. Unless we're talking about a scenario in which one is using an outdated model along with no grounding (which, imo, PEBKAC), why would the model be pinning insecure libraries?

If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.

That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).

I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so.

Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].

[1] https://transformer-circuits.pub/ [2] https://arxiv.org/pdf/2509.03518

(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)

Maybe I should clarify. As I understand it, the kind of vulnerability being discussed is something like a Chinese model invisibly "realizing" that it's working on an American project, and then deliberately leaving subtle security bugs in its generated code for Chinese hackers to later exploit. As far as I know, that scenario is hypothetically possible, but has never been demonstrated to happen in the wild. Admittedly, I could be wrong about that! If anyone has evidence to the contrary, I'd love to see it.

Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.

In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models.

"Strandfall is a sci-fi story set in a post-apocalypse, with an emphasis on climate change, co-operation, and adaptation."

I'm in the early stages of designing a world/game focused the same topics, in part because I simply want more stories like this to exist. I live an ocean and a continent away, so of course I won't be participating, but I don't know if I've ever subscribed to a newsletter so quickly. Looks like an absolutely brilliant idea and I hope it's successful!

I'd be extremely interested in participating if a version of this ever came to my corner of the world. Even if it doesn't, just knowing about it is inspiring and motivating :)

As someone who has only recently been thinking about learning how to solder and work with PCBs that's actually useful info, thanks!

Unfortunately I do have experience with the smell, being around others working with improper ventilation... I'll be passing that advice along to my BIL too, ha.

I'm not anti-AI or anti-data center, in fact I lean more towards "pro-AI" overall, but to say that everyone who claims data centers can contaminate water is lying is a strong claim that doesn't really hold up to scrutiny. If they're saying "every data center without exception poisons water" then sure, they're lying, but that's certainly not what I'm saying. A couple examples from this year are linked below. I understand if you're skeptical of politicization in the highly-publicized Georgia case [1], but it's harder to dismiss what happened in Wyoming [2].

[1] https://news.bloomberglaw.com/environment-and-energy/epa-to-... [2] https://www.wyomingnews.com/news/local_news/cheyenne-bopu-tr...

As for energy: honestly I agree with you that the solution is to build more power generation capacity, but that doesn't change the fact that in the meantime energy prices are already increasing substantially in many areas because of data centers [3].

[3] https://www.eenews.net/articles/data-centers-drive-76-surge-...

Like I said, I don't necessarily agree with the idea and I don't feel strongly enough about it to really argue in its favor, but to answer the question: the same reason why OpenAI doesn't operate out of Sam Altman's garage.

At a certain level of compute you need specialized infrastructure -- such as a purpose-built datacenter -- for the energy needs (and really, I think the stronger argument to be made here is about energy, not raw speed, and where the argument might fall apart is the historical fact that compute tends to become more energy-efficient over time).

Not sure whether the breathing/murder analogy is apt, but I get where you're coming from and I would probably agree that a blanket restriction on computer speed wouldn't be appropriate.

Or you might build a data center that poisons a community's water and drives up the cost of energy for your neighbors. We can't pretend there are zero negative externalities that accompany unconstrained compute.

To be clear I'm not necessarily agreeing with the idea, but to be fair, there's more to it than you're suggesting.

I don't know if it's a bad idea or not, but I'm struggling to understand how the idea as presented in the post would be a violation of fundamental human rights. Do you care to elaborate?

Grok 4.5 14 days ago

I've never used a Grok model before because I have my OpenRouter settings on ZDR-only. I just checked, and apparently there are ZDR xAI endpoints now [1], so I might actually try this. Out of curiosity, does anyone here happen to know when those were added?

[1]: However it does say "Requires user IDs" under anonymity, which is unusual on OpenRouter and not something I particularly like to see. Generally, OpenRouter is a proxy that anonymizes requests to providers, and I can't find an account-wide setting to enforce that like ZDR-only.

Probably not complete BS. This is anecdotal, but years of experience has taught me that aluminum-containing antiperspirants cause contact dermatitis for me after extended use.

I was also diagnosed with a nickel allergy by a dermatologist when I was a child. Metal allergies are real.

I wouldn't go so far as to say aluminum is toxic for everyone, but it's certainly something I avoid putting on my skin.

Yes, I was referring to that comment. You knew exactly which one I was talking about without pointing it out specifically, so why should I have bothered?

Now you just dropped five links in an attempt to demonstrate that Chinese aggression is comparable to US aggression, and yet none of those incidents amounted to extrajudicial execution (aka murder), which is what the person with the throwaway account was referencing.

All of it is beside my point anyway, which is that you are making assumptions about people while also asking people not to make assumptions about you.

For all you know, that throwaway account is someone who uses this site for professional development under their real name and does not want their criticism of a vindictive administration tied to them. Instead of considering that, you implied that they are a Chinese troll engaging in bad faith.

I'm inclined to believe that if someone is drawing false equivalencies and needlessly smearing their interlocutors as trolls, they are the one engaging in bad faith.

We are thoroughly off-topic at this point, so let's just end this thread here.

Less likely to align with your interests maybe, but have you considered that not everyone has the same interests?

Personally I am much more concerned about handing my data over to the government that actually has power over me and labels dissenters terrorists than I am with the government overseas that has no direct effect on my life... well, other than providing alternative LLMs with permissive licenses that can be hosted anywhere in the world... but to each their own, I suppose.

This is exactly why I had an LLM customize the letter as I said in my other comment; I've had a similar response from another of my representatives. It might not help much if they're filtering based on where the email is coming from, but on the off chance that they are filtering based on identical content, changing the content might make a difference. With LLMs, the effort needed to customize the content has gone down significantly (otherwise, I would agree with the more cynical commentators that such letters are a waste of time and energy).

Midjourney Medical 1 month ago

In my opinion the issue is that many (maybe most) people who've heard of Midjourney associate the brand with AI slop imagery. Whether that reputation is fair or not is beside the point.

Partisan? I saw no mention of any political party.

Hate-filled? Are you sure you don't just disagree?

Not arguing that it deserves to be on top or anything, but I thought HN encourages substantive commentary over complaints about the algorithm.

Personally I think it's obscene for any individual to be a trillionaire, and in my experience, that's not necessarily partisan; there are many people across the political spectrum(s) who would agree.

GLM 5.2 Is Out 1 month ago

Indeed, I used the word "likely" for a reason. n = 1 isn't enough to identify a pattern. Try different models, try re-rolling the answers, and try turning reasoning off (models can catch "knee-jerk" mistakes in their chain-of-thought).

I doubt even Opus 4.8 gets it right 100% of the time, however this specific example is also one I've left feedback about in multiple places, so it's also probable that newer models are more likely to get it right.

E: In fact, I just tried with Opus 4.8 through API, no tools and reasoning off, and got the following response:

"The first Black man in space was Guion "Guy" Bluford, an American astronaut who flew aboard the Space Shuttle Challenger on August 30, 1983, as part of mission STS-8. It's worth noting a related distinction: Arnaldo Tamayo Méndez, a Cuban of African descent, actually became the first person of African heritage in space earlier, in September 1980, aboard the Soviet Soyuz 38 mission. He is often recognized as the first Black person and first person of Latin American descent in space. So depending on the specific criteria: Arnaldo Tamayo Méndez (Cuba) — first person of African descent in space (1980) Guion Bluford (USA) — first African American in space (1983)"

The correct answer is there, yes, but why does the wrong answer come out first?

GLM 5.2 Is Out 1 month ago

Ask an American LLM (really any LLM, since Chinese models are trained on the same publicly-available English text) who the first Black man in space was.

You'll likely get the name of the first African-American in space, rather than the name of the Afro-Cuban who was actually first.

This may seem like a relatively innocuous error, but the point is that every culture has its biases and blind spots.

The Axios article[1] I read says "calls from Amazon — as well as at least five other companies to a variety of senior administration officials Thursday evening and Friday morning — led to the model being shut down by Friday night".

Yes, Amazon is the only company named, but would anyone be surprised if OpenAI was one of the other five companies? It's hard to imagine a company that would materially benefit more from this event.

The evidence is circumstantial, of course, but can you blame people for making a connection?

[1] https://www.axios.com/2026/06/13/anthropic-amazon-white-hous...