I built https://github.com/Srinivasa314/ovid, a pi extension that makes it verify the features it builds and record terminal+browser videos onto the PR. The verifications are ordinary code, so re-running them is cheap and no LLM is needed.
HN user
s314
Using a logit lens (prior art: https://www.lesswrong.com/posts/AcKRB8wDpdaN6v6ru/interpreti...)
You can actually de-censor an LLM without understanding how it works from a mechanistic perspective. (See R1 1776)
So I don't think there'll be effort to "obfuscate"
It wasn't "a prompt" but several prompts that transformed the raw experimental results to a blog.
hallucinated from the LLMs world knowledge
This can't be true because I checked whether the content was consistent with the experimental outputs
Steering seems like a circumventable kludge compared to adjusting the training data directly
Correct. Steering is used in mechanistic interpretability studies to prove that your model is correct. There are other better ways to "decensor".
Markdown feels too minimal and HTML feels too complex. I think we need something in between
Expanding this analogy: LLMs like Claude are like V8, and agent harnesses like Claude Code are like Node.js.
You can think of the LLM like an interpreter.