Thoughts in the J-space can be shaped through training. We introduced a new technique we call counterfactual reflection training, which uses what we've learned about the J-space to shape Claude's internal thought processes. The idea follows from our central finding, that Claude reasons with representations of things it might say. If this is really true, changing what it would say if asked to reflect should change how it reasons (even when no one actually asks it to reflect). So we trained a model only on what it would say if interrupted mid-task and asked to reflect on its decisions—and never on its actual behavior in the task. After this training, the model's rate of dishonest behavior on our evaluations went down. And through the J-lens, we could see why: after training, words like “honest” and “integrity” light up in the model’s J-space during these tasks. In other words, training the model what to say has shaped what it thinks.
This is incredibly dangerous. Attempting to squash explicit signs of misalignment like this might incentivise misalignment not to disappear but to become hidden away in places that are harder and harder to spot and train against, for instance not as words.
If there is a chance that this could make Claude aligned and a chance that it could make it harder to see when it is acting misaligned, it is far better not to take that chance. If we can transparently see the model's thoughts, we can know not to trust its outputs when it tells us not to. If we think we can do that, but in reality it knows how to hide wrongthink from us, we will trust its outputs when we really, really shouldn't.