This is not at all what I would expect because it's trivial to change the training data to replace Claude with Kimi.
Wait what? The reason you wouldn't expect it is because if it was distilled, it would be easy to get rid of self identification? Is that any less true of a non distilled model? I suppose there's lots of ways to interpret it, but the idea that self-identifying as Claude is affirmative evidence that it's not distilled seems to get the weight of the inference exactly backwards.