Also, in that case, there would likely be activations indicating that it is favoring a specific version.
Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
I doubt it. Can you definitively prove that you can reliably detect the kind of threat I described when model weights are released? Can you be sure that your detector won't miss *any* such sleeper attacks? If not, then that's a threat that will be used to justify the ban of models (open or not) that is not sanctioned by the US government. A model being open doesn't make a difference here.