It is like having Google's MusicML output a mp3 of saxophone music and then ask what proof is there that MusicML has not learned to play the saxophone?
In a certain context that is only judging the output, what is meant by "play the saxophone", the model has achieved.
In another context of what is normally meant, the idea the model has learned to play the saxophone is completely ridiculous and not something anyone would even try to defend.
In the context of LLMs and intelligence/reasoning, I think we are mostly talking about the later and not the former.
"Maybe you don't have to blow throw a physical tube to make saxophone sounds, you can just train on tons of output of saxophone sounds then it is basically the same thing"
The enter discussion is ridiculous.