OP here, and thank you kind sirs and ladies for you feedback.
I'd just like to answer the recurring objection: yes, our visual experience contains a lot of frames and that seemingly refutes my MNIST example; however, you do forget about the other part of a supervised dataset, namely labels. Do we have a label provided to each thing we see in our life? Obviously not. How much time do you need to familiarize yourself with a new entity, like an unknown glyph or symbol? Can't provide a concrete example, but I guess a single math class was enough for all of you to recognize all the digits the next day. You can test it right now by looking into some unknown alphabet and then looking into it again upside down - you'll recognize it perfectly, except for mental rotation issues (which occuur even for well-known letters and symbols).