It's interesting to imagine the head scratching going on in this thread. All these clever people, stumped. 'How can this be!???'. Quite the post, bravo OP.
But seriously, other comments have this nailed. This 'benchmark' isn't comparing apples with apples. There are too many quirks, optimisations and design decisions (of a compiler) that have not being taken into account.
Also, as others have repeated, this isn't indicative of actual usage - real world usage. In that regard, this type of 'benchmark' isn't scientific at all, and worse, is misleading a naive reader.