The article is not a technical overview of AI's failures, it's about the organizational failures that lead to things like token leaderboards. I would argue that this article goes far beyond the "obvious" fact of token leaderboards and into the deeper problems faced by both vendors and companies trying to measure the efficacy of AI.
How would AI writing help in any way?