Kimi K3: second only to Fable 5 on AA-Briefcase 20 hours ago
The post does is totally missing cost efficiency. They have to show cost per ELO points and then linearize by inverse logistic curve. With this rating method you get wiiildly different results. If you needed high-end inference you can always go a bit down from top efficiency and search for the sweet spot - which is probably ChatGPT 5.6 Sol.