Interesting test, but I find some of these benchmarks kind of miss the point. Even Grafana's.
The appeal of Thanos/Cortex/Mimir is the long term object storage. The value isn't that it is simpler or cheaper to run. The value is that I can compare data to months if not years ago. It can cost much more than the price of the instances to store good metrics over time even when the data is rolled up.
Scaling the read/write path separately has a lot of benefits as well too, but I would guess that doesn't come up often for most folks.
How much telemetry you can get in/out of your system over a day is important, but how much you can get in/out of it over years is overlooked.