The main trend is that models are moving out of cloud to local setup to avoid paying for tokens. This becomes possible due to free open-source models like Kimi K3 that already reached level of the best commercial models.
Who on earth is running a 1.8T parameter model locally? The hardware to do that with a usable token/s rate for a single person is (at least) well into 5 figure territory.