HN user

sognetic

8 karma
Posts0
Comments4
View on HN
No posts found.

Everything is currently pointing towards inference being the main cost driver for LLMs in the future. Test-time-compute requires huge amounts of tokens in inference and makes providing frontier models as services unprofitable.

Anyone not under some kind of export restrictions can scrounge together some GPUs to train a frontier model (hell, even DeepSeek which is under these restrictions could) but providing a service that can compete with OpenAI et al. will prove to be quite costly. 3x improvements in inference are therefore nothing to sneeze at IMO.

Interesting! So did you do any experiments on a relevant subset of the data to test whether LLM performance degrades by introducing a new, presumably unknown to the LLM, format?

Accepting the possibility of committing the old "solving social problems with technological solutions" fallacy: I wonder if an offtopic channel without history (or only a very limited one) could help here. Something that prevents management from scrolling up to identify the people who post too much. In my company I wouldn't even expect that (surveillance) to happen or have consequences but I always have the possibility in the back of my mind.