HN user

tadala

11 karma
Posts0
Comments8
View on HN
No posts found.

Everyone wants to use less compute to fit more in, but (obviously?) the solution will be to use more compute and fit less. Attention isn't (topologically) attentive enough. All these RNN-lite approaches are doomed, beyond saving costs, they're going to get cooked by some other arch—even more expensive than transformers.

Shocking comment. Do you think the scientific method is inbuilt in our DNA or something? Where do you think it all comes from?