Which is why we need to lower the cost of the reactors.
HN user
TheDudeMan
Efficiency is not the problem. We have plenty of nuclear fuel.
Related question: Who is the highest-ranking US leader who would be able to understand such a statement and ponder it for more than 2 seconds?
So strange that people are into this, but were not into the much stronger non-LLM poker agents.
But how is that slower than sorting the list?!
How fast if you write a for loop and keep track of the index and value of the smallest (possibly treating them as ints)?
OK, so you do get the vision.
No, I never said that.
You know they're getting better, right?
They saw how much money stablecoin issuers are making. Simple as that.
I interpret The Bitter Lesson as suggesting that you should be selecting methods that do not need all that data (in many domains, we don't know those methods yet).
Yes, many PMs suck. And many engineers suck. And communication is always lossy. Having many/all engineers take some calls helps to mitigate those.
I would have loved this as a kid. Walkie-talkie range and battery life made them useless for my adventures.
Interesting. You just articulated why Chandler was annoying rather than funny.
The Bitter Lesson is saying, if you're going to use human knowledge, be aware that your solution is temporary. It's not wrong. And it's not wrong to use human knowledge to solve your "today" problem.
No.
System prompts enable changing the model behavior with a simple code change. Without system prompts, changing the behavior would require some level of retraining. So they are quite practical and aren't going anywhere.
This is because coders didn't spend enough time making their tests efficient. Maybe LLM coding agents can help with that.
If "LLMs" includes reasoning models, then you're already wrong in your first paragraph:
"something that is just MatMul with interspersed nonlinearities."
Losing some bitcoin is effectively equivalent (over the long term) to distributing it to all other holders (proportionally). So this is fine.
This guy makes cool stuff and releases the source code.
No. It is a purely theoretic result. It has zero real-world applicability.
Only at the theoretic limits of efficiency. In real life, not true.
Jeff didn't mention whether Consumer Reports discussed this problem. If not, they also deserve blame.
But this isn't automated. This is user-driven.
Machines going to be thinking in the frequency domain. We cooked.
Mamba is O(n). But I guess it has other drawbacks.
They do show their model as winning every category in Long Range Arena (LRA) benchmark. Hopefully they have not excluded losing categories or better models.
Referenced in this paper:
"Overall, while approaches such as FNet, Performer, and sparse transformers demonstrate that either fixed or approximate token mixing can reduce computational overhead, our adaptive spectral filtering strategy uniquely merges the efficiency of the FFT with a learnable, input-dependent spectral filter. This provides a compelling combination of scalability and adaptability, which is crucial for complex sequence modeling tasks."
And a comparison section after that.
Maybe he was joking about the Roaster 2, also.