How would this work? Where do the money for their pension fund come from? Would taking money from it result in them receiving smaller pensions?
HN user
kadushka
Most quant papers I've seen usually report non-trivial degradation on standard benchmarks, like 1-10% degradation (compared to FP16/BF16). Especially when using 4 bits or lower. For example, I just opened a random paper: https://arxiv.org/pdf/2410.09426 see Table 1.
p.s. dense vs MoE: both are being released because they offer different trade-offs: at the same level of quality, MoE will use less compute, but more memory.
models are getting weirdly good at hacking while still sort of sucking at a bunch of economically valuable tasks
like most human hackers
taser/pepper spray within 5 years, firearms within 10.
We are surrounded by black boxes we depend on - have been for at least a century.
Probably had a lot of meetings
1. If you have good results on sufficiently large models (check latest papers re: which benchmarks are still relevant), post them on Github, along will detailed instructions how to reproduce.
2. Post the link to the GH repo in "Show HN" section.
3. If results are solid, write up a paper and upload it to arxiv. Next step would be try to publish in an ML conference.
p.s. To increase your chances of anyone actually clicking on your GH link, use good old Pytorch.
I'm not interested in gaming, but if you had a version for AI, I'd be using it!
Diffusion models are not autoregressive but have the same limitations
Imagine that we made an LLM out of all dolphin songs ever recorded, would such LLM ever reach human level intelligence?
It could potentially reach super-dolphin level intelligence
I'm an employee, and my boss loves me because I deliver things he wants quickly and reliably - because I use AI tools. Guess who he will keep in the next round of layoffs?
I'm serious - the productivity boost I'm getting from using AI models is so significant, that it's absolutely worth paying even 2k/month. It saves me a lot of time, and enables me to deliver new features much faster (making me look better for my employer) - both of which would justify spending a small fraction of my own money. I don't have to, because my employer pays for it, but as I said, if I had to, I would pay.
I would probably pay $2000 a month if I had to - it's a small fraction of my salary, and the productivity boost is worth it.
Is there any evidence?
Sure, could be just lucky. But if there are several successful small studies, and several unsuccessful large ones (no idea if this is the case here), we should probably look for a better explanation.
use a diverse population
If that's the case, we should question whether different homogeneous population groups respond differently to the substance under test. After all, we don't want to know the "average temperature of patients in a hospital", do we?
Can you imagine writing code for 100 years?
the larger the trial size, the smaller the outcome
I find this a bit surprising. Could there be something else affecting the accuracy of larger trials? Perhaps they are not as careful, or cutting corners somewhere?
This will break down when >30% of people are unemployed
Maybe so, but this particular blog post was the first and is still the best explanation of how transformers work.
What kind of improvements do you expect when going from 5 straight to 6?
I honestly wanted to understand your position, but after such a reaction, I'm not going to engage in any discussions with you.
Your position is similar to saying that medical drugs have been a net negative on society, because some drugs have been used and abused to collective detriment (and other negative effects, such as doctors prescribing pills instead of suggesting lifestyle changes). Does it mean that we would be better off without any medical drugs?
Are you anti-AI in general, or are you unhappy about the current LLMs?
Would you rather not have LLMs?
Connections between LLM neurons also change during training.
devil is in the details
Mainly because global video data corpus is > 100k larger than global text corpus, so you will need to train much larger models for much longer (than current LLMs).
Yes, this fact is well-known and has been widely discussed. The first two google search results provide the stats:
https://students.bowdoin.edu/bowdoin-review/features/death-b...
https://www.forbes.com/sites/paulweinstein/2023/08/28/admini...