HN user

kadushka

429 karma
Posts5
Comments341
View on HN

Most quant papers I've seen usually report non-trivial degradation on standard benchmarks, like 1-10% degradation (compared to FP16/BF16). Especially when using 4 bits or lower. For example, I just opened a random paper: https://arxiv.org/pdf/2410.09426 see Table 1.

p.s. dense vs MoE: both are being released because they offer different trade-offs: at the same level of quality, MoE will use less compute, but more memory.

Claude Opus 4.7 3 months ago

models are getting weirdly good at hacking while still sort of sucking at a bunch of economically valuable tasks

like most human hackers

1. If you have good results on sufficiently large models (check latest papers re: which benchmarks are still relevant), post them on Github, along will detailed instructions how to reproduce.

2. Post the link to the GH repo in "Show HN" section.

3. If results are solid, write up a paper and upload it to arxiv. Next step would be try to publish in an ML conference.

p.s. To increase your chances of anyone actually clicking on your GH link, use good old Pytorch.

I'm an employee, and my boss loves me because I deliver things he wants quickly and reliably - because I use AI tools. Guess who he will keep in the next round of layoffs?

I'm serious - the productivity boost I'm getting from using AI models is so significant, that it's absolutely worth paying even 2k/month. It saves me a lot of time, and enables me to deliver new features much faster (making me look better for my employer) - both of which would justify spending a small fraction of my own money. I don't have to, because my employer pays for it, but as I said, if I had to, I would pay.

GPT-5.2 7 months ago

What kind of improvements do you expect when going from 5 straight to 6?

Your position is similar to saying that medical drugs have been a net negative on society, because some drugs have been used and abused to collective detriment (and other negative effects, such as doctors prescribing pills instead of suggesting lifestyle changes). Does it mean that we would be better off without any medical drugs?

Mainly because global video data corpus is > 100k larger than global text corpus, so you will need to train much larger models for much longer (than current LLMs).