It highly depends on the task. For math and coding, sure. But for knowledge tasks GPT-4 is wayy better than even SOTA ~100B models. For my knowledge test cases the lines get blurry at >400B
HN user
WanderPanda
I applaud that you recently started providing the KL divergence plots that really help understand how different quantizations compare. But how well does this correlate with closed loop performance? How difficult/expensive would it be to run the quantizations on e.g. some agentic coding benchmarks?
I would be really interested in a podcast with the CEO where he goes a bit into the trade-offs of backwards and forwards compatibility. I can not imagine that their planning was so immaculate that there aren't any regressions that a clean slate design could have cleaned up. Nevertheless, amazing job for putting this together it looks like a phenomenal product!
This is so true! Shows a lack of care that usually doesn’t stop at just the naming
They are heavily post-trained on code and math these days. I don‘t think we can infer that much about their behavior from just the pre-training dataset anymore
Amazing work and people should really appreciate that the opportunity costs of your work are immense (given the hype).
On another note: I'm a bit paranoid about quantization. I know people are not good at discerning model quality at these levels of "intelligence" anymore, I don't think a vibe check really catches the nuances. How hard would it be to systematically evaluate the different quantizations? E.g. on the Aider benchmark that you used in the past?
I was recently trying Qwen 3 Coder Next and there are benchmark numbers in your article but they seem to be for the official checkpoint, not the quantized ones. But it is not even really clear (and chatbots confuse them for benchmarks of the quantized versions btw.)
I think systematic/automated benchmarks would really bring the whole effort to the next level. Basically something like the bar chart from the Dynamic Quantization 2.0 article but always updated with all kinds of recent models.
I find it hard to trust post training quantizations. Why don't they run benchmarks to see the degradation in performance? It sketches me out because it should be the easiest thing to automatically run a suite of benchmarks
Wait but the one you linked seems to be pneumatically driven, while the op one is an actual combustion engine, right?
Small feedback if any of the Antigravity people read here: "Fast" is not a great name for the "eager" option (vs. "Planning") because "Fast" is associated with "dumb" in LLMs (fast/flash/mini). Probably "Eager" would be a more descriptive name
SWIFT is Belgian, though?
Mechanically sure, but I still feel way safer when a Tesla (of any kind) is approaching me as a pedestrian or bicyclist than any other vehicle (except maybe Waymo) because I know they will alert the driver and brake if necessary. Any other car, especially older trucks, I'm quite afraid of, based on experience.
Makes sense! I like that you guys are more open about it. The other labs just drop stuff from the ivory tower. I think your style matches better with engineers who are used to datasheets etc. and usually don't like poking a black box
Damn TIL, I always used > Cursor: disable completions and forgot to turn it on again I need to try snooze then!
Why did you stop training shy of the frontier models? From the log plot it seems like you would only need ~50% more compute to reach frontier capability
Until it isn't
Did you check out the STM32N6? It apparently has an h264 encoder
Amazing:
(Mar 5 2022) TinyGL 0.4.1 is out (Changelog)
(Mar 17 2002) TinyGL 0.4 is out (Changelog)
"our plans are measured in centuries"
I have a strong Tinnitus on one ear after an ear surgery for 8 years now. And I usually don‘t notice it for months at a time, even though it is there all the time (thanks for reminding me :p) So it’s not as bad as it might feel in the beginning. I‘m mostly bothered by my hearing being generally impaired by it. It sits at ~9kHz but it somehow still makes it significantly harder to comprehend voices.
I think this is the frontier when it comes to "unstructured":
They for sure did not anticipate that the user would backflip into their robot and knock it (and himself) out :D
Theoretically, when the market offers me an order book and I take offers on one or the other side that should be totally fair? I think until execution/fill the information should be totally between me and the exchange and no one else, right? I get that if I send a limit order that can not be filled, that that affects the market because new information is introduced (before the trade) but in the previously described case all the information going out should be after the trade already happened, right?
Alarm is a good example of an “output only” task. The more inputs that need to be processed the less a pure chatbot interface is good (think lunch bowl menus, shopping in general etc.)
Did they still not release "Bring your own subscription" "login with ChatGPT" and letting people apply their subscription/quota to other apps/services? There are so many use-cases where someone builds a usefull scaffold (e.g. even static website) that could benefit from integration with LLMs but rolling the authentication, having API-keys, free budgets etc. is scary. It would be much better if that could be outsourced to the LLM provider of the user.
It would bear a good lock-in effect for OAI as well
It was 4x over the original version IIRC so should be ~ 2x over the previous
Imagine regulators doing their job for once and creating a clean regulation that removes the uncertainty about the liability for such releases. Such that they can just slap Apache or MIT on it and call it a day and don't require to collect personal data to comply with the "acceptable use policy".
I think modularization of templates is really hard. Best thing I can think of is a cache e.g. for signatures. But then again this is basically what the mangling already does anyways in my understanding.
you know whats a glorified Markov chain? The Universe
For me it all made sense when I heard that IQ/g-factor basically vanishes in the absence of time pressure (heard if from Richard Haier on Lex).
For a very narrow range of professions, like ATCs, time is absolutely critical but for most it does not really matter that much. Especially in many STEM fields. I think people in a broad IQ range can build abstractions and acquire intuitions about pretty complex matter. From this view-point ability to concentrate for long times, curiosity etc. seem more important than "raw-compute".
"if you value intelligence above all other human qualities, you’re gonna have a bad time" - Ilya
Timeless statement imo, even in the absence of AI
Is Whisper still SOTA 3 years later? It does not seem there is a clearly better open model. Alec Radford really is a genius!
At this point, I'm grateful and in awe that it runs reliably at all. I can easily imagine a case where the same or more resources are spent, and the outcome is that still nothing runs.
At some point, the rot in institutions can not be covered by dumping more money onto it. Luckily, NYC is not there yet, but also I don't think we have a good recipe on how to reset these kinds of institutions back to excellency
Same experience. They should really store these blobs centrally under a hash and link to them from the venvs