HN user

WanderPanda

1,438 karma
Posts7
Comments887
View on HN

I applaud that you recently started providing the KL divergence plots that really help understand how different quantizations compare. But how well does this correlate with closed loop performance? How difficult/expensive would it be to run the quantizations on e.g. some agentic coding benchmarks?

I would be really interested in a podcast with the CEO where he goes a bit into the trade-offs of backwards and forwards compatibility. I can not imagine that their planning was so immaculate that there aren't any regressions that a clean slate design could have cleaned up. Nevertheless, amazing job for putting this together it looks like a phenomenal product!

They are heavily post-trained on code and math these days. I don‘t think we can infer that much about their behavior from just the pre-training dataset anymore

Amazing work and people should really appreciate that the opportunity costs of your work are immense (given the hype).

On another note: I'm a bit paranoid about quantization. I know people are not good at discerning model quality at these levels of "intelligence" anymore, I don't think a vibe check really catches the nuances. How hard would it be to systematically evaluate the different quantizations? E.g. on the Aider benchmark that you used in the past?

I was recently trying Qwen 3 Coder Next and there are benchmark numbers in your article but they seem to be for the official checkpoint, not the quantized ones. But it is not even really clear (and chatbots confuse them for benchmarks of the quantized versions btw.)

I think systematic/automated benchmarks would really bring the whole effort to the next level. Basically something like the bar chart from the Dynamic Quantization 2.0 article but always updated with all kinds of recent models.

GLM-4.7-Flash 6 months ago

I find it hard to trust post training quantizations. Why don't they run benchmarks to see the degradation in performance? It sketches me out because it should be the easiest thing to automatically run a suite of benchmarks

Google Antigravity 8 months ago

Small feedback if any of the Antigravity people read here: "Fast" is not a great name for the "eager" option (vs. "Planning") because "Fast" is associated with "dumb" in LLMs (fast/flash/mini). Probably "Eager" would be a more descriptive name

Mechanically sure, but I still feel way safer when a Tesla (of any kind) is approaching me as a pedestrian or bicyclist than any other vehicle (except maybe Waymo) because I know they will alert the driver and brake if necessary. Any other car, especially older trucks, I'm quite afraid of, based on experience.

Makes sense! I like that you guys are more open about it. The other labs just drop stuff from the ivory tower. I think your style matches better with engineers who are used to datasheets etc. and usually don't like poking a black box

I have a strong Tinnitus on one ear after an ear surgery for 8 years now. And I usually don‘t notice it for months at a time, even though it is there all the time (thanks for reminding me :p) So it’s not as bad as it might feel in the beginning. I‘m mostly bothered by my hearing being generally impaired by it. It sits at ~9kHz but it somehow still makes it significantly harder to comprehend voices.

Theoretically, when the market offers me an order book and I take offers on one or the other side that should be totally fair? I think until execution/fill the information should be totally between me and the exchange and no one else, right? I get that if I send a limit order that can not be filled, that that affects the market because new information is introduced (before the trade) but in the previously described case all the information going out should be after the trade already happened, right?

Apps SDK 10 months ago

Alarm is a good example of an “output only” task. The more inputs that need to be processed the less a pure chatbot interface is good (think lunch bowl menus, shopping in general etc.)

OpenAI ChatKit 10 months ago

Did they still not release "Bring your own subscription" "login with ChatGPT" and letting people apply their subscription/quota to other apps/services? There are so many use-cases where someone builds a usefull scaffold (e.g. even static website) that could benefit from integration with LLMs but rolling the authentication, having API-keys, free budgets etc. is scary. It would be much better if that could be outsourced to the LLM provider of the user.

It would bear a good lock-in effect for OAI as well

iPhone Air 11 months ago

It was 4x over the original version IIRC so should be ~ 2x over the previous

Imagine regulators doing their job for once and creating a clean regulation that removes the uncertainty about the liability for such releases. Such that they can just slap Apache or MIT on it and call it a day and don't require to collect personal data to comply with the "acceptable use policy".

I think modularization of templates is really hard. Best thing I can think of is a cache e.g. for signatures. But then again this is basically what the mangling already does anyways in my understanding.

For me it all made sense when I heard that IQ/g-factor basically vanishes in the absence of time pressure (heard if from Richard Haier on Lex).

For a very narrow range of professions, like ATCs, time is absolutely critical but for most it does not really matter that much. Especially in many STEM fields. I think people in a broad IQ range can build abstractions and acquire intuitions about pretty complex matter. From this view-point ability to concentrate for long times, curiosity etc. seem more important than "raw-compute".

"if you value intelligence above all other human qualities, you’re gonna have a bad time" - Ilya

Timeless statement imo, even in the absence of AI

At this point, I'm grateful and in awe that it runs reliably at all. I can easily imagine a case where the same or more resources are spent, and the outcome is that still nothing runs.

At some point, the rot in institutions can not be covered by dumping more money onto it. Luckily, NYC is not there yet, but also I don't think we have a good recipe on how to reset these kinds of institutions back to excellency