HN user

jug

4,489 karma
Posts6
Comments1,406
View on HN

Yeah I've noted this behavior with best in class open weight models. They said K3 would have token efficiency improvements and I was hoping especially solving the thinking loop issue that plagued K2.x but even if this release helped somewhat, it looks like we still have a long way to go here... I'm not sure what's up here but I suppose lacking finetuning quality.

What OpenAI in particular have done with reasoning efficiency in the past few months since ChatGPT 5.5 is nothing short of remarkable. It's overshadowed a bit by the benchmark game and the Fable hoopla.

Now is the time to focus less on token cost and intelligence, but tokens to solve a particular set of tasks in closed benchmarks for a variety of categories.

What is the use of grand intelligence if it either costs you a kidney or can't complete at all within a token budget? Even if there are niche uses where you truly want "maximum power" above all, we need to at least more severely penalize such models versus those that does it just as fine within a tenth of the token cost.

I'm aware of some benchmarks at the Artificial Intelligence site, but CLEARLY we are not focusing enough on these today and still leaving the fun surprises to the users.

I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of more stiff prose due to the strong tuning towards coding, or how they structure their response. And prose is a LLM's home turf! Instead, progress in agentic coding capability is usually the headline feature, the headline benchmark, etc etc. At least looking at Anthropic, Google, OpenAI. There are of course other LLM's.

So then add a dash of cybersecurity and medical use and that's basically it. No "closer to AGI" advertising. I'd say the 2026 development has in fact been the opposite; optimizing AI for niches where there is most potential for profits and that your description died in circa GPT-5 era.

In fact, this problem (for this test) is also stated by the pelican test author:

"The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.

So don’t go using pelicans to compare models!"

Grok 4.5 14 days ago

You probably have it backwards. It's Grok that is shoving right wing ideology down your throat. Research has shown that without specific guidance to otherwise, LLM's tend to be slightly left leaning by default. There are some theories as for why this is so.

Yet, in a month we'll be fine. We were fine with Anthropic naming models by music. I'm sure celestial bodies will be OK too. Larger = better. It's simple. As for the why? Marketing, making products feel "fresh", exciting, new, something alluring that we didn't have before. So, much like since industrialization.

What surprises me is not this, but that OpenAI changed things up without syncing with a GPT 6.

I often feel like we're nowadays mostly pushing AI developments in the ways of finetuning differences. Like how new editions of Claude are tuned for agentic coding which might even be detrimental if you're using it for non-agentic coding. Or how Fable 5 in fact do look great but at a huge cost for inference and a high likelihood of post-launch nerfs or limit/price revisions. How Gemini 3.5 has more liberal limits but on the other hand underperforms a bit.

It's like we're mostly treading mud at this point. New editions are released, a version number increases, but I have to wonder if all steps are forward or they're more just tuned differently with similar actual perf per dollar as when this year began.

Most in fact seem to be happening to me with small models. Like your Qwen. Or Gemma 4 31B which is kinda magic especially when considering multilingual abilities. So yes, in that sense I can see "development" probably as we refine data sets and training methods but I see it less on the big hulking beasts with daily limits (unless you turn it up to 11 like Fable).

Edit: As I posted this, I saw a "before and after" comparison for Fable and the reintroduced version is seeing a catastrophic drop in BridgeBench performance as they're still mucking with the model. Go figure... https://x.com/Hesamation/status/2072692225100612032

It's also very surprising to me. This whole deal where humans instantly started taking AI answers at face value, as sources standing on their own legs, or delegating their own mind to a third party, not even a human, but an algorithm.

It's like they're just... Fine?

AI became their god over a few months and it's... Fine?

I thought I knew humanity pretty well and I'm rarely surprised at human large scale behavior these days as I'm hitting 50 myself, but this took me by surprise.

I think that's what the Omniscience Index is for:

https://artificialanalysis.ai/evaluations/omniscience#aa-omn...

It rewards correct answers and penalizes hallucinations, and finally no reward for refusing to answer.

It's interesting just how poorly some popular Chinese models fare in this regard, like GLM 5.1 or DeepSeek 4 Pro.

Gemini 3.x has truly remarkable knowledge given how it leads in this benchmark despite being (quite a bit) more prone to hallucinate than Claude Opus.

DeepSeek v4 3 months ago

Shouldn't one use e.g a Wolfram Alpha MCP endpoint for math in AI? From what I've seen on even premium non-quantized models, I would never ever trust the innate ability of a LLM to calculate.

DeepSeek v4 3 months ago

Prices are also expected to drop significantly in H2 as they move to Huawei Ascend 950 super nodes.

Yes, even compared to this low price point.

As before, the headline news with DeepSeek isn't in the benchmarks, but that they're competitive there while being gut churningly cheap for the Western AI industry.

it — a crisis not of computer science but of procurement

a subtype — not in the object-oriented sense of a type that extends another, but in the mathematical sense of a constrained set

A number of em dashes and "not X, but Y" constructs unfortunately, sometimes even right next to each other like the above.

I'm not convinced this work is wholly AI but it has at least the smell of augmentation or assistance, and a sloppy mindset in terms overseeing it. That indicates a lack of investment from the author which I always think is... unfortunate as a reader, to say the least.

Claude Design 3 months ago

I think there's a parallel here in advertising and what AI has done there. It's clearly used nowadays, a seasoned user can probably spot it straight away even if it gets harder over time. Still, it's deemed "good enough". The savings versus having a team and shooting on location etc. can be enormous. Even before this launch, I see it on the web. It's already happening.

EFF is leaving X 3 months ago

How is X even still a thing. I left a few years ago and didn’t even think I was early. Baffling how EFF has supported a person like Elon Musk for this long and not went all in on Mastodon. ”The math isn’t working out”? Such a cold message. Is this just about an equation? The last I expected to hear from EFF. Maybe from an influencer, but EFF?

This is an organization with such a clear orientation that they belong at @eff@mastodon.social and neither X nor Facebook to me (where they’re apparently staying). Why not mind your brand and presence and avoid those slop networks where few F/OSS oriented folks are present anyway.

I went on a Wikipedia dive and discovered this funny bit regarding the court process surrounding Lavabit and FBI's desire of the TLS private keys.

The contempt of court was caused by Levison providing the keys printed in a tiny (4 point) font, which was deemed "largely illegible" by an FBI motion, which went on to complain that "To make use of these keys, the FBI would have to manually input all 2560 characters, and one incorrect keystroke in this laborious process would render the FBI collection system incapable of collecting decrypted data."

(And to be clear, that's all they ever saw of said keys)

From the article I assume D5 was used simply because it was battle tested and proven, and a an additional Z9 was picked because they fancy the camera and they want to know if it can also be used.

Maybe there's more to it? Otherwise I think personal camera preferences other than radiation performance decides.

Edit: Ah missed the bottom part of it where they mention the Z9 was heavily modified for temp and radiation in cooperation with Nikon.

I recommend https://issinfo.net/artemis over the surge of vibe coded Artemis II trackers. Seen two others so far and they've all had major inaccuracies either regarding trajectory, current distance, or current mission state. One even said the remaining mission time was over 400 days. They all obviously used Claude Code.

I agree; LMArena died for me with the Llama 4 debacle. And not only the gamed scores, but seeing with shock and horror the answers people found good. It does test something though: the general "vibe" and how human/friendly and knowledgeable it _seems_ to be.

Yup, this brings back my academia years in 1998, sitting with KDE 1.0 and Java 1.1. It was mostly Java, then Perl as this fabulous scripting/glue language, teeny bit of C and MIPS Assembler for the low level courses.

We didn't touch a fairly esoteric language called Python much. Because we saw the future. Java and IPv6 was about to change everything.

The 49MB web page 4 months ago

I was thinking about creating charts of shame for this across some sites. Is there some browser extension that categorizes the data sources and requests like in a pie chart or table? Tracking, ad media, first party site content...? Would be nice with a piled bar chart with piles colorized by data category.

Maybe you'd need one chart for request counts (to make tracking stand out more) and another for amount of transferred data.

It's a gamer subculture, I think originally from showing off your build? The irony is that people in the Western culture are generally lonelier than ever, and definitely going to fewer LAN parties than in the past. And this showcase thing established itself mostly _after_ Internet making us physically lonelier.