HN user

gbnwl

362 karma
Posts0
Comments106
View on HN
No posts found.
Grok 4.5 14 days ago

We can assume (outside of Ollama) that they meant the strongest model from each lab. If you limit yourself to just looking at the literal strings in the list, literally none of these are models. What model is "Deepseek" or "GPT"?

How is this any different than what we have already? We've had this ability for ages (6+ months, decades in the AI world), you can literally today easily prompt CC or Codex to use subagents to accomplish tasks and they'll do it well. My entire workflow is one top level orchestrator chat creating tickets to dispatch to subagents to implement, and other subagents to verify. Why is this being sold as a new thing? Have HN users never tried tried asking CC or Codex to use subagents?

As usual HN posters are hyper aware of other's credentials while ignoring that their BS in CS (if that) doesn't magically qualify them to assess everything in every domain.

"I'm a software engineer, I'm sure if I had the time to study Neuroscience, I'd figure out what all of these researchers failed to realize all these decades! I (alone) have the magic of critical and logical thinking"

Frontier as in "Frontier Model" is a legitimate vocabulary term you should probably be aware of in 2026. It's not something the author made up or chose randomly, it's common parlance in the space.

Can you really look yourself in the mirror and say with a straight face that fundamentally nothing has changed about the relationship between the US and its allies? Do you really think Europeans will be quick to forgive the wrongs of this administration? They’ve lost faith in our political system and will, rightfully so, do everything in their power to disentangle with us. The problem with your theory is that they know even if Trump is replaced by someone closer to European social values, our electorate could just as easily completely reverse course in 4 years. It literally already happened. Bush never threatened to annex European territory with military force as far as I can tell. But I understand why in these chaotic times you’d want to gravitate towards hopeful fictions.

Probably a better way to phrase it would be keeping “pace“. Yes, they are still behind but by about the same amount as they always have, they aren’t drifting further behind. Like two marathon runners, one a a mile ahead, but both maintaining the same mile time.

Of course the motivation makes sense on the surface. What I'm getting at is that the supply of capital vs the supply of potential "control of the future" plays feels incredibly imbalanced. Money seems to be so desperate to move into AI it's lost all prudence (the particular people and company mentioned in the OP nonwithstanding, maybe they do deserve 1B).

"not wanting to risk missing out" is essentially just FOMO right? "Smart" money has feels more like FOMO money these days. We literally have shoe companies savying they're going to pivot to AI and having their market cap increase in multiples as reward.

Not the first to notice this I'm sure but it feels like there's an insane amount of pressure pushing capital towards anything with a hint of AI legitimacy. It's as if asset owners across the planet have come to a consensus that the only industry that will matter going forward is this one (fair enough I guess), but this intense systemic pressure squeezes insane amounts of money toward litearlly any AI shaped outlet that opens up. It's just starting to feel like "scared and desperate" money more than "smart money".

Never thought I'd see the day ragebait made it to HN. Yes, let's pretend doing a long jump on the moon is comparable to running a marathon at its prescheduled time at its prescheduled location. Weather is always a factor in sports that take place outside. Might as well put asterisks on all accomplishments that took place on sunny days by your logic right?

DeepSeek v4 3 months ago

I didn't express this well but my interest isn't "who is in the top spot", and is more _why and _how various labs get the results they do. This is also magnified by the fact that I'm not only interested in hosted providers of inference but local models as well. What's your take on the best model to run for coding on 24GB of VRAM locally after the last few weeks of releases? Which harness do you prefer? What quants do you think are best? To use your sports metaphor it's more than following the national leagues but also following college and even high school leagues as well. And the real interest isn't even who's doing well but WHY, at each level.

DeepSeek v4 3 months ago

I’m deeply interested and invested in the field but I could really use a support group for people burnt out from trying to keep up with everything. I feel like we’ve already long since passed the point where we need AI to help us keep up with advancements in AI.

I'd wager that being conscripted in Norwary carries a different level of risk of deployment than being conscripted in the US, given the fact that we've been essentially been nonstop involved in wars for my entire lifetime.

When you were conscripted did you fear you might be sent to Iraq or Afganistan? It just feels like given our history an American conscript will litearlly always have some active warzone to possibly be sent off to. Our contries and our armies are not the same. Is Norway today chomping at the bit to send its soldiers to Iran? Or, per Trump, "our next conquest" Cuba? I really don't think you can think of being drafted into the American army the same way you think of the compulsory service of countries like South Korea or your own.

Being conscripted in a defensive army is materially different than being conscripted into one that takes every opportunity to engage in conflicts across the globe.

NASA Force 3 months ago

Genuinely sorry he let you down and you're left holding the bag dude. But please understand people aren't going to accept your weak rationalizations anymore.

Agreed. The confidence people have to predict what these tools will be capable of two years down the line, when it's barely been over a year since Claude Code was first released, is astounding.

Liked the article in general, but

These apps will win awards at the next all-hands. In two years they’ll be unmaintainable tech debt some poor soul inherits and rewrites from scratch.

Huge assumption/prediction that I think is actually just wrong. There's this weird assumption from a certain crowd, never justified or explained, that tech debt accrued by AI is now, and will forever be, impossible for AI to address, and will for some reason require humans to fix. Working at pace with agents I accrue tech debt every day, then go through the code nightly, again with agents, to clean and tidy everything up.

The more I see this view espoused the more bizzare it seems. People's assumptions seem to be "if AI couldn't one shot this perfectly the first time, then it's useless to try to have it go back over the codebase and identify and address issues". This doesn't match my personal experience at all, second or third passes over code with CC or Codex are almost always helpful and weed out critical issues, but I'm open to hearing from the rest of the HN crowd on their experiences on this.

You can search videos within a channel, go to the channel page and look for the magnifying glass all the way at the end of the nav bar that has

Home | Videos | Shorts | Playlists | Posts | *Magnifying glass here*

Well at least in browser its there, I can't find it on mobile for whatever reason.

[dead] 4 months ago

Their comment reporting stats here:

“Provider: OpenAI (gpt-4o / o1)”

Uh so is it 4o or o1? Very different models. When you read this, how did you interpret this?

    “Suite: 11-task core suite (atomic coding tasks)”
- OK ill take your word for it
    “Configuration: autoroute_first=true, single_file_fast_path=false
Run Variant Token Delta (per call) Step Savings (vs Baseline) Task Success Baseline (2026-03-13) -18.62% — 11/11 Hardened A +8.07% — 11/11 Enhanced (2026-03-27) -6.73% +27.78% 11/11 Key Takeaways:

- What useful information do you glean from this vova_hn2? Perhaps Im just ignorant.

    “The ROI of Precision: While the "Enhanced" run used roughly 6.73% more tokens than the baseline per request, it required 27.78% fewer steps to reach a successful solution.”
So it actually takes MORE tokens but less “steps”? This could all use actual discussion feom the creator. A blog post or detailed comment. Instead we get this.

What sets me off is projects like this that throw random numbers and technical jargon at you because the user simply asked their LLM to do so. It gives the veneer of “oh it must be legitimate look at all the data” to people mentally stuck in 2024 not realizing anyone can generate junk and pass if off in a way that (used to be) convincing.