HN user

PeterStuer

10,876 karma
Posts2
Comments3,658
View on HN

All depends on context coherence. E.g. in Claude Code with the "1M context" models I am reluctant to push past 30%. So if that Codex 272K is 100% coherent than it is less of a difference. Still, even then the "hard fail" boundary has moved. Longer running agentic processes will surely run into this more often.

Codex Resets 4 days ago

Regulatory capture. You get them to outlaw the part of the competion (safety!) that is unwilling to pricefix and participate in your margin and market division agreements.

"We could have designed our protocols to be minimally compatible with “a nation of laws,” but the tech bros insisted that compromise was treason, and, as a result, we will lose more privacy than necessary"

Unfortunately, no, you can't have a prophilactic that just makes you a little bit pregnant. We used to know this.

The Memory Heist 7 days ago

The 'BOTH WAYS' in the original stood for 'both walking to the school and walking back from school'. ;)

The Memory Heist 7 days ago

Blaming this on Claude is a bit of a reach. The user very carefully goading Claude seems to be key here.

Also interesting buried a bit: user creats a website. Cloudflare immediatly erects a toll gate on it, without even asking.

Humanity has a proven infinite capacity for 'make work'. What that will mean for you depends on whether you will be lounging, poisoning and backstabbing in Capital City, or breaking large rocks into little rocks in the Districts for your next cup of porridge.

Next step? Maps Pro subscription 'blue' plan will guide you through the fastest routes, Maps Free users will be routed on the fast 'green' plan, but might ocasionally be diverted to other routes to 'balance' (free up blue routes') traffic under extreme congestion conditions.

Yep, the virtual universal fast toll road option.

"Ask the model" is usually the point where I realize that either what you need is not a specific answer to their specific question, but a tutor to guide you on a personal journey through the underlying fundamentals of more than one prerequisite areas of understanding, or, the question is too niche or domain specific that I would have to 'ask the model' as wel as I do not know. Now I did observe, not just with LLM but also with Google search, that many people struggle with this and never seem to get good results. They need meta-learning on how to ask the right questions, which I can offer, but often I notice at that point their eyes starting to glaze over as they want a 'direct' solution, me searching or llm guiding for them.

"My main project right now is to establish a framework for large-scale, unsupervised code generation in our codebase... sifting through the unsupervised agent’s (Qwen’s) output"

Ouch. I love my local AI setup with Qwen, but that is a mismatch right there. That model is not the right match for that project. It's like trying to develop a major software solution by just throwing in hundreds of fresh junior programmers and have them spew out random code bits, while what you needed is a good PM, a great architect and a handfull of senior engineers. Might as well pack it in for a year until your model has grown into the ability for those rolls. There is a reason why Opus 4.6 and now Fable dramatically changed the SWE capabilities, and IMHO Qwen is not there yet.

My development platform of choice is Opus 4.8 --effort max. My 'in production' general model just switched from GPT-4o to GPT-5.2. I probably should test local inference for that part (some translation/knowledge extraction/summerization) next but it has not made it to the top of the priority list yet, even though 3 other small specialized task models in the app already run on prem.

Fable 5 is Back 21 days ago

Remember the Clipper Chip back in 1993 because personal computing was considered too dangerous a technologie to be accesible un-controlled/suprvised to the general public?

Fable 5 is Back 21 days ago

So a model named 'Fable 5' is "back". Both excited, as the previous Fable 5 I had access to for just 3 days was fantastic, and anxious as (yes, in my 1 person anecdote) the model refered to as 'Opus 4.8' was stealth nerfed over the last 2 weeks to a degree I had not experienced since the massive nerfs back in the GPT-3.x days. Fingers crossed.