I gave GPT-5.6 Sol, Fable 5, Grok 4.5, Sonnet 5, and GPT-5.5 the same greenfield spec: build the Basecamp 5 frontend and API. Fable won both tracks at $85.87 in 2:06:40. Grok reached 84% of Fable's frontend score and 87% of its backend score for $9.30 in 36:48. Five reruns exposed meaningful variance: the best run beat the median by up to 0.46 points, while Sol’s best frontend beat its first score by 0.72. The full report shows where each model excels, where it is okay, and where it fails.
HN user
aethelyon
Co-Founder <a href="https://klu.ai">https://klu.ai</a>
100% — publish the hidden research, the value is in the discoveries, not in the dividend. With all due respect to the author, it feels like he missed the entire lesson of history.
Built ScreenCommander to solve a personal gap: individual app integrations severely limit what local AI agents can actually achieve. This macOS CLI tool captures your desktop as screenshots, then allows local agents like Codex to interpret and perform actions (clicks, keystrokes, navigation) visually. Requires local permissions for accessibility and screen recording. It's like giving your agent actual eyes and hands, augmenting or bypassing rigid app skills altogether.
Works well with Codex, but not great with Claude Code or Gemini CLI yet (both are bad at novel CLI tools despite having a skills file). Also works well in conjunction with other skills (Atlas, or Apple Script), especially with non-vision models like Spark.
It was initially one-shotted from a GPT-5.2 Pro briefing into Codex-5.3-Codex-xHigh, then iterated on to fix performance issues and expand capabilities.
"some" or a single file?
this is fake news, the xml tags break the output when the model output is the system prompt with the example tags, see screenshot: https://x.com/0xSMW/status/1944624089597137214
same as what happens with claude
comparing o3-pro reasoning to gemini 2.5 pro and claude 4 opus on a speculative, open-ended prompt
No, I’ve seen this pattern as well. Will apologize and then when you ask to continue it will have a change of mind and refuse again. It’s a bad RLHF/AIF loop that it gets stuck into.
Bloop is amazing. Once you use it you stop building your own DIY codebase QA setups.
2 years
This is cool, but the data collection is the hard part, right?
Spoiler: it's fast, cheap, overly protective, and has Kafkaesque DX
Spoiler: it's fast, cheap, overly protective, and has Kafkaesque DX
This is awesome, but there were a couple of great laptop interfaces from that movie too. Spent some quality time in the 90s getting AfterStep/Litestep to look like them.
I used to be worried about face scanning. But sometimes I wonder if it's an inevitable evolution of technology.
Which – to be clear – is not support for it, but a question about what is emergent from the new things we create.
100% agree, I think the 26% will greatly increase over time... or the ones that don't will decline as a business over time.
the 13.4% is likely leaders in ML for some specific use case like fraud or recommendations. it would be great to have access to raw data with anonymized demographics.
we are very, very early.
great data – wish they provided the raw information to slice the respondent audience more, but aligns with what I've seen in the market re: concerns and models.
this is cool
We benchmarked retrieval, GPT-4 turbo vs GPT-4, and fine-tuned several models: https://klu.ai/blog/openai-devday-2023
You can use the result of one here https://huberman.klu.ai/
We benchmarked retrieval, GPT-4 turbo vs GPT-4, and fine-tuned several models. You can use the result of one here https://huberman.klu.ai/
We benchmarked retrieval, GPT-4 turbo vs GPT-4, and fine-tuned several models. You can use the result of one here https://huberman.klu.ai/
check out https://klu.ai – we built it for this reason – sign up, book some time, and I'll help you however I can
Microsoft brought GPT-4 to GA for all customers on Azure OpenAI this week. This removes the endless waitlist for some. Wrote up a few notes from our experience with it.
we built https://klu.ai/ for this
======
outside of us, here's what I see happening
80% of folks aren't building in prod
if you pull apart the 20% that are building, I've seen this from largest to smallest population:
1. most people are not monitoring, followed by 2. home-grown solutions logged into existing observe/analytics platforms, followed by 3. LLMOps tooling like Klu
the 2 cents on the unfortunate truth: I think that many of the AI bolt-on features are living the classic feature lifecycle in that they are launched, no one is monitoring them for improvement, and the feature retention sucks so there's no top-down push to prioritize. the people measuring and improving are exceptional builders regardless of LLMs/RAG.
I started compiling all of the known public information (ala geohotz, semianalysis, et al) in an attempt to build a model card for GPT-4. Am I missing anything?
seems like it, but no one is talking about it – everyone I ask IRL says performance is bad, but not seeing in benchmarks
It depends on what consulting you are doing. You need to match the experience/expertise to the services. Most organizations need help on strategy, implementation, and hiring.
This article seems directionally correct, but poorly written