As part of the Holistic Agent Leaderboard (HAL) initiative at Princeton CITP, we evaluated more than 220 agent runs across 9 benchmarks, the equivalent of over 20,000 agent rollouts across 9 models and 9 benchmarks for a total cost of $40,000. The benchmarks are: AssistantBench, CORE-Bench Hard, GAIA, Online Mind2Web, Scicode, ScienceAgentBench, SWE-bench Verified Mini, TAU-bench Airline, and USACO.
In that process, we “burned” 2.6 billion prompt tokens and learned a lot along the way. In this article, I’d like to share some of the insights we gained, with a particular focus on the GAIA benchmark.
By that definition, the ChatGPT app is now an AI agent. When you use ChatGPT nowadays, you can select different models and complement these models with tools like web search and image creation. It’s no longer a simple text-in / text-out interface. It looks like it is still that, but deep down, it is something new: it is agentic…
https://medium.com/thoughts-on-machine-learning/building-ai-...
Exactly. I think the study is a good reminder that we really have to be careful about the productivity gains attributed to AI. Main takeaway imo, despite limitations from the study, is AI is not a panacea, it can increase productivity, but only if used 'well' and with the good workflows in place, and in the right context.
I can be frustrating at times. but my experience is the more you try the better you become at knowing what to ask and to expect.
But I guess you understand now why some people say vibe coding is a bit overrated: https://www.lycee.ai/blog/why-vibe-coding-is-overrated
and setting up postgresql on a simple VPS is so easy... You can literally ask Gemini 2.5 Pro or o3 or Sonnet 3.7 and do it in 15-30 minutes...
Learned helplessness is really something and vibe coding is overrated imo:
https://www.lycee.ai/blog/why-vibe-coding-is-overrated
why would anyone accept to expose sensitive data so easily with MCP ?
also MCP does not make AI agents more reliable, it just gives them access to more tools, which can decrease reliability in some cases:https://medium.com/thoughts-on-machine-learning/mcp-is-mostl...
true. SAP is more of an entrenched SaaS. Once you have it, the switching costs are too high later on. What an amazing kind of biz. The stickiness is crazy.
so you are telling me that hallucinations (that by definition happen at the model layer) are an engineering problem ? so if we just spin up the right architecture, hallucinations won't be a problem anymore ?
I have doubts
the architecture astronauts are back at it again. instead of spending time talking about solutions, the whole AI space is now spending days and weeks talking about fun new architectures. smh
https://www.lycee.ai/blog/why-mcp-is-mostly-bullshit
So what is it with MCP? It’s just the latest hype-infused hysteria of the architecture astronauts. And please, “don’t let architecture astronauts scare you.” https://www.lycee.ai/blog/why-mcp-is-mostly-bullshit
SAP is kinda of a monopoly no ? The software is bloated and complex to use. The UI is dated but everyone still uses it because it is so embedded into the core arteries of businesses.
So why are people so excited about MCP, and so suddenly? I think you know the answer by now: hype. Mostly hype, with a bit of the classic fascination among software engineers for architecture. You just say Model Context Protocol, server, client, and software engineers get excited because it’s a new approach — it sounds fancy, it sounds serious.
https://www.lycee.ai/blog/why-mcp-is-mostly-bullshit