HN user

pants2

3,254 karma
Posts9
Comments861
View on HN

Right - one would assume that if they're on the verge of an ASI that will consume the entire economy, advertising is only a distraction.

Near the end of 2025, Google was on top - best LLM (Gemini 3 Pro), best video model (Veo 3.1), and best image model (Nano Banana). They seemed unstoppable.

And then, they stopped. Why? Are they having trouble with TPU?

When even Meta can come out with a better LLM than Google, something must be terribly wrong.

Kinda funny that their "cost-vs-performance" chart looks the same as the one for Composer 2.5[1], except that it includes Composer 2.5 at a completely different spot.

What are the chances that CursorBench ranks Cursor's model highest, and Cognition's bench ranks Cognition's model highest? Both are to be RL'd from Kimi as a base model, BTW.

I'd posit that it's not deliberate deception, but for both companies their training data and benchmarks come from the same dataset (Devin/Cursor interaction logs) so they naturally overfit.

1. https://cursor.com/blog/composer-2-5

If you're developing on top of LLM APIs directly, this is definitely not true. There are differences in how context caching works, in what's available through native harnesses, the types of tools you're fine-tuned on (GPT uses apply_patch while Claude uses edit, with different formats), the API surface (Agents SDK, Responses API, Managed Agents), cost structures, and best-practice guidance all around.

Not to mention the meta of account limits, billing, ZDR contracts, etc.

When you're a student in a competitive program at a top university, graded on a curve, and you know your fellow classmates are cheating with AI, you have little choice but to do the same. Especially when jobs for new grads are harder to come by and there's more pressure to also go above and beyond with internships and side projects during your time in school. There's no way to compete without cheating.

Claude Tag 29 days ago

I built something similar at my org. Users simply connect the agent via OAuth and that inherits all of their permissions, so it acts as them.

What's cooler is then it can view/add/remove people from channels, so it can conduct access reviews -- overall I consider it a security improvement.

This seems crazy low to me. AWS has default 3K IOPS and 125 MB/s throughput, meanwhile my Macbook Pro has 700K IOPS and 14.5GB/s throughput.

Is Amazon running on super outdated legacy networking?

Looks very impressive if the benchmarks translate to real world usage!

$1.4/$4.4 pricing and actually served at a pretty comparable price from DeepInfra or CloudFlare! Approx. 1/4th what Opus costs per token.

Looking forward to the artifical analysis test.

CrankGPT 1 month ago

You also need to consider the energy released during the big bang as a prerequisite for creating that food and gasoline. The big bang released about 10^70 J of energy, roughly equivalent to eating 10^63 big macs

LinkedIn offers no way for $company to disavow users who claim to work for $company - they will appear on the official company page as long as it's in their profile.

We've had fake recruiters that claim to work for us running basically the same scam. These are great fake profiles: LinkedIn Premium, tons of relevant posts, etc... but they don't work for us, and we get angry messages from people saying our recruiter tried to scam them. No, they're not our recruiter despite showing up on our company page on LinkedIn. No number of reports could get them taken down.

I finally got it solved by buying drinks for a buddy of mine that works for LinkedIn, but not all startups have that connection!

CrankGPT 1 month ago

If you're comparing raw calories to output, yes. Even gasoline has a caloric value, but humans can't drink gasoline. Growing and preparing food for human consumption uses a lot more energy than pumping and refining gasoline, so at the end of the day, human efficiency gains are not that impressive.