HN user

oofbaroomf

579 karma
Posts1
Comments57
View on HN

Wow. Hopefully, Ternus will bring what he brought to Apple's hardware to their software. The hardware is leaps and bounds ahead of anything else, but their software gets worse and worse every generation. I'm glad to hear this.

ChatGPT recommended me some good hard drives for price per TB, and one particularly cheap one had direct checkout with Walmart, so I tried it, because why not? It let me get all the way to the payment step before it told me it was out of stock. Walmart's website told me it was out of stock when I decided to click on the link. This is probably part of why it doesn't convert.

I'm currently using a fully vibe-coded, personal River window manager that works just how I want it to. I switched to it after I realized I couldn't do everything I wanted in Hyprland (e.g. tile windows to equal areas instead of BSP by default).

Simple example of how impactful this separation has been for me.

"Prediction" markets were supposed to be great because of insiders: they make the probabilities much more accurate and actually useful for forecasting.

But they ended up just being for gamblers and there is no more signal.

Ok, but something like Zed is almost as snappy as native GUI frameworks AND has a consistent user experience. It doesn't seem like they are making any tradeoffs there.

I think AI Studio uses the API, so rate limits are extremely high and almost impossible for a normal human to reach if using the paid preview model.

Claude 4 1 year ago

Mathematicians don't do high school math competitions - the benchmark in question is AIME.

Mathematicians generally do novel research, which is hard to optimize for easily. Things like LiveCodeBench (leetcode-style problems), AIME, and MATH (similar to AIME) are often chosen by companies so they can flex their model's capabilities, even if it doesn't perform nearly as well in things real mathematicians and real software engineers do.

Claude 4 1 year ago

The improvement from Claude 3.7 wasn't particularly huge. The improvement from Claude 3, however, was.

Claude 4 1 year ago

Nice to see that Sonnet performs worse than o3 on AIME but better on SWE-Bench. Often, it's easy to optimize math capabilities with RL but much harder to crack software engineering. Good to see what Anthropic is focusing on.

Claude 4 1 year ago

Interesting how Sonnet has a higher SWE-bench Verified score than Opus. Maybe says something about scaling laws.

Claude 4 1 year ago

Wonder why they renamed it from Claude <number> <type> (e.g. Claude 3.7 Sonnet) to Claude <type> <number> (Claude Opus 4).

Devstral 1 year ago

No. I am referring to Claude 3.5 Sonnet New, released October 22, 2024, with model ID claude-3-5-sonnet-20241022, colloquially referred to as Claude 3.6 Sonnet because of Anthropic's confusing naming.

Devstral 1 year ago

The SWE-Bench scores are very, very high for an open source model of this size. 46.8% is better than o3-mini (with Agentless-lite) and Claude 3.6 (with AutoCodeRover), but it is a little lower than Claude 3.6 with Anthropic's proprietary scaffold. And considering you can run this for almost free, this is a very extraordinary model.

Things like Hunyuan 3D are nice for game assets and the like, but they aren't able to really do CAD well. That would be like using Stable Diffusion to code.