Takes time to move your hand to it.
HN user
oofbaroomf
Proton is a tool Valve made, based on Wine, to easily run Windows games, on Linux [0]. GP meant Proton.
Have you ever been forced to code something impossible and tedious by a user who keeps getting more and more frustrated as you keep trying?
Wayland is great and I generally prefer it, but it's worth keeping X around for KDE Plasma I think. Things like remote desktop are nicer on X, X is much easier to use on Android compared to Wayland, etc.
If DDG got 28% more visits, Google lost about .6% of their visits.
Egeustimentis.
source: https://www.anarchyishyperbole.com/p/significant-digits.html
Wow. Hopefully, Ternus will bring what he brought to Apple's hardware to their software. The hardware is leaps and bounds ahead of anything else, but their software gets worse and worse every generation. I'm glad to hear this.
Seems like it's down right now. I guess that's the "State of Homelab"? :)
ChatGPT recommended me some good hard drives for price per TB, and one particularly cheap one had direct checkout with Walmart, so I tried it, because why not? It let me get all the way to the payment step before it told me it was out of stock. Walmart's website told me it was out of stock when I decided to click on the link. This is probably part of why it doesn't convert.
I'm currently using a fully vibe-coded, personal River window manager that works just how I want it to. I switched to it after I realized I couldn't do everything I wanted in Hyprland (e.g. tile windows to equal areas instead of BSP by default).
Simple example of how impactful this separation has been for me.
Wait... weren't there many ARM Chromebooks already?
Ugh, I just wish there was a deterministic and formal way to tell a computer what I want...
"Prediction" markets were supposed to be great because of insiders: they make the probabilities much more accurate and actually useful for forecasting.
But they ended up just being for gamblers and there is no more signal.
Is this essentially a cloud-managed specialized subagent with an LLM-friendly API?
Seems like an interesting new category.
20% loss isn't too bad if you start out at double the capacity though.
Ok, but something like Zed is almost as snappy as native GUI frameworks AND has a consistent user experience. It doesn't seem like they are making any tradeoffs there.
a "clear" opinion... :)
I think AI Studio uses the API, so rate limits are extremely high and almost impossible for a normal human to reach if using the paid preview model.
Mathematicians don't do high school math competitions - the benchmark in question is AIME.
Mathematicians generally do novel research, which is hard to optimize for easily. Things like LiveCodeBench (leetcode-style problems), AIME, and MATH (similar to AIME) are often chosen by companies so they can flex their model's capabilities, even if it doesn't perform nearly as well in things real mathematicians and real software engineers do.
The improvement from Claude 3.7 wasn't particularly huge. The improvement from Claude 3, however, was.
Nice to see that Sonnet performs worse than o3 on AIME but better on SWE-Bench. Often, it's easy to optimize math capabilities with RL but much harder to crack software engineering. Good to see what Anthropic is focusing on.
Interesting how Sonnet has a higher SWE-bench Verified score than Opus. Maybe says something about scaling laws.
Wonder why they renamed it from Claude <number> <type> (e.g. Claude 3.7 Sonnet) to Claude <type> <number> (Claude Opus 4).
No. I am referring to Claude 3.5 Sonnet New, released October 22, 2024, with model ID claude-3-5-sonnet-20241022, colloquially referred to as Claude 3.6 Sonnet because of Anthropic's confusing naming.
The SWE-Bench scores are very, very high for an open source model of this size. 46.8% is better than o3-mini (with Agentless-lite) and Claude 3.6 (with AutoCodeRover), but it is a little lower than Claude 3.6 with Anthropic's proprietary scaffold. And considering you can run this for almost free, this is a very extraordinary model.
I wondered if Microsoft would...do this properly
No.
A little bit unrelated, but the rise of Transformers in 2022 was because of the compute available - in 2018, it would have been almost impossible to make something like GPT-4.
GP didn't say Cline/Roo charged anything on top.
How big do you think your index is compared to Google?
Things like Hunyuan 3D are nice for game assets and the like, but they aren't able to really do CAD well. That would be like using Stable Diffusion to code.