Is this an open source tool? Is it something we have to pay for? The site really doesn't tell me what I should be expecting.
HN user
mdasen
"Just as most issues are seldom black or white, so are most good solutions seldom black or white. Beware of the solution that requires one side to be totally the loser and the other side to be totally the winner. The reason there are two sides to begin with usually is because neither side has all the facts. Therefore, when the wise mediator effects a compromise, he is not acting from political motivation. Rather, he is acting from a deep sense of respect for the whole truth."
~Stephen R. Schwambach
But to run the entire benchmark it cost $727 with Gemini 3.6 Flash and $925 with GLM-5.2, $198 (21.4%) less. I tend to look at the cost to run the whole index rather than the weighted average cost per task.
LMArena's "code" leaderboard is really skewed since it's a front-end JS code and design leaderboard. It generates a demo app with two models and then asks "do you prefer A or B". People can look at the code, but most of the time it's just going to be which one looks nicer.
Models that people like the design aesthetic of (Claude, GLM) tend to do better in LMArena than they do on other benchmarks. Design matters, but you look at a model like GPT-5.5 and it's behind Kimi K2.6, Sonnet 4.6, Qwen3.7 Max, and GLM-5.1 on LMArena's code leaderboard. Then you look at benchmarks like DeepSWE and GPT-5.5 blows them out of the water with only Fable and GPT-5.6 beating it.
I'm not saying that the LMArena leaderboard isn't useful, but I'm not sure how much weight I'd give it as a "code" leaderboard. I think often times it's a design comparison of simple front-end React apps rather than a coding comparison. GLM-5.2 is a very good model, but when you look at DeepSWE or Terminal-Bench v2, GPT-5.5 is well ahead.
It also depends on how many tokens it needs to burn through to accomplish something.
At this point, I always look at things like Artificial Analysis' total cost to run their tests. It'll take into consideration the cost of tokens, how many tokens it burns through, and how effectively it uses caching (and the price of that caching).
If a model "costs the same" but its reasoning ends up going through a ton more tokens, it doesn't really cost the same in real world usage.
When OnePlus started, they were considerably cheaper than flagship phones from others. At $299, the OnePlus One was a ton less than the $650 you'd pay for an iPhone 6 or Galaxy S5. You were getting a 95% flagship phone at half the price. You could get a OnePlus One with the latest Qualcomm Snapdragon or you could get a Samsung with 30% the performance and a 640x480 low-res display for the same price.
I feel like the "something new" was price. Over time, that price kept creeping up. Yes, it went from being a 95% flagship to being a 100% flagship, but it also went from being half price to full price.
It was also cool that it used Cyanogenmod which meant you got a community OS that actually got updates, but over time other manufacturers started offering updates for their phones (rather than abandoning them soon after manufacturing). And that was something new other than price. But I think the big thing was that it was a half-price phone when it launched. In 2014, it was just such an amazing deal. Today, it's the same price as Samsung phones.
I'll add:
- Telegram had usernames in 2014 before Signal added them a decade later, allowing people to chat without sharing their phone number
- Telegram has unencrypted chats which allow for giant chat rooms of 200,000+ and channels with millions of subscribers. Signal warns about performance issues when you have more than 150 people in a group. Telegram isn't just a messenger - it's often used as a social publishing platform like Instagram.
I don't use Telegram and use Signal a lot, but I also understand why other people use Telegram: the same reason they use Instagram.
I'm a bit skeptical.
Cursor's benchmark finds that Cursor's model (Composer 2.5) is basically as good as Opus 4.8 max and GPT-5.5 xhigh, but at a fraction of the price.
Artificial Analysis' testing shows Composer 2.5 to be pretty far behind: https://artificialanalysis.ai/agents/coding-agents. You look at the DeepSWE benchmark (which is probably the hardest to game at this point) and GPT-5.5 xhigh gets a 64, Opus 4.8 max gets 56, and Cursor 2.5 gets 16.
I don't doubt that Cursor works well for some people. It's beating DeepSeek v4 Pro in the DeepSWE benchmark and that's a very capable model. But I'm skeptical of the claims that it's a competitor for Opus 4.8 and GPT-5.5. It just seems convenient that their model does so well on their own benchmark while third party benchmarks have it far behind. Maybe it's a really great benchmark and a better measure than third party ones - I'd love for a cheap model to do as well as the expensive ones.
What it's saying is that the M6 will be released, but not the M6 Pro or M6 Max. Instead, Apple will wait to release new Max/Pro chips for a future generation.
It's not simply marketing since the Pro/Max chips of a generation use the same cores as the regular version, just more of them or different combinations of performance and efficiency cores.
You have the ability to move, as long as Bluesky Social PBC allows it.
They hold the keys for your DID. If they don't allow you to move to another PDS, you can't move. The original theory was that you'd hold the private keys, but that's something that would hugely limit adoption so they decided to hold the keys themselves.
In terms of moving your backlog of posts to a new server, part of the issue is liability (not merely legal liability, but reputational as well). When you have a user on your platform and they're posting stuff, you're moderating them in real time. If they turn out to be a horrible troll, you've get the reports. Let's say a horrible troll has been on EvilServer and EvilServer has been ignoring the reports against them. They now want to move to your GoodServer and bring all their post history with them. As an admin of GoodServer, you can't see that everyone has been reporting this troll for years. They're now moving over lots of horrible, inflammatory, potentially illegal posts to your server.
Where do you see that? I see they have GPT-5.5 (xhigh) at 55, GPT-5.5 (high) at 53, and Muse Spark at 43. Muse Spark does beat GPT-5.4 mini (xhigh) which scores 40, but the key there is "mini".
In the coding index, GPT-5.5 gets 59.1, 58.5, 56.2, and 52.1 for xhigh, high, medium, and low while Muse Spark is behind at 47.5. For agentic, GPT-5.5 gets 74.1, 72.0, 69.4, and 59.7 (xhigh, high, medium, low) while Muse Spark gets 62.0 (beating only GPT-5.5 low).
GPT-5.5 only gets beaten by Opus 4.8 in their general index, is the top spot for coding, and is #3 behind Opus 4.8 and GLM-5.2 for agentic (excluding Fable 5 which takes the top spot, but is unavailable).
Not really. You could allow private.icloud.com only if they're using Apple's SSO. If someone tries to create an account not using Apple's SSO, then you don't allow private.icloud.com email addresses.
I find that I don't use a ton of output tokens. I'm usually around 95% cached input, 4% input, and 1% output.
For me, the big thing with MiMo-V2.5-Pro and DeepSeek V4-Pro is that cached inputs are practically free. Kimi K2.7 Code is 53x more expensive for cached inputs which is 95% of my costs.
If I use 95M cached input tokens, 4M input tokens, and 1M output tokens, that'd be: $18 for cached input on Kimi K2.7 Code vs $0.34 with MiMo/DS; $3.80 for inputs on Kimi vs $1.74 with MiMo/DS; and $4 for output on Kimi vs $0.87 with MiMo/DS.
Of all the places where I'm accumulating costs by using Kimi, it's the cached inputs. The real savings with MiMo/DS's price cut is the cached inputs.
Yes, it's a "smaller" (137B) model that competes with Haiku, but it's basically the performance of Qwen3.6-35B-A3B which is 75% smaller and 98% smaller in terms of active parameters (since it's a mixture of experts model). Microsoft should be comparing its model to good smaller models, not Haiku 4.5.
Qwen-3.6-27b is closer to Claude Opus 4.7 than it is to Haiku 4.5 in a lot of benchmarks - and it's way smaller than Microsoft's new model.
Sure, it competes with Haiku, but it shows how far Microsoft is behind lots of other small models that are available.
.
The point of the article isn't "Malawi is poor compared to Europe," but rather comparing Malawi to other countries that were similarly colonized and how other colonized countries have done a lot better than Malawi - despite often having more adversity in their post-colonial existence.
Yes, being exploited will leave you in a bad state, but it's also important to learn why other similarly colonized countries have done a lot better over the past 30 years - what are the conditions and policies that improve things
I too don't want to write OS-specific stuff, but here's some counter arguments.
With egui, it's an immediate mode GUI rather than retained mode and that has trade-offs: https://github.com/emilk/egui#why-immediate-mode. It's going to use more CPU (and battery power), there can be jitter and things shifting after the initial rendering, and other stuff. I think egui is very different from most cross-platform and platform-specific libraries.
With .NET MAUI, you're getting native controls, but you're now using a layer that's trying to use native controls on the underlying systems that don't always align completely. A lot of things act mostly the same across systems, but some things don't totally.
With Flutter, your app is going to be larger in part because you're shipping a rendering engine, runtime, widgets, etc. Does it have the look and feel you want? Maybe. That's a bit subjective. Does it handle all the little things correctly? When I'm using an app, I want it to scroll like how I'm used to scrolling working on my system. If you have differently styled buttons, I don't care, but if the scrolling feels wrong, it's going to annoy me. And there's so many little things.
Frankly, one of the reasons why Electron often does well is that a lot of the little things "feel right" because the UI is essentially a Chromium-rendered web page which users are used to interacting with. But that has downsides too - shipping a web browser with your app and the memory usage.
Heck, Qt apps in Gnome or GTK+ apps in KDE can look/feel "off".
And it'll all depend on your ecosystem. Often cross-platform solutions are lacking in accessibility - sometimes completely missing, sometimes half-baked and it works in some parts and not in others or just is janky. Memory usage is often higher. Many little things that make an app feel right might not be there. Many have slower startup times since they're loading a bunch of stuff that native apps don't need to. And it really depends on what approach the cross-platform library is taking to determine what is going to cause pain.
So you kinda have to pick your poison and what's acceptable to you will vary depending on your goals and tastes. Maybe React Native is the way to go for you with lots of native controls available and the feel that provides and the performance and size is acceptable.
If you create a Flutter or Kotlin Compose Multiplatform or AvaloniaUI app and put it on the web, it's not going to feel right as something like HN does. Right-click, text selection, etc. are all going to be different or missing. If you're creating a solitaire game, maybe that doesn't matter - you get desktop and web in one go and it's not a big deal.
But you have to know what you're building to know if the trade-offs being made are good ones. This isn't meant to sound anti-cross-platform, but as someone who has suffered some pain in this area, I guess I just wanted to impart that it isn't all sunshine and rainbows. Some times it can still be worth it, but just go in with your eyes open.
electric-drivetrain with onboard gasoline generator
Generally speaking, it's more efficient to power a car using a series-parallel hybrid system than an electric drivetrain with generator (series hybrid) while not really being any more complicated.
In a series hybrid (electric with generator), you're losing energy converting the rotational energy into electric energy. It's better to use the engine's output to power the wheels while it's in an efficient range. It's why Toyota's series-parallel hybrid design offered better mileage than vehicles that (primarily or fully) operated as series hybrids like the Chevy Volt.
No screens
You can't really sell a car without a screen due to government regulations which require backup cameras (since 2018 in North America, since 2022 in the EU and Japan).
no assists
Automatic Emergency Braking is going to be required in the US in 2029 (detecting frontal crashes about to happen and automatically braking, including pedestrian detection).
The EU requires even more including blind spot detection and lane-keeping assist.
I certainly agree that cars need knobs and buttons for controls like AC/heat, music, etc. However, it'd be hard to make a car where you aren't putting in a screen and assistive technology. I think a better argument would be to make a car where the screen was simply Apple CarPlay/Android Auto and a backup camera - rather than shoving a lot of garbage UX into it.
It was about x64 being unable to keep up - independent of Intel’s Fab capabilities which have improved lately.
But the big reason x64 couldn't keep up was that Intel's fab capabilities were horrible. Intel got stuck and couldn't get smaller nodes out and competing fabs caught up and left Intel in the dust.
Apple was able to ship 22nm Intel processors in Summer 2012 while their iPhone processors were 32nm that Fall and 28nm in Fall 2013. Spring 2015, Apple shipped 14nm Intel laptops and later that Fall 14/16nm iPhones. Competitors had caught up and soon TSMC started surpassing Intel.
Yes, Intel's fab capabilities have improved lately, but Intel's fab failures were causing x64 to fall behind. If Intel had retained fab supremacy, x64 wouldn't have fallen behind. I think Apple still likes the idea of being able to build exactly the parts they want (so they can optimize for power, thermals, etc), but Intel fell behind because their fabs stopped being competitive.
It's really interesting how much the AI harness seems to matter. Going from 48% via Google's official results to 65% is a huge jump. I feel like I'm constantly seeing results that compare models and rarely seeing results that compare harnesses.
Is there a leaderboard out there comparing harness results using the same models?
Sure. Dirac is just a fork of the Cline harness and obviously OpenCode could take the same techniques and implement them. I don't know how difficult it would be to implement them in OpenCode, but given that Dirac and OpenCode are both open source, a future version of OpenCode could always be a re-branded Dirac (I'm sure there are ways to implement Dirac's techniques without having to completely replace OpenCode's underlying code base, but this illustrates that at the extreme, they could clearly just take Dirac in its entirety to get the same results).
Bad actors can strip sources out
I think the issue is that it's not just bad actors. It's every social platform that strips out metadata. If I post an image on Instagram, Facebook, or anywhere else, they're going to strip the metadata for my privacy. Sometimes the exif data has geo coordinates. Other times it's less private data like the file name, file create/access/modification times, and the kind of device it was taken on (like iPhone 16 Pro Max).
Usually, they strip out everything and that's likely to include C2PA unless they start whitelisting that to be kept or even using it to flag images on their site as AI.
But for now, it's not just bad actors stripping out metadata. It's most sites that images are posted on.
In the early days, Netflix benefited from other media companies not recognizing streaming for what it was: their replacement. They licensed content to Netflix cheaply without thinking about how it would impact DVD sales or cable tv subscriptions.
It's kinda like how IBM didn't see the value in software and that let Microsoft become Microsoft.
If this does bypass their own (and others') anti-AI crawl measures, it'd basically mean that the only people who can't crawl are those without money.
We're creating an internet that is becoming self-reinforcing for those who already have power and harder for anyone else. As crawling becomes difficult and expensive, only those with previously collected datasets get to play. I certainly understand individual sites wanting to limit access, but it seems unlikely that they're limiting access to the big players - and maybe even helping them since others won't be able to compete as well.
As someone clumsy, I'm so grateful that my MacBook Air can take a beating. It has one slight dent of about 1mm in the 4 years I've had it and I definitely drop it or knock it off a desk or something a few times a year.
I'll take the extra weight of aluminum (0.3lb, 130g). Yes, someone might say the ThinkPad X1 Carbon is 14", but the 13" MacBook Air actually has a 13.6" screen.
If I were in the market for a PC laptop, I'd definitely take a look at the ThinkPad X1 Carbon, but I'm also not worried about the weight of my MacBook Air. The X1 Carbon Intel ones are on sale right now since Panther Lake will be a huge upgrade coming soon, but even on clearance they aren't cheap. An X1 Carbon with 32GB RAM and 1TB storage (Ultra 7 268V, the cheapest one due to the sale) will cost $1,679 while a similar MacBook Air will cost $1,699 - and the M5 has 48% better single-core performance and 56% better multi-core performance (Geekbench). A 16GB/512GB (Ultra 5 225U) X1 Carbon is $1,538 compared to $1,099 for a MacBook Air - and the M5 has a 74% single and multi core advantage there.
Panther Lake might narrow the performance gap, but early indicators don't seem like that's the case. Even the top of the line Ultra X9 388H sees the M5 with a 36% single-core advantage while the Ultra X9 388H gets 3% faster multi-core. And I'm not sure the higher wattage "H" processors work for something like an X1 Carbon.
The highest non-H Panther Lake processor (Ultra 7 365) sees the M5 get 51% better single-core and 58% better multi-core. Maybe we'll see better, but it looks like Intel isn't closing the gap in 2026.
For fascism, it's not always about getting something you think is a lot. It's about a power relationship. Trump has demonstrated that Nvidia will bow to his will.
It's also potentially an implementation of the foot-in-the-door technique (https://www.simplypsychology.org/compliance.html). It's a common manipulative strategy where you get someone to do a small favor for you which makes them much more likely to do a large favor for you later.
No. When you go into a Costco, Costco is a retailer who bought merchandise to sell to you. When you go to Amazon, a large amount of the products are being sold by third party vendors while Amazon is taking a large cut.
That doesn't seem to be the case across Europe based on current sales.
Looking at marketshare in the EU+EFTA+UK 2025 to 2026:
VW Group went from 26.8% to 26.7%. Stellantis went from 15.5% to 17.1%. Renault Group went from 9.8% to 8.7%. Hyundai Group 8.4% to 7.6%. BMW Group 7.0% to 6.9%. Toyota Group 8.0% to 7.2%. SAIC Motor was flat at 2.0%. BYD 0.7% to 1.9%. Tesla 1.0% to 0.8%.
So it doesn't really seem like BYD is eating into the sales of European manufacturers yet. VW + Stellantis + Renault + BMW + Mercedes + Volvo + Jaguar Land Rover was 66.9% in 2025 and it's 67.1% in 2026, an increase of 0.2 percentage points (looking at just VW + Stellantis + Renault, it was an increase of 0.4pp).
We'll see what happens going forward, but Chinese cars aren't killing it yet. SAIC Motor is flat. BYD is doing very well, but it's a lot easier to grow when you're small. I think that Chinese cars will present challenges, but I'm less sure that it's over for European automakers. Right now, European automakers are marginally increasing their marketshare (probably more noise than anything, but not evidence of decline).
I think BYD is a strong company and I think they'll continue to gain marketshare, but will others? SAIC has seen modest European growth since 2024, but nothing really threatening and they're sitting at 2% marketshare and their modest growth seems to becoming no growth. Chery is really small. Geely is ultra small without Volvo.
So it feels like it's really the BYD story. BYD is the company actually making inroads and growing at a significant rate. And I don't think that a single company can destroy the European auto industry. It's possible BYD could become 10-20% of the European market and that would be a major win for them and make a significant dent in competitors. But do you see them becoming more? Are there other companies that seem promising?
I think the big thing keeping Blazor back is that C# doesn't work well with WASM. It was built at a time when JIT-optimized languages with a larger runtime were in-vogue. That's fine in a lot of cases, but it means that C# isn't well suited for shipping a small amount of code over the wire to browsers. A Blazor payload is going to end up being over 4MB. If you use ahead of time compilation, that can balloon to 3x more. The fact that C# offers internal pointers makes it incompatible with the current WASM GC implementation.
Blazor performance is around 3x slower than React, it'll use 15-20x more RAM, and it's 20x larger over the wire. I think if Blazor could match React performance, it'd be quite popular. As it stands, it's hard to seriously consider it for something where users have other options.
Microsoft has been working to make C#/.NET better for AOT compilation, but it's tough. Java has been going through this too. I don't really know what state it's at, but (for example) when you have a lot of libraries doing runtime code generation, that's fine when you have a JIT compiler running the program. Any new code generated at runtime can be run and optimized like any other code that it's running.
People do underappreciate the JS/TS ecosystem, but I think there are other reasons holding back stuff running on WASM. With Blazor, performance, memory usage, and payload size are big issues. With Flutter and Compose Multiplatform, neither is giving you a normal HTML page and instead just renders onto a canvas. With Rust, projects like Dioxus are small and relatively new. And before WASM GC and the shared heap, there was always more overhead for anything doing DOM stuff. WASM GC is also pretty new - it's only been a little over a year since all the major browsers supported it. We're really in the infancy of other languages in the browser.
I definitely agree with the first point - it's not meant to be the best.
On the second part, I think the big thing was that they needed something that would interop with Objective-C well and that's not something that any language was going to do if Apple didn't make it. Swift gave Apple something that software engineers would like a ton more than Objective-C.
I think it's also important to remember that in 2010/2014 (when swift started and when it was released), the ecosystem was a lot different. Oracle v Google was still going on and wasn't finished until 2021. So Java really wasn't on the table. Kotlin hit 1.0 in 2016 and really wasn't at a stage to be used when Apple was creating Swift. Rust was still undergoing massive changes.
And a big part of it was simply that they wanted something that would be an easy transition from Objective-C without requiring a lot of bridging or wrappers. Swift accomplished that, but it also meant that a lot of decisions around Swift were made to accommodate Apple, not things that might be generally useful to the lager community.
All languages have this to an extent. For example, Go uses a non-copying GC because Google wanted it to work with their existing C++ code more easily. Copying GCs are hard to get 100% correct when you're dealing with an outside runtime that doesn't expect things to be moved around in memory. This decision probably isn't what would be the best for most of the non-Google community, but it's also something that could be reconsidered in the future since it's an implementation detail rather than a language detail.
I'm not sure any non-Apple language would have bent over backwards to accommodate Objective-C. But also, what would Apple have chosen circa-2010 when work on Swift started? Go was (and to an extent still is) "we only do things these three Googlers think is a good idea", Go was basically brand-new at the time, and even today Go doesn't really have a UI framework. Kotlin hadn't been released when work started on Swift. C# was still closed source. Rust hadn't appeared yet and was still undergoing a lot of big changes through Swift's release. Python and other dynamic languages weren't going to fit the bill. There really wasn't anything that existed then which could have been used instead of Swift. Maybe D could have been used.
But also, is Swift bad? I think that some of the type inference stuff that makes compiles slow is genuinely a bad choice and I think the language could have used a little more editing, but it's pretty good. What's better that doesn't come with a garbage collector? I think Rust's borrow checker would have pissed off way too many people. I think Apple needed a language without a garbage collector for their desktop OS and it's also meant better battery life and lower RAM usage on mobile.
If you're looking for a language that doesn't have a garbage collector, what's better? Heck, what's even available? Zig is nice, but you're kinda doing manual memory management. I like Rust, but it's a much steeper learning curve than most languages. There's Nim, but its ARC-style system came 5+ years after Swift's introduction.
So even today and even without Objective-C, it's hard to see a language that would fit what Apple wants: a safe, non-GC language that doesn't require Rust-style stuff.
Is there a congruent DGGS that you would recommend?