Just to nitpick the math. If you are going to fire 50% of the company, the AI tools should actually make the remaining people 100% more efficient, not 50% :)
HN user
Snuggly73
it has been pretty much a benchmark for memorization for a while. there is a paper on the subject somewhere.
swe bench pro public is newer, but its not live, so it will get slowly memorized as well. the private dataset is more interesting, as are the results there:
I mean, its right there in their blog - https://cursor.com/blog/scaling-agents
"We've deployed trillions of tokens across these agents toward a single goal. The system isn't perfectly efficient, but it's far more effective than we expected."
if it’s too hard for you to write, it’s too hard for you to understand and comprehend. how are you going to take responsibility for that code and maintain it if needed?
Well, could it be because it was instructed to kinda "study" Servo?
https://github.com/wilsonzlin/fastrender/blob/3e5bc78b075645...
I've watched them today work in the new repo - https://github.com/wilson-anysphere/fastrender/tree/main , adding another 50k lines trying to optimize scroll/rendering performance (spoiler: not really)
At this point, its 1.5mlocs without the vendored crates (so basically excluding the js engine etc). If you compare that to Servo/Ladybird which are 300k locs each and actually happen to work, agents do love slinging slop.
Not sure - if it works, then who needs Cursor (and all other IDEs). You just ask for a browser and it comes out of the thin air.
This is from the "official" build - https://imgur.com/fqGLjSA
The "in progress" build has a slightly different rendering but the same result
Noticed that as well - I think it was “manual”
The latest commit now builds and runs (at least on my Mac). It’s tragically broken and the code is…dunno…something. 3m lines of something.
I couldn’t make it render the apple page that was on the Cursor promo. Maybe they’ve used some other build.
And there is the thing about the cost. The blog post says that they've spent trillions (plural!) of tokens on that experiment.
Looking at OAI API pricing, 5.2 Codex is $14 per 1 million output tokens. Which makes cool $14m for 1 trillion tokens (multiplied by whatever the plural is). For something that "kind of works".
Its a nice ad for OAI and Anysphere, but maybe next time - just donate the money to a browser team?
The only thing that I got to actually run on WSL2 was the "Excel" (couldnt get anything actually to compile on Mac or Windows).
It a broken mess that probably implements 0.00001% of Excel. And its 1.2m locs.
With codebases developed in this way - either they need to figure out how agents are going to maintain them (in which case SWE as we know is dead - it will only be limited to those that can spend trillions of tokens, or they are going to remain weird demos.
error: could not compile `fastrender` (lib) due to 34 previous errors; 94 warnings emitted
I guess probably at some point, something compiled, but cba to try to find that commit. I guess they should've left it in a better state before doing that blog post.
I mean...the naive approach for a prime number check is o(n) which is linear. Probably u've meant constant time?
Well, for some reason it doesnt let me respond to the child comments :(
The problem (which should be obvious) is that with a/b real you cant construct an exhaustive input/output set. The test case can just prove the presence of a bug, but not its absence.
Another category of problems that you cant just test and have to prove is concurrency problems.
And so forth and so on.
I mean "have been bad" doesnt exclude "getting worse" right :)
Test cases are great, but not a total solution. Can you write a test case for the add_numbers(a, b) function?
Yes, and for some cases no.
The models are gotten very good, but I rather have an obviously broken pile of crap that I can spot immediately, than something that is deep fried with RL to always succeed, but has subtle problems that someone will lgtm :( I guess its not much different with human written code, but the models seem to have weirdly inhuman failures - like, you would just skim some code, cause you just cant believe that anyone can do it wrong, and it turns out to be.
Thanx. More of a "faster keyboard" so far then?
And yeah - if I had a crystal ball, I would be on my private island instead of hanging on HN :)
The article is arguing that it will basically replace devs. Do you think it can replace you basically one-shotting features/bugs in Zed?
And also - doesn’t that make Zed (and other editors) pointless?
Ok, if its almighty, then why is not the benchmarks at 100%? If you look at the individual issues, those are somewhat small and trivial changes in existing codebases.
(note that if you look at individual slices, Opus is getting often outperformed by Sonnet).
This type of comment implies that it’s going to stop with “them” and somehow “us that adopted the LLM” will be the winners. The goal is full automation, there is no “adapt or be left behind”.
I am going to prefix this with that I could be completely wrong.
Simon - you are an outlier in the sense that basically your job is to play with LLMs. You don't have stakeholders with requirements that they themselves don't understand, you don't have to go to meetings, deal with a team, shout at people, do PRs etc., etc. The whole SDLC/process of SWE is compressed for you.
i thought it might be something like this (still a weird overkill), but if you are effectively replacing the parser with new peg and replacing the backend with something new - then there is nothing left - just start from scratch :)
looking at the "att" branches (excuse my unhealthy curiosity) I can only say - "jesus fucking christ".
from the old parser ast -> to json -> to new ast representation (that is basically again copy of the old one) -> to some new incomplete bytecode generation
im sure there is some good explanation, but....why?! :)
reflection seems slightly wrong as well
Neither :(
LCB Pro are leet code style questions and SWE bench verified is heavily benchmaxxed very old python tasks.
Ignoring the tests, the first change was adding a single parent id column and the second "more complex" refactoring added few more hash columns to the table (after you've specified that you wanted them, i.e. not an open-ended question)
Its a very impressive model, but I think we have different views on what is complex.
If you like CC - I'll just leave this here - https://github.com/sst/opencode
Not sure how it went in their tests - I've tried Opus and GPT5 and it was few lines of react + tests, so I guess 'no'