It sounds to me like an incentive for new companies to make RAM.
HN user
vectorhacker
So it turns out the comment paper was a joke: https://lawsen.substack.com/p/when-your-joke-paper-goes-vira...
Makes me wonder if LLMs get tired. /s
Yeah, I no longer consider the SWE-bench useful because these models can just "memorize" the solutions to the PRs.
I'd also be interested to know what happened in the last 2-4 months.
I think you've hit the nail on the head there. If these systems of reasoning are truly general then they should be able to perform consistently in the same way a human does across similar tasks, baring some variance.
Stanford was setup by people who came from that tradition.
I'm still not convinced that it's not going through approximate reasoning chain retrieval and that's self-triggered to get more reasoning chains that will maximize it's goal. I'm seeing a lot of comments from other SWEs using it for non-trivial tasks in which it fails at but is just trying harder to look like it's problem solving. Even with more context and documentation, it fails to realize details an experienced SWE would pick up quickly.
I think the fact that it was caught is a testament to the power of open source and the quality of engineering. It speaks volumes about the promise of OSS.
"Good, EU overreach is getting out of hand" my boss said when he saw this article.