HN user

vectorhacker

35 karma
Posts5
Comments10
View on HN
OpenAI O3-Mini 1 year ago

Yeah, I no longer consider the SWE-bench useful because these models can just "memorize" the solutions to the PRs.

I think you've hit the nail on the head there. If these systems of reasoning are truly general then they should be able to perform consistently in the same way a human does across similar tasks, baring some variance.

I'm still not convinced that it's not going through approximate reasoning chain retrieval and that's self-triggered to get more reasoning chains that will maximize it's goal. I'm seeing a lot of comments from other SWEs using it for non-trivial tasks in which it fails at but is just trying harder to look like it's problem solving. Even with more context and documentation, it fails to realize details an experienced SWE would pick up quickly.