Exactly right. There are many different domains that AI can be helpful in, and I am frustrated by how it's used to process information in a lossy manner.
HN user
deanalyzer
Yes, and what's scary is that I can easily imagine HR departments loving these tools and using them at scale.
this looks incredibly useful! I've also run into similar problems with code review and been building similar tooling with the same idea (reconstructing the changes into entities, and finding focal/important changes), but haven't gone as far as this.
This is nicely put.
One possible implication is that this calls for a new wave of developer tools designed to introduce friction between human cognition and code. This would be a very different philosophy from that of developer tools in the past.
This could be interesting, but it badly needs systematic benchmarking results. It is not difficult to get Claude Code or Codex to install and run a solver locally, so the tool’s current value proposition is fairly muddled.
If there were evidence that it offered better performance, I might consider running larger workloads on it.
Yes, I wonder how the verdicts would hold under a blinded test. This analysis read like the authors going out of their way to be supportive of Grok.
I learned about Benford's law over a decade ago, and I always found it beautiful and elegant. But surely, fraudsters have become more sophisticated by now. I wonder if you asked an AI to commit fraud, if it would be clever enough to avoid such mistakes.