Makes me wonder what the value of a human on a PhD course is
HN user
data_maan
Dm me at jellymax@protonmail.com
Michal Valko's paper should be mentioned. It's discussed here https://youtu.be/UUq4ixTmye8?is=K75EsFIYgyrPKdCI
Another famous dude dumping his thoughts on HN who is gulping it up like an addict.
Add this to the long list of names like Terence Tao, and others who seem to be intellectually incontinent lately in the sense that one cannot navigate this space anymore without encountering their thoughts
Aside from the sad life events, little information is shares about his "system", the thing HN is interested in.
Is this more than a harness built on top of a SOTA commercial LLM?
Was this in the GPT2 paper?
If LLMs lie as much as the OP claims in the article, why can they then solve Olympiad math problems they never saw during training, consistently?
There's the aimoprize.com on Kaggle for example that shows this
More "American minds": https://en.wikipedia.org/wiki/Hartmut_Esslinger
Chief designer at Apple war German.
built with American capital and mostly American minds.
I would say "built with American agency and commercial spirit", not minds.
Most of the things that we have were first built elsewhere (Germany being a prime supplier here with the mp3 or the Zuse), but turning them commercial was the input that came from America.
To be fair, Iran is not pretentious either, killing a few thousand people because they dared to protest.
There are no good guys in this conflict.
A model to whose internals we don't have access solved a problem we didn't knew was in their datasets. Great, I'm impressed
Strategic? Yes.
Moral? Hm. From a moral POV this would be about who has the right to terrorize the Iranian population: the Iranian government or the US/Israel government.
Opinions differ: hobby coders love it, but domain expert secretly despise it because it narrows the gap between the skills they spent years honing and the average Claude, I mean Joe, that just uses this mental exoskeleton.
The people getting pushed out are the intermediates and seniors who aren't high performers.
Also the people that can't market themselves. There are very average programmers that have a large following on X that seem to do very well.
I love these posts that are so on the edge that I can't tell if it's sarcastic or for real :)
What do you mean ? These are top-notch mathematicians
YeS. I didn't dispute that. I disputed that they are NOT top notch ML specialist and have made one of the worst benchmarks of 2025-2026. Benchmarks like these would have worked maybe in early 2024 at latest. The field has moved on significantly since.
And yes, many many other benchmarks don't use toy problems -- their names are just a prompt away.
You are kidding right ? FrontierMath benchmark [1] is produced by a startup whose incentives are dubious to say the least.
They did 1) open source some of their datapoints (on a similar order of magnitude) and 2) they carried out detailed evals. Here is much to learn from their blog posts, much more than from the current dataset.
But fair. If you don't like them, have a look at IMProofBench. Have a look at the AIMO competition. Have a loom at HardMath. It's quite a landscape of datasets already.
Unlike the AI hypesters, these are real mathematicians trying to inject some realism and really test the boundaries of these tools
As mentioned above, realistic benchmarks that are bigger and better exist. Unfortunately, from a benchmarking POV, these mathematicians are the hypesters with a preprint that wouldnt even make it to the AI&Math workshops at ICML or NeurIPS.
If it's the latter case (which it has to be), it seems that attention credit (via, e.g., articles in NY Times) is very unfairly distributed.
None of the people that advanced the state of benchmarking and did the hard work on much bigger benchmarks got any, but a ridiculous benchmark of 10 question scored big.
We will learn if the magical capabilities attributed to these tools are really true or not.
They're not. We already know that. FrontierMath. Yu Tsumura's 553th problem, RealMath benchmark. The list goes on. As I said many times on this thread, there is nothing novel in this benchmark.
This fact that this benchmark is so hyped shows that the community knows nothing, NOTHING, about prior work in this space, which makes me sad.
These problems are representative of the types of subproblems research mathematicians have to solve to get a “research result”. They are finding that LLMs aren’t that useful for mathematical research because they can’t crush these problems along the way. And I assume they put this doc together because they want that to change :)
Same holds true for IMProofBench problems. This dataset shows nothing new.
But everything has been explored in other datasets already.
If only a bunch of mathematicians learn something, why are so many people talking about this, why is the NY Times posting about this?
This is the attention economy at its worst.
It's not angst. It's intense frustration that they 1) are not doing the science correctly, and 2) that others (e.g. FrontierMath) already did everything they claim to be doing, so we won't learn anything new here, but somehow 1stproof get all the credit.
If you want to do this rigorously, you should run it as a competition like the guys at the AI-MO Prize are doing on Kaggle.
That way you get all the necessary data.
I still think this is bro science.
There are some experiments which cannot be carried out more than once
Yes, in which case a very detailed methodology is required: which hardware, runtimes, token counts etc.
This does none of that.
It wasn't like this in any way.
CASP relies on a robust benchmark (not just 10 random proteins), and has clear participation criteria, objective metrics how the eval plays out, etc.
So I stand by my claim: This isn't scientific. If CASP is Japan, a highly organized & civilized society, this is a banana republic.
Yes, but people at those labs may be running those problems because a Fields Medalist is in the paper, and it got hype.
Not because of the problems, and not because this is new methodology.
And once the labs report back, what do we know that we didn't know before? We already know, as humans, the answer to the problems, so that is not it. We already know that LLMs can solve some hard problems, and fail in easy problems, so that is not it either.
So what do we really learn?
Because the companies have the data and can solve them -- so providing the question to a company with the necessary manpower, one cannot guarantee anymore that the solution is not known, and not contained in the training sample.
How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science.
Science should be about reproducibility, and almost nothing here is reproducible.
On the website https://1stproof.org/#about they claim: "This project represents our preliminary efforts to develop an objective and realistic methodology for assessing the capabilities of AI systems to autonomously solve research-level math questions."
Sounds to me to be a benchmark in all but a name. And they failed pretty terribly at achieving what they set out to do.
these are problems of some practical interest, not just performative/competitive maths.
FrontierMath did this a year ago. Where is the novelty here?
a solution is known, but is guaranteed to not be in the training set for any AI.
Wrong, as the questions were poses to commercial AI models and they can solve them.
This paper violates basic benchmarking principles.
Nothing prevents them, and they are already doing that. I work in this field and one can be sure that now, because of the notoriety this preprint got, the questions will be solved soon.