HN user

data_maan

651 karma

Dm me at jellymax@protonmail.com

Posts9
Comments413
View on HN

Another famous dude dumping his thoughts on HN who is gulping it up like an addict.

Add this to the long list of names like Terence Tao, and others who seem to be intellectually incontinent lately in the sense that one cannot navigate this space anymore without encountering their thoughts

If LLMs lie as much as the OP claims in the article, why can they then solve Olympiad math problems they never saw during training, consistently?

There's the aimoprize.com on Kaggle for example that shows this

built with American capital and mostly American minds.

I would say "built with American agency and commercial spirit", not minds.

Most of the things that we have were first built elsewhere (Germany being a prime supplier here with the mp3 or the Zuse), but turning them commercial was the input that came from America.

First Proof 5 months ago

What do you mean ? These are top-notch mathematicians

YeS. I didn't dispute that. I disputed that they are NOT top notch ML specialist and have made one of the worst benchmarks of 2025-2026. Benchmarks like these would have worked maybe in early 2024 at latest. The field has moved on significantly since.

And yes, many many other benchmarks don't use toy problems -- their names are just a prompt away.

You are kidding right ? FrontierMath benchmark [1] is produced by a startup whose incentives are dubious to say the least.

They did 1) open source some of their datapoints (on a similar order of magnitude) and 2) they carried out detailed evals. Here is much to learn from their blog posts, much more than from the current dataset.

But fair. If you don't like them, have a look at IMProofBench. Have a look at the AIMO competition. Have a loom at HardMath. It's quite a landscape of datasets already.

Unlike the AI hypesters, these are real mathematicians trying to inject some realism and really test the boundaries of these tools

As mentioned above, realistic benchmarks that are bigger and better exist. Unfortunately, from a benchmarking POV, these mathematicians are the hypesters with a preprint that wouldnt even make it to the AI&Math workshops at ICML or NeurIPS.

First Proof 5 months ago

If it's the latter case (which it has to be), it seems that attention credit (via, e.g., articles in NY Times) is very unfairly distributed.

None of the people that advanced the state of benchmarking and did the hard work on much bigger benchmarks got any, but a ridiculous benchmark of 10 question scored big.

First Proof 5 months ago

We will learn if the magical capabilities attributed to these tools are really true or not.

They're not. We already know that. FrontierMath. Yu Tsumura's 553th problem, RealMath benchmark. The list goes on. As I said many times on this thread, there is nothing novel in this benchmark.

This fact that this benchmark is so hyped shows that the community knows nothing, NOTHING, about prior work in this space, which makes me sad.

First Proof 5 months ago

These problems are representative of the types of subproblems research mathematicians have to solve to get a “research result”. They are finding that LLMs aren’t that useful for mathematical research because they can’t crush these problems along the way. And I assume they put this doc together because they want that to change :)

Same holds true for IMProofBench problems. This dataset shows nothing new.

First Proof 5 months ago

But everything has been explored in other datasets already.

If only a bunch of mathematicians learn something, why are so many people talking about this, why is the NY Times posting about this?

This is the attention economy at its worst.

First Proof 5 months ago

It's not angst. It's intense frustration that they 1) are not doing the science correctly, and 2) that others (e.g. FrontierMath) already did everything they claim to be doing, so we won't learn anything new here, but somehow 1stproof get all the credit.

First Proof 5 months ago

If you want to do this rigorously, you should run it as a competition like the guys at the AI-MO Prize are doing on Kaggle.

That way you get all the necessary data.

I still think this is bro science.

First Proof 5 months ago

There are some experiments which cannot be carried out more than once

Yes, in which case a very detailed methodology is required: which hardware, runtimes, token counts etc.

This does none of that.

First Proof 5 months ago

It wasn't like this in any way.

CASP relies on a robust benchmark (not just 10 random proteins), and has clear participation criteria, objective metrics how the eval plays out, etc.

So I stand by my claim: This isn't scientific. If CASP is Japan, a highly organized & civilized society, this is a banana republic.

First Proof 5 months ago

Yes, but people at those labs may be running those problems because a Fields Medalist is in the paper, and it got hype.

Not because of the problems, and not because this is new methodology.

And once the labs report back, what do we know that we didn't know before? We already know, as humans, the answer to the problems, so that is not it. We already know that LLMs can solve some hard problems, and fail in easy problems, so that is not it either.

So what do we really learn?

First Proof 6 months ago

Because the companies have the data and can solve them -- so providing the question to a company with the necessary manpower, one cannot guarantee anymore that the solution is not known, and not contained in the training sample.

First Proof 6 months ago

How is that interesting for a scientific point of view? This seems more like a social experiment dressed as science.

Science should be about reproducibility, and almost nothing here is reproducible.

First Proof 6 months ago

On the website https://1stproof.org/#about they claim: "This project represents our preliminary efforts to develop an objective and realistic methodology for assessing the capabilities of AI systems to autonomously solve research-level math questions."

Sounds to me to be a benchmark in all but a name. And they failed pretty terribly at achieving what they set out to do.

First Proof 6 months ago

these are problems of some practical interest, not just performative/competitive maths.

FrontierMath did this a year ago. Where is the novelty here?

a solution is known, but is guaranteed to not be in the training set for any AI.

Wrong, as the questions were poses to commercial AI models and they can solve them.

This paper violates basic benchmarking principles.

First Proof 6 months ago

Nothing prevents them, and they are already doing that. I work in this field and one can be sure that now, because of the notoriety this preprint got, the questions will be solved soon.