HN user

horhay

24 karma
Posts0
Comments40
View on HN
No posts found.

It's remarkable, its not out of the bounds of the pattern of success that AI has had with math recently to the point that people should sound alarm bells.

A lot of the weight this holds is the fact that it's an old problem and that its difficulty hinges on the lack of investigation the disproof side of hypothesis. The model basically took a contrarian path and found tools and methods that support that a disproof is viable. So the (unquantified amount of) mathematicians out there were all dedicating their resources on the notion that this can be proved. Some with hindsight would say that if they a had team of experts who are driven to the goal of disproof that this would have been achievable by humans, and one of the mathematicians of the paper state as much,this still has value in terms of reliability measurement, and possibly human-aided endeavors when the methods scrounged by the model can be used in other solutions.

I think the intention of this paper is to build some type of culture of "math generalists" that don't quite exist in today's academia. The thing is, is that a good half of the people in that paper were actually very pragmatic on the implications of such a success and present questions in terms of the measurability of the difficulty of the problem and the generalizability of the solution provided for other questions. Gowers in particular offers no resistance and in fact resorts to the theatrics of "being the bearer of bad news" on Twitter for some reason.

As with Tao, he's always been a measured optimist even before the tools were consistently usable for his work. And even still nowadays, he adds stipulations to his statements on the successes of AI. Yes, he's part of Math Inc. now and is in close contact with Google Deepmind for some projects but his interest lies in using the tools today. Gowers has been hypothesizing on the future of math in the tone he has taken now ever since o3/GPT5. There's no comparison between the two who should attract more scrutiny.

This part of the announcement holds no value besides maybe taking a shot at the Deepmind Co-Mathematician paper. Nearly every mathematical success they've achieved around the GPT 5.2 generation has been done with general (and even public) models. Their last bountied problem solved was done with 5.4 Pro, also a general model.

There may be years of investigation as to how far you can generalize these methods. As to how central it is, it's a longstanding problem that Erdos loves to cite for that branch of math.

The thing is is that it seems a lot of the effort through the years (which is unquantifiable in scale as to how much time was spent and how many people focused their entire worklives on it if any) has gone for trying to look for the proof, and the search for the disproof seems minimal.

It's a very complicated matter honestly. This is a new height that AI has reached, even though it follows the usual methods of success that it has had.

What strikes me as unusual though is that they do make a point of saying things like "this is a general purpose model that wasn't trained on the problem" among a few other things as if that's new. The last bountied problem they accomplished used a public model that ALSO didn't rely on specialized training. And that didn't make their blog.

I'm not sure your characterization of Tao is accurate lol. In that companion paper, only Gowers seems to extensively show no pragmatism in the implications of this accomplishment. Even the younger math experts in that paper were a lot more cautious with their statements. Tao seems to follow that same tune most of the time even though he uses AI for first-pass inspections of solutions brought to his attention.

Nano Banana Pro 8 months ago

It still has some artifacts more often than not, they are a lot subtler in nature but they still come out, whether it's texture, proportion, lighting, or perspective. Now some things are easier to fix on second pass edits, some are not. I guess it's why they consider image editing to be the next challenge.

Gemini 3 8 months ago

They ran the tests themselves only on semi-private evals. Basically the same caveat as when o3 supposedly beat ARC1

TTS still sucks 8 months ago

The Gemini models and Eleven V3, and whatever internal audio model Sora 2 uses are about neck and neck in converging performance. They have some unexplainable flavor to them though. Especially Sora.

Sora 2 10 months ago

It's the skin textures. It's the slightly better lipsyncing. Maybe it will be different when us normal users get it but so far the demos with Sam don't make him look waxy.

Sora 2 10 months ago

So far the true progress it has made is getting textures right close up. It still fudges how skin looks like the more it pans away from the characters.

Eleven v3 1 year ago

Training "high" points in voice inflection has been the priority, we've seen this in the 4o voice outputs and to some degree the Google NotebookLM podcast outputs. I would assume it's because they're trying to make it "act", but now it's a problem of swinging too hard on one end of the spectrum.

Lol and the Playstation was already in the public conscious as a product that a lot of people found easy to understand. With AI tools only being presented this way, I'm slowly becoming less surprised why the less informed public has a level of aversion about it.

Dude. Have you been paying attention to even the first Veo or even the first few iterations of Kling? They've HAD facial expressions that follow the prompt pretty well. You're being fooled by your own senses now because now you can't think they've existed before speech and sound effects have been integrated into the output. They've been there. You just couldn't hear what they were saying. You're paying attention now to how the words they are speaking make sense because lipsync actually adds relevant context to the output. But people have been making similar outputs just with a different workflow prior to this.

I don't need to create anything for you. Go visit r/aivideo and go look at the Kling or even the Hailuo Minimax (admittedly worse in fidelity) attempts. Some of them have been made to even sing or do podcasts. Again. They've been there for at least 6-10 months ago, this happens to generate it as one output. It's not nothing, but this really exposes a lot of the people who aren't familiar with this space when they keep overestimating things they've probably seen a months ago. Somewhat accurate expressions? Passable lipsyncing? All there. Even with the weaker models like Runway and Hailuo.

Again. Use the products. You'll know. Hobbyists have been on it for quite sometime already. Also. I didn't say they were just adding foley, though I can argue the quality of the sound they're adding, that's not my point. My point is, is that everytime something like this comes out there's always people ready to speak on "what industries such thing can destroy right now" before using the thing. It's borderline deranged.

I'm not underselling it. I'm reminding people who get swept up by headlines to actually use the products and be an objective judge of quality when it comes to these things. Because when you lose that objectivity, you start saying things like what you just said. Veo 3 level tech is basically Kling 2/Veo 2 fidelity with native sound generation, so was it that the last generation of these things were already decimating production houses? Be for real. With the tech they had 6 months ago, all they needed to do was add sound manually, which they could have pretty much also generated. A new layer of abstraction isn't "decimating" anything. I'd really take it easy from professing things like that. These things are great for what they are, but let's be actual objective consumers and not fall for these talking points of "oh industries are gonna change in x-months".

I mean I'm gonna say this with the hype settling down. But it's pretty on par with visually Kling 2 and Veo 2, it happens to output sound pretty ok but having it be one general output along with the visuals is the gamechanger. Beyond that, eh. I've kinda seen people try to take it to the limit and it's pretty much what you'd expect still from their last model

I think what I genuinely have an issue with is that extrapolation ignores the finer details of research work being taken to understand the current issues that may affect the future goals of a technology. It basically just ignores the science of it in favor of "well we've seen this level work before". The thing is, sometimes the more these experts learn about the current goal to beat in their product, the more things previously unknown and unplanned for present themselves as new issues. Extrapolating is a bit past-facing when you see it this way, because those new problems are entirely new challenges that haven't been considered in the equation and may require an entirely new approach. Every success is different, every failure, and every challenge is different. So extrapolating seems pointless.

Yeah, no. That's not my point. Saying something will get better in /any/ amount isn't unreasonable. That's being positive towards innovation and it's harmless. That's not exactly what the modern tech consumer is being made to expect nowadays. It's always the next iPhone moment. The next GPT. Always a huge leap. Like if we landed on the moon and we've been promised we'll have colonies outside the solar system by 2040

The thing about that recent rollout of self-driving trucks is they picked a stretch of road connecting Houston to another shipping point in a very straight line. And they bragged about 2000 "unassisted" miles on that stretch of road which is ~250 miles in length. So they're basically championing the idea that their trucks which have driven about less than 10 trips without an on-vehicle driver in that area is competent enough to be relied upon in a real work capacity.

Whether this amount of success is proof that there won't be any issues with the tech in that area, remains to be seen. Hell, they're not even interested yet in talking about how this may pan out outside their Houston trial runs.

I generally think that Kling or even Runway has achieved the visual fidelity of Veo (flaws and all, physics problems and direction of action and such), but now people are basically experiencing sensory bias where they think that some things about the visuals make better sense because nw it has sound as an added context. Visually, yeah. Probably on par with Kling, possibly worse on the depction of dynamic action

We're discussing the implications of this here when this has presented nothing novel in this medium/genre? I gave it a shot and it still has the same pitfalls for video genAI. Dealing with chains of dynamic actions is the biggest challenge, moreso with anime with its several fight scenes. No, it did not do good, and none of the non open-source models can do a good job of it for the most part either