Agreed. The benchmark closest to my experience is FrontierMath Tier 4. Fable and Sol (90%) are very far ahead of Kimi K3 (not even 40%). Kimi is trained heavily to basic agentic tasks, like all the other open models right now.
HN user
hodgehog11
They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.
It really does depend on your application. In my domain (math research), it is substantially better. Fable can solve really hard tasks with surprising consistency. It makes mistakes, and occasionally refuses, but honestly, at the top level, ideas are the currency and the rigor is the busywork. The other models cannot come close in this domain.
If you couple Fable's idea factory with Sol's rigor, you get a real game-changer. It puts the emphasis on top-level ideas, and nearly trivialises the intermediate layers.
Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.
I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here to stay in one form or another.
The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.
That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, because it is all too dangerous in the hands of anyone else.
Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companies are still hanging in there.
I love convex optimization and there are a few SciML projects I am on where I really need results from there. But in AI research with deep neural networks, it's become a liability, because people will just not let go. I'm getting tired of reviewing convex optimization theory papers in ML conferences that are still trying to wave away the obvious issues with their application to deep learning. It's harsh, but I do feel we can only start talking about an intellectual debt once that stops being the case.
It's not a matter of whether the theory "works"; it's a matter of whether one is asking the right questions. Convex optimization studies how quickly an optimizer can reach the optimum. In the non-convex case, there are many basins containing their own local minima. The more sensible questions there are "which basin is it likely to go into?" and "how do I steer it to go where I want?". Global convergence rates are largely irrelevant by comparison.
No, I have to push back as well, sorry. It takes a very long time to get to the "near-minimizer" stage when training a neural network, and in practice, you never get there (see neural scaling law regimes). What you are saying is the viewpoint from 6-7 years ago. Things have changed.
The reasons why optimizers work well for neural networks in their highly nonconvex landscapes has absolutely nothing to do with their performance in convex landscapes. If that were true, everyone would be using Newton-CG. These optimizers were born in the convex optimization literature as a consequence of the genetic optimization nature of incremental publication (and because that was all we had), but their modern study is through the lens of implicit regularization (their preferences for certain minima) and their stepwise vs. continuous rates for feature learning in multilayer models.
This is completely new theory by the way, and requires painful reinvention of the field. It does not stand on the shoulders of convex optimization. The nonconvex setting is assuredly not a perturbation of the convex setting, and those that do continue to work on deep learning optimization from the convex optimization perspective are well behind the times.
Very confused by this comment. The older (poorer) parts of the ML literature focus on models with convex and (gradient-)Lipschitz objectives, but that's not representative of reality, not even close. Modern objectives for AI models are famously nonconvex (catastrophically, from the point of view of classical optimisation theory), and that's where the interesting research is.
I really envy you. There is a clear divide amongst my colleagues now in terms of who is using Fable and who isn't (this is math work, so it is well and truly better than all the other models). Everyone is quickly becoming reliant on this, and with Anthropic constantly changing their mind, it's just miserable.
For the particular task I'm working on (a mathematical task in validated numerics), even Sol has generally just repeatedly given up. I asked it for the main problems it could not solve, gave them to Fable, and Fable solves them, every time.
There is a clear hierarchy here and Fable is well and truly at the top for very tough tasks. It just sucks that Anthropic is so incompetent at delivering a cost-effective experience without the uncertainty of it getting pulled. It doesn't matter how strong their model is, customers will turn away.
Overengineering is the name of the game with Fable. Sometimes you don't want that, sometimes you really do, especially as a researcher. It's a very nice tool to have around for those special tasks.
For some tasks, there is no amount of "steering" that will produce sensible code. The model needs to be sufficiently capable as a baseline; this is the "intent" that people are referring to with Fable.
I'm sorry to hear you are unable to use Fable; my partner is in the same boat and it frustrates her immensely to see what I've been able to do with it. As someone who is working with developing new linear algebra routines, Fable is so far ahead of GPT-5.5 and Opus that it's obscene. Massively better insights and far better at handling delicate corner cases without needing to mention them. I would be stunned if GPT-5.6 is at that level, but one can hope.
I would be absolutely stunned if this were really the case in general given how irresponsibly large Fable is, and 5.6 Sol most definitely is not. It depends on what your problems are though, I suppose, since there are those that swear Fable is at best a minor upgrade over Opus, which has not been my experience.
What type of work is this for in your experience?
I don't agree that my argument was a "No true Scotsman", since the argument made above (as far as I read it) was that "all PhDs are a waste of time". My counterargument was that the PhD, if done right according to its purpose, is not a waste of time. That is not the same as asserting that all PhDs are going to be good, and then refining the definition.
i grew up in math where novel work meant a truly original theorem
I am in the math department and work with geometric analysts and combinatorialists, but mostly with CS researchers. I work with math and CS students. My entire education was analysis, then pure probability theory. Just for some context. And while CS has a lot of problems, math does too, unfortunately. A theorem may be original, but that doesn't make the work good.
bro you are so out of touch it's laughable. there are departments full of your so-called "not true advisors".
Your experience, not mine. Regarding "out of touch": I've supervised across five different universities, all (apparently) in the top 100 worldwide. There has been a quality supervisor in each department I have come across. They may be in the minority, but they exist. They tend to not have many students, because it is virtually impossible to give adequate supervision to more than ~5 students at any one time.
you recommend to them they what? transfer schools?
I now live in a country where that is easy to do, so yes I do, and yes I have. I have had friends that were "forced" to grind in a bad supervisory relationship for a while until they found another better supervisor and then had a good experience. I understand it is more difficult in a country where you cannot easily transfer. In those cases, I would 100% suggest dropping the degree early before wasting time and money on it, but naturally it depends a lot on their situation. A bad PhD can be a lot worse than no PhD.
that kind of quality is 1/1000 in neurips.
I've found the proportion is more like 1/100 in the theory domain, but yes, there is a lot of rubbish in NeurIPS. There is a lot of rubbish in every venue.
attend a US T10
I did a lengthy postdoc at one of these. They are largely paper mills by construction. A top 500 most cited researcher got that way because they run a paper mill. I did not associate with the students there very much, and I doubt that they thrived there either, other than getting a nice item on their CV. On the other hand, that institution benefited me at the time. I'm glad I did my education somewhere else.
there is literally no program. you're completely full of shit. there is no "standardized phd". every single department around the world completely makes it up to be whatever they want.
Yes, which is why your experience is not necessarily indicative of that of others. That's a good thing.
the illusion that academia was a priesthood in pursuit of truth/knowledge/beauty/whatever
I won't disagree that academia is a cesspit and I would exit if I felt I had the opportunity to do good elsewhere. I actively do not recommend it to anybody that feels they could do anything else. But there certainly are people that benefit from the PhD, so calling it useless is a bit much.
have you looked at the mirror recently? the frustration comes not from the craven/mercenary individuals who admit their cravenness - it comes from gaslighters like you who claim there's some idealized version of it that exists that everyone supposedly falls short of (hint hint: have you ever heard of this convenient concept of original sin?).
I love the insinuation that I must be an evil prick. This is tantamount to hearing from a victim that all X are evil. I respect your position and your experiences, but also recognize you do not speak for the entire world in this matter. The bad supervisors tend to excel in their metrics nowadays (including how many students they have), so their perceived presence is disproportional.
in addition i got a job in FAANG so that was a nice consolation prize
Good for you! If that was always an option, I would question the PhD too. Telling others your own experiences to warn those away from academia is perfectly fine, unless having that degree did help get you to that point. If you honestly know that having the PhD on your CV did nothing at all for you, and did not positively influence your chances, then this is fair, and I would recommend teaching others how they can avoid doing the PhD to get into those roles. Yes, people get FAANG jobs without doing higher education, but was everyone going to take that route? Otherwise, you are in the ivory tower telling others that they shouldn't come in.
Yes, I did get it within the last ten years (with an excellent supervisor fortunately), and I take part in student supervision. I get to talk with the students and determine what is most valuable for them in the long-term. Considering where those students have ended up, and that we still keep in contact, I'd say I've done alright so far.
today it's about pumping out questionable papers that your advisor tells you to pump out and targeting the right conference.
That's called a garbage supervisor. I'm sorry that was your experience, nobody deserves that. There has been only one student I have personally experienced where I felt some of the papers were questionable, but they were not strong, and should not have done the degree in the first place (unfortunately, I have no choice in the matter and just need to get them through).
it's not supposed to be suffering. it's supposed to be challenging
Immense challenge is suffering because you encounter failure over and over and over again. Eventually success will come, but the interim is tough on the mental state. No human likes to fail repeatedly, but it is important to experience that at least once in a research career. It only needs to happen for one project; afterwards, the student tends to develop an incredible autonomy because they've been through the worst of it. But of course, this is where the supervisor is supposed to be responsible to find the right balance, since the difficulty of the task needs to be judged to ensure the student can finish it, while being as ambitious as possible within reason. It's pretty difficult to get right, and many supervisors don't even bother, choosing instead to push their own career. To any PhD student out there experiencing that, find another PhD supervisor please, before it is too late.
there are no problems like this in academia (at least not CS) because it's absolutely impractical (re graduation, tenure, etc.) to set out tackling problems which are intractable.
No, this is precisely what the PhD is supposed to be for. It should be the most challenging topic of your entire early career, or your supervisor wasted your time. You have a handful of years to impress people, so they have to count.
this entire comment is "phd virtue signaling"
Honestly, it just sounds like you were unfortunate enough to have been in a supervisory relationship where you were used as part of a paper mill. I'm really sorry to hear that; it isn't uncommon, but it isn't what the degree is supposed to be, and it seems you did not experience what you should have. The program doesn't work for everyone, but it does have a purpose, and it is frustrating when selfish academics bastardize that purpose to give this kind of false impression of the degree.
I won't try and argue the merits of Bachelor's and Masters. But if you honestly believe that you can pick up the same experience from PhD on the job, then is seems like you learned little from that PhD that you were supposed to and your supervisor failed you. PhD is not about learning content. It's about picking up research independence, confidence, and a strong capacity for critique. It is supposed to be suffering of a special kind, an experience in grinding and radical unproductivity when the problem really is that hard. Maybe you did get that but you aren't using it in your job; if so, then maybe you didn't need that PhD, but that's hardly the fault of the degree.
Basically we are saying that if a hobbyist really wants to operate and maintain something, they should be allowed to after some amount of time if the studio stops.
I don't believe that is the problem as defined by SKG. I think the issue is less about empowering users to make use of the property afterwards (although that is nice), and more about the legal ambiguity in the process where products are removed from users without effective prior notice (from purchase date). Games without subscription fees are traded as effective goods by commerce law, but publishers operate as though they are services without a defined end-date through a EULA. IANAL, but my understanding is that whole concept is not legally tested. A one-time purchase should include a contract where the terms for revocation (ideally none, but we can't have nice things) are agreed upon and cannot change at the discretion of one party. You cannot have a license that essentially says "we can do whatever we want at any time". That has never been okay in the history of commerce. A subscription is different, since you know exactly how temporary it is at any particular time. Of course, both a defined lifetime or a subscription are tactics that have been tried and were not as popular with users, so companies resorted to effective trickery while users looked the other way for a time. The social contract is changing as users are watching games they loved die. So overall, the better solution for every party now is a minimal EOL. This isn't a precise threshold and it isn't an expectation that the game would continue to function as normal. It's a minimal effort taken to ensure that the customer still has something left to play around with which is in the spirit of the intended customer agreement; in the California bill, either patching for reasonable offline play (the command-line switch), providing server binaries (I can see this as more of a problem more often), or refunding (no-one wants this).
if a rights holder decides not to continue to license a film or TV show or book for publication, we let them.
We let them, provided they don't rip the product we buy out of our hands. That's not new for games, but is only now happening for movies. It's unacceptable in every domain. It's like a user being told they can rent a movie, but it isn't clear how long for, so the movie can be requested back after only 5 minutes or could be after 5 years. It is fundamentally unfair. That comparison is quite real too, since some have bought a game, only to find it to be shutting down soon. That feels like theft. Sony removing movies from users libraries and from their computer feels like theft. I would argue it is theft, but that is part of the legal battle here.
I want studios to take risks, honestly the bigger the better, because that's how really interesting ideas get made
I agree that creative risks should be taken, and also that bad economic strategies should not be rewarded. That's the natural course, and it seems that is happening right now just as one would expect. Many AAA studios are not taking many creative risks according to users (it's almost in the definition of AAA at this point), but are making poor economic decisions by not handling volatility with diversification. It's a recipe for disaster and the developers suffer most it seems.
having a reserve and required auction at shutdown might be a good idea
I think that's fine, but I'm not sure whether it addresses the legal issues if the new rights holders do not uphold the original understanding of the purchase.
Thanks for continuing the discussion! It's not often an outsider would get to speak with someone with your experience.
I definitely agree that it is not tenable in any studio to have a full time staff member dedicated to packaging software to be in line with this legislation. The solution has to be similar to a toggle in the engine code itself at the very earliest design phases. That means there has to be sufficient advanced notice and can only apply to future titles. There is hope that an inexpensive industry of third parties may arise to easily handle this aspect upfront with new software, but it would be great if there was a tangible demo of this.
Good point about publishers providing the cash to get the studios through the bad times. I think where this falls apart now is in the modern big-budget live service model itself, since these projects are expected to be so long term, so expensive, and gain so much income, that their lack of success now seems to end up in a studio turning the lights out altogether. Concord comes to mind here. Bungie's reliance on the new Marathon is not something I would wish upon any studio either. These kinds of game development strategies do not seem to be healthy for the industry anymore, and some diversification is really needed. These comments here are not really an SKG thing, it just feels like a far bigger issue from the dev side right now. I hate that devs are constantly losing their jobs in the current market, and I think gamers do too. I just can't see the status quo as sustainable. I guess that's why I treated the "back breaking" as binary, but it is true that there is a whole range of suffering inbetween.
A studio should never be forced to keep something going when it is deemed a financial failure. Again, it has to be an upfront design decision during the early planning years so that the end of life version is mostly a compilation target. Latest version goes out, no more updates, no more servers to run. There were more specific solutions discussed by other devs in a video on Ross Scott's channel, but that's the general idea. Does this seem even remotely possible from your viewpoint? Maybe not on current projects, but for projects five to ten years in the future?
If it is too hard to implement in this way for the team, the legislation may force to accept that the nature of the project may be so inherently risky with the current staff resources that it should not come to fruition until new software developments have made it less risky. But you are right that it might just end up as one less person with an actual dev job at the studio, which isn't great in the current economy. Either way, I think many users feel their hands have been forced by the publishers as I understand.
Absolutely! Happy to learn from an industry veteran; hard to argue with that pedigree (part of why I like HN). I figured you might have been, but wanted to push back a little because I think it is important.
Here is my impression. 2001 feels like it was still within a golden era of game development; "AAA" games of that time would have been made by smaller studios still, and budgets could be very large, but not catastrophically so. The industry was still expanding. Post-GFC, once graphics scaled, demands seemed to scale, and costs blew out. Games had to reduce risk as a consequence, become more consolidated, more live-service. But the model was never sustainable at that scale. The tech improved so that costs for basic games went down, but big-budget AAA live service costs went to the moon. Volatility skyrocketed, leading to rapid hiring-firing phases. Now it is at its most extreme and the AAA side of the industry is in crisis. Demands for long-term support could be the straw that breaks the camel's back, but it always seems like that back was going to break eventually anyway. What I hope is happening is that talent is falling into the hands of smaller publishers, but that might not be true. What I do think is that the nature of development may need to change so that studios are able to facilitate these requirements while remaining profitable. Some have shown it can be done, anyway.
That's my impression from a semi-outsider perspective. Happy to be corrected though.
Small like 30 people? I don't think so. Genuinely small studios are able to preserve their games for the long-term, so there is no excuse.
AAA gaming isn't something that needs to be protected. Anything of sufficient magnitude and budget should be made responsibly or not at all. If a game is unable to exist without screwing over the user in a very legal sense, it has no right to exist. The requests of SKG are not only sensible, they impose a serious legal ambiguity in the current system that needs to be corrected one way or the other.
AAA was always an untenable monster that brought obscene risk; this was clear even in the mid-2010s and the industry is rapidly moving away from that model for good reason. The user-base clearly does not look kindly on AAA anymore either. That's why they are not profitable anymore. The SKG requirements would make very little impact overall compared to this.
Judging from the decisions and outputs of the last decade or so, the leadership at Meta, including Mark Zuckerberg, have got to be among the most incompetent I have ever seen. They go all in on the worst decisions; not just the worst in hindsight, but also the worst at the time. The only thing keeping them afloat is their monopoly from past purchases. They are a posterchild for why the US is no longer a properly capitalist nation.
It just requires game engines to develop a tool to make this much easier to do. The rule won't apply retroactively, so it just means different design choices from the start.
For a small game studio it may be incredibly prohibitive.
Alright, you lost me here. All of the games that do not follow what this law would suggest are AAA (or scams, essentially). The smaller studios always have some acceptable end-of-life plan from my experience.
I don't think that's what working state means here. It's not a full snapshot or even necessarily multiplayer support. It means minimum functional, which to me is just barely enough to see the assets used in the game. Not necessarily networking, matchmaking, online services, or the relevant tooling. Developers have already proposed making helpers that would introduce an effective switch to do this.
Also what you said about releasing the code and letting the community figure it out was explicitly an example SKG said was okay, from memory.
I think we know from using Athens of Ancient Greece as an example that true democracies of this kind are not a good idea at scale. Enough of the general public can be so easily swayed on a clearly catastrophic idea.
Not sure if you know this, but this is literally the Stop Killing Games movement. Despite the recent apparent setback as reported in the news, the organisers are in talks with EU MPs that are writing sweeping legislation to address this sort of thing across all digital mediums.
You only need legislation like this to hold in one major market to make a big difference.