When I say distill I also mean mine it architecturally for insights but I doubt seriously the model training is entirely distillation of Claude, it’s almost certainly a mixture of both original corpus and reinforcement as well as distillation. I think it’s a little condescending to imply that these new open models are cheap ripoffs with nothing original to them. These teams and labs are top tier as well, working under unreasonable constraints imposed by the USG. That’s a powerful combination for creativity.
HN user
fnordpiglet
Oink.
Because the open model might have been distilled from outputs of the closed model doesn’t mean they are architecturally equivalent. There is almost certainly innovations in architecture present in the open models that the closed labs didn’t think of. It also is almost certainly true that they aren’t completely built out of a distilled corpus, that reinforcement is equivalent, etc. Therefore closed labs will also benefit from being able to inspect in totality the architecture, activations, weights, and be able to train against it at scale in an ensemble of other models and their internal work.
The only way there is no benefit would be is if the open models are literal copies of the closed model, which unless there was direct theft, is highly improbable.
Interestingly the reality of the open source release is they are opening up to full distillation by the closed source model providers at a deeper and more fundamental level. If anything the open sourcing will help Anthropic and open ai ladder up faster. Open source has always been about mutual cooperation towards a goal and has never closed the door to commercial success. All the hand wringing about open weight models putting closed providers at a disadvantage doesn’t get what working in the open actually does for commercial interests - it is like science in the open - it enables and lifts all boats. Likewise commercial success doesn’t close the opportunity for competition or more open source work - it’s the economy of activity and competition that matters overall. When things stagnate is when closer concerns turtle up and collude on not competing for each others turf.
The future is good and better for everyone the more work is in the open and the more work is in the commercial space. It’s good all around.
Come now. The purposes for which the data is used is relevant. I am much more concerned with my internal corporate IP being actively used against me than passively used to train the model. I would also note you and sign agreements that prohibit the collection of data for use as well, which is also one of the key selling points of bedrock. In the west you can actually enforce such an agreement in court and win.
I wouldn’t lump this into a west vs east thing as well. This is particularly PRC. I feel comfortable doing business in Japan, Korea, Singapore, Thailand, Malaysia, etc. But it requires some particularly strong willfulness to pretend the PRC isn’t actively and structurally built around economic espionage, and funneling IP through PRC for short term economic gain has been one of the primary factors in their growth over the last 30 years. This just scales it faster.
I wouldn’t expect the USG won’t compel AI companies in the US to disclose and retain data as well - however it’s not a simple thing, the companies are hostile to it themselves, courts are often unsympathetic to the government, and the “machine” for converting it into actionable economic advantage is non existent - and there’s a very significant human component in that all links in the chain are culturally uncomfortable with such things. While it happens and it’s possible it’s very difficult, fraught, and does not scale. The PRC is the opposite - the courts, government, and business culture are all aligned in the goals and processes.
I think a better framing is the marginal utility of the models capability growth. At a certain point frontier models will only be needed for frontier problems. The demand for that capability will decrease with time. The hand wringing about not understanding is to my mind anthropomorphic - AI of today lack agency and awareness. Even the constructed stuff Anthropic puts out there in the model docs involve contrived scenarios to elicit “scary” behaviors. It’s unclear that as models become more sophisticated whether they’re better at instruction following or not but it certainly feels that way - even if it’s through better alignment or just an artifact of scaling. However I think the malign actors of humans using powerful models for bad stuff isn’t unreasonable to be concerned about.
The marginal utility problem is a real one for AI companies. I think the current generations are already saturating marginal utility for 95% of the population. Almost everyone I know outside of my career has no use for a more powerful model. This is a serious problem for the economics of AI and semiconductor investment. This is a bigger problem than Chinese models. It leads to a demand curve problem - that supply outstrips demand.
I’d also note that running a 2.8 trillion parameter model at scale efficiently is not simple. I would expect when open weights land getting it running fast, efficient, and at full capability will require sufficient resources it’ll be expensive outside of Chinese hosting. Which I think almost no western corporation would use for any internal work. You have to anticipate your use won’t just go towards training but will be actively mined for IP, trade secrets, MNPI, etc, or anything of use to the Chinese government or Chinese companies. I don’t say this to crap on the Chinese - but this is the playbook for the last 30 years.
That said I fully intend to use deepseek hosting for operational agents that are making decisions about non sensitive material. The economics are astounding.
Kimi? The economics aren’t that amazing to merit switching from 5.6. I expect fable will rapidly reappear in subscriptions. Competition is good.
Lost me here - I was there - I worked on open sourcing communicator at Netscape and starting Mozilla. We were the one company that tried to own it all, but were out spent and out maneuvered by a more insidious company. This totally misunderstands the situation and the history. Plus the “our CTO” and “CTO letter” at the start was too pretentious. For gods sake; CTOs are just engineers that aren’t worth trusting because they’ve either maneuvered their way to their seat or they are serving the investors on their knees. Any CTO worth a damn knows they aren’t worth a damn.
“”” We have been here before. Mozilla exists because one company tried to own the front door to the web, and an open community rose up to make sure it never could. Twenty-five years later, someone is running the same play. We bet on open the first time. Open won. Together, we can do it again. “””
And this demonstrates this benchmark does not necessitate achieving those goals to achieve a perfect score. You seem to miss the point that almost all of math, physics, computer science is built on constructing an objective then cheating to attain it, and that demonstrates that there is an equivalency. Maybe the benchmark is flawed, or maybe the goals are not strictly necessary to attain.
For instance, is solving a math proof by enumerating all permutations exhaustively on a computer cheating? Does it matter that it is not a proof by construction? That its not descriptive? Of course not. The proof of the four color theorem is all that’s necessary and sufficient to prove it. Calculator at an algebra exam? Who cares. This isn’t an exam, this is the real world. The fact an AI can use a physics harness to perfectly achieve ARC-AGI-3 without attaining those goals demonstrates the power of the technique and that the goals are not necessary for that class of problems. Then find another benchmark that actually demands the goals be necessary and sufficient to achieve the benchmark goals. But don’t denigrate the fact we have technology today that yesterday was a fantasy.
A great deal of mathematics is transforming nonlinear problems into linear ones and solving them with linear techniques. Others are solving non linear problems through stochastic methods. In almost all cases most non trivial math is done by transforming a harder problem into a simpler one.
I get what you mean in terms of testing the model itself to see its improvement in some domain. However if you can transform the domain to be better adapted to the model and achieve the desired results, this is indeed an accomplishment because a whole domain of problems is shown to be practically feasible with this technique without expensive model improvements. Of course the benchmark still exists without the harness, but the harness also exists which allows these problems to be solved.
As noted elsewhere the models themselves were used to build the harness, which means the models can in fact score this scores without intervention but building a harness for themselves adapted to the domain and using it. Is this cheating by the goal posts you’re setting?
There’s a real tension between “I want to solve problems and this technique shows how to solve the problem domain,” and the “I want to measure how something performs unassisted with other techniques.” Fortunately it’s not a mutually exclusive situation. You can do both simultaneously, gain the benefit of the technique to transform the problem into something tractable and keep measuring using the benchmark.
2008 in the EU and last year in the US.
Not necessarily - the compliance side would open up stripe entirely unless they very carefully structured as a bank holding company. Stripe has been trying to gain a charter in Georgia for years and just got it last year. I think this taught them it would be easier to acquire one. Additionally PayPal has been a EU bank for 12 years and just recently got its US banking charter. A global bank is pretty compelling.
What I find interesting about this isn’t acquiring the flow from PayPal but the fact that PayPal has a bank charter and stripe doesn’t (not a real one, all that Georgia bank stuff is shenanigans). Properly structured this gives stripe a lot of interesting opportunities that they’ve had to delegate to partners and that dilutes not just their margins in the fee stack but restricts the types of transactions and merchants they can host, beyond the minimal regulatory constraint.
The next move would be to offer a stripe managed card network that captures the full MDR interchange fee stack, from issuance to processing to network to bank. They could offer deep discounts to merchants hosting their accounts on their bank issuing cards through them and processing through them by cutting out every middle man. It would instantly become the third largest network in terms of points of sale behind visa and Mastercard. They could further carve out heavy rewards for customers through direct relationships with airlines, hotels, and merchants through their existing issuer services. They would become the first fully integrated vertical in payments, and the efficiency that would give them in fees would allow them to crush the hodge podge payment stacks out there that exist. I’d wager in five years they would dominate the market.
Agreed. It’s also a lot better at following instructions. If anything codex can get caught into an overly literal adherence to instruction while Claude you can barely trust it to sit still for 5 minutes. Tell Claude to use an MCP for a task, 30% chance it’ll do it. Provide it skills, 10% chance it’ll use them when appropriate. Codex is almost the mirror of that. It’ll almost always use the MCP and recall the skills.
The challenge I think for codex is the restriction on context size and the constantly rolling compactions. They are less aggressive or disruptive but it is still annoying you can’t force a 1mm context window on a 1mm context window model.
But it’s recall beyond compaction boundaries is much better than Claude code. The impact of compaction is much less noticeable.
Performance on task work is between fable and opus, but the marginal utility between that gap is not enough to pay extra.
Yeah that was my instinct too. What sort of career defining trends are visible with this much historical data? Feels like someone wrote clickbait research to get published.
Hmm ok. The fact 5.6 Sol performs around Fable level and is included without mega token spend in the subscriptions means I’ve promoted codex to my primary harness and model. The latest release of the CLI, app, and desktop fills a lot of the gaps.
Anthropic painted itself into a corner with fable at many turns and this latest twist is one of the more interesting. Either fable is too expensive to run at scale, or they’re trying to incentivize mega spend on tokens, or whatever - but them locking the frontier model away for the few enterprises willing to spend top dollar while codex is including frontier in the subscription (and I’ve found it also is both less token hungry and the limits are much higher for codex) has finally made me put Claude aside and use it as my backup for very specific tasks, where codex has filled that spot for a long time now.
50% more weekly limit, but no fable. Ok. I might have a refactoring job somewhere for you Claude for those extra tokens.
But a proof isn’t an explanation it’s a proof. Proof by assuming the opposite is true and demonstrating a contradiction is very indirect and not at all directly explanatory yet it’s a proof non the less. The goal of proofs is to demonstrate something to be provably true, not expository knowledge gathering.
In fact most mathematicians (myself included!) think the more clever the trick the better the proof! The trick itself being clever is interesting because it often yields a new way of tackling or thinking about your own proofs. A bland explanatory proof that elicits some conceptually “why” is only preferable if it has a reason for doing so - does understanding why yield a new avenue of research? Often then the “why” is quite a clever trick too.
I think it’s a bit the opposite of programming. There you want your solutions to demonstrably not be clever and the code be its own documentation. It’s a different discipline.
It’s a real condition. For me it’s jet liners of various makes. I had to rewrite the quote as “0.005 Boeing 777’s” to be able to comprehend just how strong those snails teeth are.
Given the game is stable and the changes would be at the integration points, and Fable was able to do the direct integration, why would the answer not be “it’ll maintain itself” at some abstract level. The decision to maintain open source is up to the maintainers and I think the answer is “no one” 99.99% of the time, but I’ll wager if someone is willing to spend the tokens on it, a CI reintegration agent would do just fine in keeping it working as the underlying dependencies have required changes (which would really be only major changes in apple apis that aren’t backwards compatible.”
Pylint is different because it’s working against a necessarily dynamic wavefront that it has to keep parity with as it advances. All python changes, ecosystem adaptations, etc - and maintaining that with an AI harness in CI would never work. It would require a concerted effort and thought along the way.
So it’s sort of a different beast all together. In fact I think this is a great demonstration of using AI to resurrect technology built for X to work with Y, where X is dead and Y is current. Automating this feels like a net positive and because the original software is “finished” there isn’t decision making and strategy required.
I prefer hytelnet and MUDs but I don’t count, I’m just too old.
Reagan didn’t do it alone, and the philosophy that created this monstrosity inexplicably keeps marching along as if all it needs is more patches and it’s the fault of the demand side and supply side for their philosophical view that “free markets are magic” under every scenario, trundling through the wreckage holding hope as a strategy. As the only country with this problem, maybe the philosophy is wrong.
I blamed Reagan because he managed to effect the philosophy in law and regulation, and therefore create the inertia. I can blame him for his role until the end of time, just as I can blame Julius Caesar for ending the Roman republic long after his death (even that happened shortly after his actions!)
But it’s ok - the current slow motion disaster unfolding will make the hellscape of reaganomics look like a brilliant insight as we play out early stage idiocracy.
This is wrong - the training data is necessary but insufficient. There are a lot of other parts of the architectures used that add a lot of value - otherwise Markov chains would be all you need. There are layers upon layers with non linear activation functions, learned residuals, etc. They still absolutely must interpolate but the space they interpolate through is much more complex than the training data, and they can definitely create things not in their training data. What they can not do is wander outside their non linear parameter space’s convex hull. But this is a really permissive constraint on what they can do “creatively.” People generally under estimate the advantage the architectures confer on that constraint. This is why there was a step function change in expressive power as the architectures (attention, self attention, transformers, diffusions, others) evolved given the same training data. Generally though I challenge you to define “creative” in a way that is precise enough to measure and isn’t self referential or refer to concepts ill defined.
The key tho is can they solve problems not easily solved before with prior techniques. Further can they identify problems not readily presented. Then identify novel solutions. Etc. The answer is emphatically yes they can. These features don’t have to literally exist in their training data, but the supporting highly convoluted network of associations of all their training data does have to in some complex space allow for it to produce these answers. It’s not the same as they’re stochastic parrots at all.
Are they creative? No, because they don’t have awareness. My personal imprecise definition of creative requires both self and awareness as well as free will. There is no driving awareness in all AI architectures, it all derives from extrinsic impetus. Creativity is derived, IMO, from a layer of our minds that is not readily assessed or measured and is only indirectly expressed through language, art, and music. Hence it is not directly trainable and therefore a learning model can’t learn it by reinforcement. It can learn the proxies, but the proxies are not, as we all deeply know, the same as our experienced awareness. We are not our words, our art, our music. We try hard to bridge it, but it’s impossible and you and I know this to be true from experience. In fact we can not even examine our own awareness because it’s not directly observable or possible for us to directly reason about. This is core to a lot of philosophy, especially mid and far eastern philosophy of the mind, the self, the five aggregates of Buddhism, etc. Psychology points at it, and modern psychology avoids it because it’s practically difficult for outcome oriented treatments.
The moral hazard is making a product with nearly totally inelastic demand a multi layered adversarial free market with structural price opacity. Thanks Reagan!
It’s actually likely fair use on the way in as well. What’s not fair use is the production of copyright material with the model and the question is the extent to which model providers have to prevent it. These topics came up with the photocopier, VHS tapes, etc. The training side is more subtle because they are clearly unlicensed and used in the model but this is actually similar to taking a book and photocopying sections and using them for handouts and in training materials or other uses. The crucial part is they effectively destroy the original material in training and no where in the model is the copyright material, even if they can produce something similar when deliberately induced to do so. However you the user induced it, and depending on what you do with what you induced, you can violate the original copyright holder. (N.b., IANAL, but these are my summaries of discussing with a law professor at length who specializes in copyright, open source, etc)
Whether it’s moral or not to not remunerate everyone who produced the training material is of course important but a different question. I sort of agree with Sanders et al that Ai should be a public trust like the Alaskan oil reserves. But good luck.
Mission accomplished!
ai generated imagery can’t be copyrighted while all other photography can and generally needs to be treated as it is. Therefore you likely have to pay a royalty to Getty or other asset outlet. Of use AI.
There’s even a skills over MCP WG; and I typically deliver skills via tools in my MCPs intentionally. I find Claude and codex recall of skills via MCP tools to actually be higher than skills themselves, which I feel have an (unmeasured) less than 30% recall rate. I have to typically force skills to be loaded explicitly through / or $ depending on the flavor and skill graphs are very unreliable.
The catch is they’re still the administration.
My struggle is despite giving it a good honest go I couldn’t find the use case for OpenClaw. Maybe I and my life is just too simple, and I’ve admittedly not invested a huge effort into spending countless hours watching breathless YouTubers vying for my attention to figure it out. I have built a few bits and bobs, I feel like I’m ideal in that I’m pretty deep into agentic workflows for work, my house is totally home assistant, I’ve got current tool following models pinned to local 4090’s, I’ve been fine tuning my smaller models, i even have a custom chatterbox based tts/sst pipeline with voice nodes everywhere.
I’d love some OpenClaw master to opine on meaningful use cases beyond clawbook, checking the weather, and telling you about crypto news, as it genuinely feels like something I should find utility for but am just too old or something.
I think for many practical purposes the frontier open weight models are almost universally good enough for most things. There may be greater and greater frontiers but at q certain point it becomes like IQ. Having a 150 IQ doesn’t mean you’ll be more successful at any particular task over someone with a 125 IQ. Indeed there’s a diminishing return on intelligence on many utility functions where being more intelligent yields more be same or worse ultimate outcomes. It might very well be the person with a 150 IQ could understand some extraordinarily complex and esoteric concepts faster, but it doesn’t mean with more effort the 125 IQ person can’t either; and sometimes that extra time spent yields better outcomes overall.
I suspect AI will be somewhat similar where even if the linear scaling laws continue to hold the practical utility of a model flattens for almost all conceivable use cases.
In some ways I already feel this has begun to happen. The marginal utility of opus class models and fable has in my perception begun to flatten. While I can tell the differences they aren’t earth shattering. I could continue to use the present models for the rest of my life and be ludicrously more productive simply by adapting within their constraints through ever more sophisticated applications.
What holds back the open weights IMO is hardware scaling and industrial production. As the enormous transfer of wealth in debt and equity markets unfolds with semiconductor and adjacent companies and the corresponding capital investments are made, and the eventual bubble pop leading to over capacity and market flooding, as well as advances in technology, math, techniques, and efficiencies, will make very large open weight models more directly attainable. This will also lead to chimera models that MOE very large models to get very close to the 1-2T parameter dense models, at which point I suspect utility for almost all uses is nearly fully saturated.
There will be areas where more capable models are needed but they will be frontier models on frontier problems. This, IMO, is inevitable, and without some criminalization of weights (see the attempts to criminalize encryption algorithms in the 20th century and all the wonderful tshirts that emerged). It’ll be harder to print a trillion parameter model on a shirt but I’m sure someone will try, as will governments try to keep us in our boxes slaving for food coupons and basic rights like health care.
The problem is there’s a real wall on the vram side. While fused main memory is ok the inference speeds on larger models are impractical. With vram on a GPU the machine class, power requirement, GPU costs, and other factors put them out of most people’s reach. Cloud GPUs require a second job to keep available and hot. What closed providers offer is packing and scale advantages as well as infrastructure. The scaling laws here aren’t the same as Moore’s law - in fact they predict more required hardware and more scale over time. Moore’s laws isn’t keeping up with expanded needs and the ability to fab and produce at scale the specific things that weren’t needed a few years ago are lagging. So it’s not a 6-8 month lag; it’s a lag that will be induced by hardware scarcity and an ever increasing lag until something fundamentally changes with matmul.