Immediate reaction is that it seems to be a bit behind Meta Muse Spark 1.1 performance at approximately the Deepseek v4 Flash price point. That's quite good given Muse Spark benchmarks a lot better than Deepseek v4 Flash (assuming benchmarks mean anything, which they don't).
HN user
drob518
Typically, the inner loop of an interpreter is all about keeping the branch predictor happy and trying to fit in L1 cache as much as possible. So, falling through makes sense. I’m also curious whether a simple check of a flag and a conditional branch (which is highly predictable for a mode flag like whether you’re tracing or not) wouldn’t be faster than one of the indirect branches. That would keep the tracing code out of L1 when not in use (you can locate it in a remote function).
That’s the fear.
But they make up for it by shipping it late.
Yep, agreed. They still are not releasing anything frontier-class (Gemini Pro) at this point. Feels to me that they keep getting scooped by others (e.g. Kimi 3) and then are retrenching.
I agree that there’s huge value in being the generic category name (see Kleenex). That said, Google and Apple are going to roll this into every phone and tablet and laptop/Chromebook. Yea, OpenAI can release a phone app (already have), but I’ll bet you dollars to donuts that the Google and Apple integrations will be superior, and even if the EU or somebody forces them to create an “AI provider neutral API” like they have for web search in browsers, most people will just roll with the default like they do for Google Search. You might be allowed to choose OpenAI, but most normies won’t.
My personal feeling is to not move to ASICs just yet. Things are still pretty frothy right now, so I would probably wait 6-12 months. At some point the froth always calms down. At that point, commit to ASICs. That said, I’m also not totally sure exactly how much the ASIC hard codes vs having some wiggle room. The Taalas site is a bit vague as to exactly how they encode the model.
It’s owned by whoever directed the AI to write it.
Right, but the AI isn’t the one who would actually claim copyright here. It would be the human using AI to accelerate the coding. And the human does have standing.
Exactly. A fast, cheap model, particularly with the right harness and loops, might take us a lot farther than we might guess.
Of course not. But many tasks won’t require Fable 7 level intelligence and many people won’t want to pay for it. Honestly, I’m using Deepseek v4 Flash a LOT lately to do more mundane tasks because it’s so nearly free and I don’t need Fable or even Opus. Serving those mid-level models at high speeds and low prices is a definite winner for lots of applications. And sure, the frontier models will continue to drive the frontier forward.
Sure, but you’re undoubtedly the tip of the spear. Lots of people don’t need that.
The churn is an issue. For a while there, it felt like we were getting a new number format every month or two (e.g., fp4, ternary, etc.). That level of innovation works against moving things into hardware, or at least you need to be willing to spin hardware constantly.
I see a future where lots of models can plan a vacation for me, not just ChatGPT. Saying that there is brand equity in the ChatGPT brand feels a lot like saying there’s a brand equity in the MySpace and AOL 25 years ago. Google in particular, via Android and its relationship with Apple to power Siri, has a much better shot at grabbing the “plan a vacation for me” consumer market, IMO.
I could easily see OpenAI become irrelevant in 2 years if they stumble at all and don’t keep up with the other frontier models.
Right. Arguing with the original article, not you, my counterpoint would be, “Sure, if they’re able to do that. But so could any of the other model-only competitors.” The best you can say today is that they have announced an intention to do those things, but in no way have they established themselves as being successful, yet. And the model-only Chinese providers have access to lots of e.g. wearable and consumer tech.
More importantly, as sustainable long-term businesses, model-only providers are particularly at risk. Knowledge Atlas, Moonshot Labs, and Anthropic face defensibility challenges versus OpenAI, Alibaba, SpaceX, Meta, and Google.
Hm. How is OpenAI not a “model-only” provider just like Anthropic? Seems like they are vulnerable in the same way.
Yea, Flash is quite fast, though looking at model data on Open Router some of the other models are quite fast (Muse Spark, Grok, etc). I’m sure all these models have been trained on my conventionally published books as well, but I don’t care.
Exactly. If I could upvote you twice, I would.
If you read the media from 1998 - 1999, there was than overarching message that EVERYTHING was changing. EVERY business needed to be on the Internet right now. Any business that wasn’t investing in the internet was going to be slaughtered. Brick and mortar stores were viewed as a negative on your balance sheet. All commerce would be e-commerce in just a couple of years. Startups with no revenue and no prospects for profitability for at least a decade were IPOing for hundreds of millions of dollars, sometimes with 2x-3x the valuation of brick and mortar peers in the same segment who had established brands going back decades. Whenever anyone questioned the lunacy, everyone responded that, “This time it’s different because the Internet changes everything. None of the old rules apply.”
And then the crash happened and we learned that all the old rules still did apply, and while the Internet was hugely transformative, the simplistic idea that brick and mortar was bad and e-commerce was the only way forward was fundamentally wrong. Yea, surely some segments did change (Amazon killing Borders and Waldenbooks, Netflix crushing Blockbuster, etc.), but investors ultimately look at profit even if they’re willing to ignore it for a while to grab market share.
So, I suspect that AI is going to play out similarly. Right now everyone is frothy, screaming that anyone not token maxing will die. Yes, AI is going to be impactful, and, yes, it’s not going away. But the hype is surely overblown. When was the singularity supposed to arrive? Three years ago, Altman was saying it was going to be 2025, right?
Yea, I forgot a big one: blockchain.
Most of us on some level felt confident that AI would completely revolutionise our society.
Seriously? Maybe I’m just getting old, but having lived through the dot-com hype in the 1990s, the XML hype in the 2000s, and the cloud hype in the 2010s, this has been an utterly predictable hype cycle around AI. This does not mean that AI is completely useless. It will impact the world but not in the ways the snake oil salesmen are claiming. This is very similar to the internet. The best bubbles are built on a kernel of truth that is then blown way out of proportion and wrapped in layers of snake oil fabrication.
I personally can't wait for this to end, and for everyone to collectively get back to waiting for whatever the next golden ticket that's supposed to solve all of our problems turns out to be.
So, you’re committed to falling for the next one, too, eh?
I’d like a “Bonsai 2.8T.” That is, something that is near the Fable/Sol/K3 class, but capable of running locally on consumer hardware.
Yea the performance/price ratio for Deepseek is off the charts. I’ve been using V4 Flash a lot lately and it’s quite good.
I don’t think you’ll get full Fable performance at that level, at least for a while, but I’ve been watching some of the 1-bit models (e.g. Bonsai) with interest. Perhaps we can drive parameter count up on local models while still keeping memory consumption reasonable for consumer hardware. So, for instance, running models with 1T parameters in 128 GB systems.
That and the actual algorithms that power AI are well known and there doesn’t seem to be any secret sauce at that level.
I think there are two things happening here.
1. There is a “space race” mentality happening at the national level with respect to AI. So, China is committing to the race.
2. Even if the race turns out to be a dud, China is hoovering up massive amounts of data as customers throw everything into their prompts. This is useful for all sorts of national objectives. Why hack when you can just put up a shingle that says “Artificial Intelligence” and customers hand over their data willingly?
Either way, China wins.
Better than average chance I’d say. I suspect they are hoovering up EVERYTHING that gets sent to them. Whether that’s a problem or not depends on your data. I do wonder how many security tokens they get in the stream on a daily basis.
Cf Microsoft v Intel circa 1995
Thanks for the link. I need to evolve my own system in this direction.
Interesting, thanks. Looking at the model card on Huggingface, it’s combining the Qwythos and Qwable fine tunes from Empero.