Well it seems to be hug-of-death'd - how are you loading 41k points of data?
HN user
Catloafdev
I wonder how they detect this kind of thing. Seems like this is going to be a perpetual issue until it stops being worth doing.
Side note, didn't they stop releasing real thinking tokens for Fable? Or is it still part of some subs or API usage?
It's so full of bullshit marketing prose I don't even know what it actually is.
Well that surely isn't a good sign of their financial health.
I would have expected this to happen post-IPO.
If you read the rest of the article, it breaks down the cost, it's not a misleading title.
I expect this will kick off a series of 'GLM helped HuggingFace where Fable wouldn't' type headlines.
I mean not only does this not actually address P vs NP but the solution provided isn't even a valid sudoku grid.
There are fully open models, it's just not as common because that's basically giving away the sauce, which is non-viable for many.
As someone that's generally for the proliferation of open models, I want to take this seriously, but it's really difficult when it was clearly written by AI.
I guess they fired whoever used to write copy for these things.
Edit: to be clear, I'm not trying to just dunk on them, I think it's actively hurting their own point to do this, and counter-productive when people can easily clock it - it makes some percent of the audience immediately tune out.
This is just wrong on multiple levels, the open source model ecosystem is very much a thing.
It's a LLM model, not a phone app.
Available on HuggingFace: https://huggingface.co/collections/prism-ml/bonsai-27b
Doing some naive math, the F16 filesize is ~53.8gb, the 1-bit version is ~3.8gb, about 7% of the original size. The F16 size is roughly 2x param count, so that gives a rough ballpark of ~110B.
DDoS protection is pretty essential. I highly encourage you to read through what they offer if you're not familiar.
So the complaint is that one day they have such a strong monopoly that they can freely turn evil?
Just want to make sure I understand the real issue here, because that sounds like a lot of fearmongering to me.
They're creating optional opt-in protection layers for the services they operate.
I genuinely don't understand these generic complaint comments.
Are you complaining that they offer too much? Or do you believe nobody is offering similar services?
Are you genuinely asking "does anybody even need any of their services"?
What Cloudflare competitors offer a similar range of services?
How is it hyperbolic when it's literally the subject of the post? Did you not read the OP?
Those are people without a better option.
Big difference vs xAI, where the sentiment is valid.
Well that's not what's happening here, lol.
If it weren't a real problem, these types of articles and services wouldn't exist.
That's pretty accurate to what I've seen, I'd definitely recommend a smaller model for active agentic use on that hardware.
It definitely seems to be the leader on 'general intelligence' on this hardware from my casual usage, but the newer Qwen or Gemma series models are much more usable speed-wise for agentic use, and often is as good or better than DS4 on that front.
Heads up, you can absolutely run DS4 Flash on a 128gb machine - I have it running on my Strix Halo box right now.
It looks like it has 4 tiles on it, no?
The demand is not coming from 'normal Apple customers' it's coming from people who want a machine that can run local AI.
It has nothing to do with Macs being especially good at AI. It has everything to do with being one of the last 'cheap' devices being sold with that much unified RAM.
For most coding or agentic tasks, Qwen 3.6 27B likely outperforms, yes.
For 'general intelligence', DS4 Flash seems to be a noticeable step up still.
Oh, it is. I was looking at the Huggingface repo which listed the lower number at the top of the page, looks like that's wrong.
299B for Hy3 vs 284B* for Flash
Edit: fixed, got bad info
Curious how people feel about this compared to DS4 Flash, given they are pretty close in size. Also curious how well it holds up to heavy quantization.
DS4 Flash can currently run reasonably well on systems with ~96gb+ RAM, I wonder if Hy3 can compete there.
Looks great - is there any info on what server resources are actually required per feature or user count?