HN user

marci

511 karma
Posts1
Comments329
View on HN

But with Apple's AFM 3 architecture, we might end up with huge SOTA adjacent on devices with limited RAM.

They use a technique where you only load between 1B and 4B of a 20B dense model for an entire prompt run, not token by token like a MoE, and use mostly the low power ANE instead of GPU cores.

Now, imagine if/when they scale up to 100B or more? On a chip using 2W?

A few words on DS4 2 months ago

"That’s where EMO comes in.

We show that EMO – a 1B-active, 14B-total-parameter (8-expert active, 128-expert total) MoE trained on 1 trillion tokens – supports selective expert use: for a given task or domain, we can use only a small subset of experts (just 12.5% of total experts) while retaining near full-model performance."

https://allenai.org/blog/emo

Unfortunately, the most extreme is that it's the new normal that now, there's >0 chance that someone, whether they are a US citizen or not apparently, child or adult, can end up in a camp, with no due process.

Claude Sonnet 4.6 5 months ago

Imagine, a llm trained on the best thrillers, spy stories, politics, history, manipulation techniques, psychology, sociology, sci-fi... I wonder where it got the idea for deception?

Qwen3-Coder-Next 6 months ago

Their issue with the mac was the sound of fans spinning. I doubt a dedicated gpu will resolved that.

If your family members ever had to mount an ikea furniture or equivalent, they'll probably have an as easy or easier time replacing a part on a fairphone. Especially for the battery. At least for version 3 and older. I don't know for later models. If you know how to swap batteries in a tv remote, you know how on this phone.

It makes more sense to word it like this when you take into consideration historical trends, like drowned towns for lakes or dams, highway system along redline, thriving neighbourhoods erased to create parks… often preceded by violence and little to no compensation.

Any repo? even if not production ready. I'm curious about how you approached replication, compared to mnesia or couchdb, especially now that erlang natively supports json.

Qwen3-VL 10 months ago

My precedent post should have answered this question. But since it didn't, I think I'm ill equipped to answer you in a satisfactory fashion, I would just be repeating myself.

Qwen3-VL 10 months ago

it's sometimes not really a matter of which one is better but which one fits best.

For example many have switched to qwen3 models but some still vastly prefer the reasoning and output of QwQ (a qwen2.5 model).

And the difference between them: those with "plus" are closed weight, you can only access them through their api. The others are open-weight, so if they fit your use case, and if ever the want or need arise, you can download them, use them, even fine-tune them locally, even if qwen don't offer access to them any more.

It all depends on the scale you use. At the individual, sure. But it's like cars. They keep getting more effecient, yet total energy consumption keeps increasing.

The further we can go, the further we will go.

The more CPU power we get, the more JS heavy websites get.

The more images we can generate, the more we will generate.

The more we can do, the more we do, whether we should or not.