Yeah MoE is a little worse for the same size, but you can often run bigger MoEs at respectable speeds even on cpu ram offload. The dense models really need to be 100% vram
HN user
electronsoup
It gets into loops quite often, and surprisingly often gets the edit tool call wrong
I find that running better quantization, like Q8 tend to prevent this even though its a bit slower to run, it saves overall time with less churn
Using 3.6-27b is even slower again than 3.6-35b, but I find the accuracy really pays off
in secret is impossible without the whole world knowing.
I'm curious about why this is
Outside of an actual test detonation, presumably this could all happen in a secure place?
They didn't say they had never traveled south though
If you put that behind an API, you could sell the service much like the AI providers
Why this is not a PR for llama.cpp
likely not well thought out
Or it has been, and cruelty is the point
Surely this will get arbitraged like anything else, where fans who get picks will onsell tickets
So now we need to run farms of spotify accounts playing songs to get our concert tickets?
You may need to move on to other services like Apple Music
So how many of their employees are now familiar with the codebase? zero?
At what point does spanish internet become too unreliable? There was a thread the other day about someone's CI jobs failing due too this.
the GPU apps we are building with them are
I can't help but get the feeling you have use-case end-goal in mind that's opaque to many of us who are gpu-ignorant.
It could be helpful if there were an example of the type of application that would be nicer to express through your abstractions.
(I think what you've shown so far is super cool btw)
worth their salt
That's a big assumption. Often there's no time to do things right, or no money, or lack of oversight, and so on.
Not every company is staffed by empowered and highly motivated staff
and most OS do enable it by default
Whenever I see a chip like this, I think "why wont my company let me use a decent computer"
Which mainboards are cheap and have 4 pcie16x (electrical) slots, that don't need weird risers to fit 4 GPUs
If it was so important, wouldn't he just filibuster it till he got what he wanted?
but then again you need to plausibly explain why was someone operating your car while you were not aware of it.
There is no such requirement.
and they're just as capable as using AI as anyone
Wouldn't the assumption be the opposite, in that AI is magnifying the decision making of the engineer and so you get more payback by having the senior drive the AI?
I guess the end of copyright is near if this is fine to put on a corporate website
That's some weird gatekeeping. The hardware they do support is supported well.
Perhaps they mean ISAs
oh awesome! I had assumed they were just targeting M1/M2 for the time being
The only way to accomplish this at scale is to build something that is legit better and let the market decide. Anything else is just principled wishful thinking.
No they need to tariff/ban things that are non-EU
How effective would this setup be if the parent company in the US is ordered to order the EU subsidiary to do something not in the interests of the EU?
I'm curious about the iOS situation