In absolute terms, you are absolutely right but relatively i don't think so. if we were to compute private money divided by state resources for AI, i think china might have more share than the US. also, even if government spend in the US on AI is so high, shouldn't we get then some models for free? maybe thinking machines is doing that, but its funded privately by a16z.
HN user
dwa3592
Could I get some extra humor on the side please? Thanks.
In the end its really VC money (US) versus State resources (China). In my personal opinion, building reliable LLMs is kind of a fundamental science problem which if done right has the potential to help everyone regardless of the background, so it should definitely be funded by states resources (taxes etc), which is what China is doing. In them doing so, the rest of the world also benefits, I think its a net win.
this is good to see. i also trained a stt under 500kb for sub dollar chips. it had about 20 words that it could understand(like start, stop, left, right, go, up etc) and then the spell mode where you could say the word spell and then say the individual english alphabets and close with spell. it was super fun to work on. these tend to be extremely unstable though, like confusion between p and t (at least for my accent). will have to try this one now.
lobotomizing is what i was thinking. i don't have a million dollars.
Can I just say that they look awful?
waiting for - "Running Kimi K3 on X years old hardware".
This is super exciting. I really need to buy better hardware to try this stuff.
lmao, wasn't xAI caught doing this recently? moreover at least moonshot is being honest about it.
I am not even sure if you are being sarcastic.
https://www.zillow.com/homedetails/80-Prentiss-St-San-Franci...
check this out.
Antigravity sucks so bad that I have started to feel that google really doesn't wanna compete, they just wanna hang in there at number 2 or 3, to just annoy the number 1 and 2.
glad to know. thanks :)
dario has been saying open source models are dangerous. who knows who is listening to him.
Its really easy to argue against local models because when it comes to quality, you can argue using the tokens/sec. and when it comes to speed, you can argue using the parameter count. This is not compared to the frontier stuff but it is the frontier of last year that now runs on a local machine. It was impossible to do this last year.
It will, but the process at this point is SSD bound rather than compute bound. On a bigger machine, Apple silicon must help but I don't have a bigger machine. I can think about this more and will make changes if that helps.
>Yes, it’s technically running, but not in a way that would be useful by normal LLM standards.
What are the LLM standards?
Do you know how many people use perplexity? I know many people who are not software engineers or tech workers and have a LLM subscription for rewriting their stuff (non-native english speakers) in english. There are many use cases for running good models locally. Maybe not for you, but someone might find this beneficial.
Why isn't there a video of it?
agreed!! in my heart i really wanted to say by the end of 2026 but wanted to add some wiggle room in case they start to ban open source AI development.
it's a 16GB machine. i am proud of this machine so far.
i have been optimizing for that. for now samosa is capped at using half of the avaiable cores and switching between them, which keeps the system 'less hot' as it would have been. i will also release better thermal control in the next release. at this point its basically sacrificing about 20% of the speed to keep the hardware less stressed (and hot).
i am working on making it faster but to me 7-9 tokens/sec feels very good. it was 0 tokens/sec a year ago.
>It's quite telling you didn't use Qwen3.6-35B-A3B locally to build that
that would have run into a race condition unfortunately ;)
but there is a sample landing page + a python function on the repo which shows what the model produced. my goal is to integrate the local model in my workflow so that claude/OAI can call this model for basic stuff.
I have a prediction. By the mid of 2027, we will have >200B MoE models running on basic consumer hardware.
I am running Qwen3.6-35B-A3B locally on my 16GB mac with 7-9 tokens/second. Link - https://github.com/deepanwadhwa/samosa-chat
This is a GPT4 level model running locally with a decent speed on a 16gb ram macbook air.
Thanks for testing. Actually if you open the samosa app; it will show you memory consumption, tokens/sec etc. Let me know what you think.
me and my wife made up a word in 2024 for this. the word doesn't exist in any language. we say it to each other all the time. even if i give you the spelling for it, you will say it wrong. i recommend everyone to do something similar. i should do it with my parents too.
>One reason I am not running local models on my Mac now a days, is My Mac book pro is getting heated more
Samosa was optimized for keeping mac's temp in check. This is one of the important reasons for creating this - no excessive wear/tear to the machine. https://github.com/deepanwadhwa/samosa-chat#the-three-princi...
maybe give it a try and let me know if your macbook is getting heated while running this?
i was pleasantly surprised to find out it hadn't been done until now.
my iphone is already kind of dumb. i don't have any social media apps. i have all notifications turned off (except ringtone for call). i only use my phone for calling and texting. it feels very expensive for those 2 operations. but i have the choice of downloading whatever shit i want to download.
This marketing video on the page is nice!! can't wait for the hardware to get cheaper to live the AI life i wanna live.
What I agree with in this article is that there is shit ton of slop on the internet today. What I disagree with is that if pangram says it's human, then it's not AI. That is hardly the case. Pangram fails spectacularly in detecting AI.
https://github.com/deepanwadhwa/ai_detector_fails/blob/main/... - I posted this link in another thread the other day. I might do more of these tests in the upcoming days and put universal jailbreaks for people to fool these 'AI detectors'.
I like the timer but just give the option - play with timer (pro), play without timer (relax). split the results as well.