I miss the days of 4changpt.
HN user
3abiton
Yeah, Qwen3.5-122B-A10B-NVFP4 produces better responses than Qwen3.6-27B-NVFP4 (both from unsloth), but I'm mostly using them for programming in various ways, mostly Rust, Clojure, Python and JavaScript, and some translations tasks, but not much more than that, so YMMV. > Edit: as a concrete example, I'm working on a "optimization framework via agent harness" right now, Qwen3.6-27B-NVFP4 is often unable to actually complete the optimization within 100 turns, while Qwen3.5-122B-A10B-NVFP4 has no issues finishing within ~50 turns or so.
But why are you using Qwen3.6-27B-NVFP4 compared to the FP8 or full version? In my experience the Q8 of 27B is on par sometimes better than 122B. I am experiemnting witb higher quants for 122B to fit on my Strix Halo, but still, the difference honestly for my workflow is not that much. I just wish they released 3.6-122B version.
in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far.
I have used both Qwen3.6-35B and Qwen3.6-27B locally (both Q8 quantized with llama.cpp). I have also used antirez's quant of DS4-flash. They all performed within the same tier, DS4 being a bit more efficient, but they all gave really good results, mainly used for bash scripting, debugging, python and some C++. I am curious what type of applications/langauges failed with Qwen? One thing to note, the chat templates were "broken" for qwen models and had to debug it, there are already effort on this. Tbh, the same with gemma.
Also at scale Kimi (or DeepSeek and others) don't have to worry about local deployment. The number of param keeps growing, I highly doubt consumer hardware will catch up. I am just glad Qwen team is still putting up great bangers.
It would be so cool as a feature to allow people on HN to make some of their favorited posts/comments public.
Another crucial difference is: ROCm is open source, while cuda isn't. Yes it's tough to port things to newer gpus, but in theory people did it (therock spear heading the ROCm patch for strix halo before the offical support)
Honestly, root detection is a cat and mouse game. So many ways to spoof it. Same with Play Integrity. If the goal is to prevent bad actors, it will never work really. It will be just a big headache for the citizens. Look at how gatekeeping certain websites is working. VPNs became mainstream. So are proxies too especially for bad actors. Malicious actors have just to pay up a service to do so. And I am sure it will be the same with Android (it already is with keyboxes you can purchase to spoof play integrity)
This was a very enjoyable read, and loved the distribution inclusion, rather than point statistics. I hope there are more tests on different factors like GPUs (nvidia vs amd vs intel), screens, mice, and ultimately windows to see if there is really a difference. I don't play competitve gaming, but careful data driven approach would put a lot of the debated questions to rest.
Funnily the current high end Mac Studio are not suited for current LLMs. M3 Ultra is "quite an old" chip for AI, despite its bandwidth. The issue for running local models (especially LLMs), you need few things to align really well: 1) compute power (affecting PP) 2) VRAM capacity (affecting model size you can load) 3) Bandwidth (somewhat affecting decoding speed).
The issue with the M3 chip is the compute performance, as it doesn't fit well the transformer architecture. This changed with the M5 (apple baked their own matmul into the chip), which would significantly speed up PP (and video/image generation btw), making the M5 Ultra significantly faster than the M3 Ultra and in practice much more usable. You can try to load Kimi or GLM on M3 Ultra, but it's not usable. Now the M5 Ultra is not out yet, but undoubtedly it will be a superior offering, and shilling 15k on 512GB version is actually reasonable (if it's every priced remotely around that tag).
This is technically impressive, but is it usable in practice?
I just don't think that I can ever trust an xAI model knowing that they are actively trying to shape its replies to fit a political narrative. How can you trust their models to be reliable in a business setting with the foreknowledge that their models are being nudged around in the backend?
https://www.nature.com/articles/s44387-025-00048-0
Large language models reflect the ideology of their creators.
It has very interesting insight from analysis of LLMs political leanings. Spoiler alert: they all have political bias.
I still don't fully get the additional value over tmux, beside notification regarding the agent status?
I run OpenWrt on 2 of my routers. It's really amazing the level of control. Although, I am now building a better approach: a mini PC as a managed linux router replacing my ISP (no wifi). Then my 2 wifi routers for wifi.
The target audience is different. Coding is mainly a trade of the tech savvy, who like many on r/localllama users do not hesitate to deply on 16GB Vram gpus. Even if so, it is estimated that within 2 years we will be able to run Claude 4.8 on consumer hardware give the rate of improvement of open-weight LLMs, which will put more financial pressure on "paid" labs. It's just a matter of rate of improvement which is shrinking between open-closed models.
The "trick" is well documented in their Deepseek-OCR paper, that builds on plenty of other work. It's just not simple to just switch a commonly used LLM architecture to a new one, but I don't doubt most frontier labs are already experimenting with it. This by itself a very active field of research.
They are heavily bogged down by bandwidth unfortunately. The macs are on another level. If Apple decides to release AI dedicated hardware, it would dominate this space (consumer AI).
No wonder why deers are seen as snobbish. All the illiterate ones were shot.
Honestly it's just a hierarchy difference between the two countries. In the US, tech/fin/military companies have the upper hand compared to the government (fragmented between 2 parties). Despite the sharades with Anthropic, Tech-fluencers are in control. Compared to china, the government (dictatorship) has more control over Tech companies (take any example from the past 10 years). For them, undermining the US AI supremacy is an objective, and releasing open weight models is the way, and I'm all for it.
One day, maybe not far from now, a breakthrough will allow huge LLMs (say 200B in size) to run well on an old 5 year old Dell desktop.
I think there will be specialized hardware (beside GPUs) that would be custom made for LLMs. Yes TPUs exist, but mainly for datacenter. GPUs exist, but they are adapted from mainly graphic application. Once all the demand from data center dries up, innovation will kick in.
And a big thing that's missing is ... the harness comparison. Ot plays a very big role. I use forge, and I have been inpressed with what it can do given all the limitations of local models.
I think nearly everyone mentioned Qwen, so my turn I guess. Qwen 3.6 35B Q8 (MTP), on a Strix Halo, with llama.cpp. Around 40-50 t/s. Really great pefromance, I get always suprised by its capability. I used with forge-code directly in zsh. For long context 150k+) it start degrading and forgetting.
I have the same. The difference is, if you do email verification, you will "verified" status. If not, you can still add the company to your linkedin, just unverified, which is not a label.
Are there evidence that this approach helps maintain "accuracy" performance when quantized? It sounds a bit like mxfp4 with gpt-oss, which was a confusing model upon release.
Podman has lots of underappreciated features, and it's fully open-source!
Not to mention the competition: chinese open-weight models and open-source harnesses. Qwen3.6-(27B and 35B) have proven to be worthy and capable of running locally. I am confident more SMEs would look into this as a solution given the ballooning costs of API usage. You get a decent setup with an RTX 6000 Pro.
* That can still yield useful "discoveries" in certain fields, absent the discovery of new mechanics that exist outside said training data
One can argue, new knowledge is just restructured data.
I think the main concerns about LLMs is the inherent "generative" aspects leading to hallucinations as a biproduct, because that's what produces the noi. Joint Embedding approaches are rather an interesting alternative that try to overcome this, but that's still in research phase.
Step 1: Have a workshop space Step 2: ? Step 3: Profit
Qwen3.6 35b a3b is still my local champion but I may use this for auto complete and small tasks.
I second this! Using the Unsloth Q6 (I forgot the exact name). Currently using it with forgecode (with zsh), on my Strix Halo, and it's suprisingly really good. I would say slightly Similar to Haiku 4.5, plus additional privacy, minus speed. It's surprisingly really fast for the hardware, given the speculative decoding, still PP is on the slow side.
They patched the "non-existent" issue it seems. And totally denied it happened in the first place. Honestly, someone should do a dump of redacted client documents to teach them a lesson. Short of a class action lawsuit would be an understatement. This is really huge.
Even though all can be replaced by a decent mini pc with beefy memory, with lots of VMs.