HN user

pferdone

257 karma
Posts2
Comments104
View on HN

The consensus right now is that Qwen3.6 in its 27B and 35B-A3B versions is better for coding whereas Gemma4 is stronger when it comes to OCR, audio transcription and the likes. Margins are slim though and the harness at these model sizes is the most important factor.

I can see that and I don't know your setup, but there are people pushing >70t/s with MTP on a single 3090, with big contexts still >50t/s. 64k is not a lot for agentic coding, and IIRC 128k with turboquant and the likes should be possible for you. r/LocalLLM/ and r/LocalLLaMA/ are worth a visit IMO.

EDIT: just found this recipe repo, may wanna give it a go: https://github.com/noonghunna/club-3090

EDIT-2: this can also shave off a lot of context need for tool calling -> https://github.com/rtk-ai/rtk

The Guardian, especially for their podcasts, is the only news website I am paying and have ever payed for. And I pay more willingly than any other newspaper would get from me for their paywall stuff. It‘s that valuable for me to support this approach.

I mean good for them for improving their user‘s experience. I am just surprised that maintaining your own webpack config seems like such challenge. Enabling chunks to split code to improve loading times or dynamic imports, it‘s not exactly rocket science. Yes DX and all improved as well, but CRA is such a bad choice for production code. It gets you going fast(er), but as soon as you have to tweak it, it seems like devs jump to the "next" framework that already has those tweaks builtin…until they hit the next roadblock. Thus you never learn what actually makes it all work.

When I installed my DIY backlight solution (ala AmbiLight) to my TV I also had a red shift for some of my LEDs. But that was mainly due to the fact that not enough voltage reached the later LEDs on the strip to power green/blue because they need higher voltage than red.