HN user

mft_

5,112 karma
Posts2
Comments1,670
View on HN
Laguna S 2.1 1 day ago

Thanks for flagging. From the few benchmarks I can find, it looks there or thereabouts with Qwen 3.6-35B-A3B, or maybe a touch below. I'm interested to compare a model that is a big jump larger with pretty impressive benchmarks, but more heavily quantized to fit.

Laguna S 2.1 1 day ago

Looks impressive, and this size fits achievable home hardware.

That said, if someone would kindly quantise this down for the 64GB paupers, that would be appreciated. (I know there’s likely degradation, but some people reported good results with a 2 bit version of Qwen 3.5 122B, and this is starting from a higher point. Would be interesting to try, at least.)

Edit: someone in the process of doing so: https://huggingface.co/vcruz305/Laguna-S-2.1-GGUF

I think the premise is wrong: I don't think Google is more hated than all of the organisations you list.

Within the tech world, I believe people generally dislike meta, Palantir, Oracle, and Microsoft more than Google.

Out in the real world, in the US at least I'd bet money that UnitedHealth must be more hated. And anyone with at least a modicum of rationality would probably dislike Big Oil and Big Junk Food more too?

---

That said, I think this is an excellent cautionary lesson:

We secretly started to believe... ...we hate them more because we dared to believe, and Google let us down.

Not from a single data point. You have no way of knowing whether this is, on the one hand, a single reasonable request or, on the other, the start of a pattern of bureaucratic pain.

If your entire business decision is driven solely by an emotional response to Hetzner having a policy that needs ID, maybe you’re not the right person to be making these sort of decisions?

Qwen 3.8 4 days ago

I think everyone is hoping this!

It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.

Transcribe.cpp 4 days ago

I tried this on my Mac soon after launch and it was consuming a significant amount of processor cycles even just sitting idle in the menu bar. (From memory, ~20% of an M1 Max.)

It may have been an early issue but with no obvious way to interact and report the issue and, eh, Google’s general attitude around customer satisfaction, I just gave up and deleted it again.

Transcribe.cpp 4 days ago

I’m really glad to hear this!

A while ago, I auditioned about 10 different STT apps on my Mac, with this realtime/streaming transcription as a goal. I failed to find that feature in an app I was happy with, but settled on Handy as the best option otherwise. So if Handy adds this, it will be perfect!

Same here, in Germany.

Some categories are especially exaggerated: when needing a number of new wardrobes a few years ago, we struggled to find anything that was remotely similar in price or value to IKEA. The big retail park furniture stores were 2-4x the price for similarish quality (i.e. chipboard; modular construction). Smaller higher quality stores (e.g. plywood/real wood rather than cardboard or chipboard) were roughly 4-8x more expensive. And one really nice local option was roughly 10x more expensive. [0] (We didn't even explore bespoke/handcrafted for obvious reasons.)

Wardrobes, in particular, are begging for disruption - essentially a relatively modular approach, higher quality materials than IKEA, (much) lower cost than Moormann.

[0] https://www.moormann.de/de/schrankone.html

You just need a sufficiently self-interested actor that sees open ecosystems as a necessary part of reducing their own risk profile, relative to the alternative of complete reliance of a third-party business that can take an exorbitant cut and/or Sherlock them at any time.

This would be an argument for an organisation developing its own model; but not per se for releasing the weights openly.

The possible explanations (I'm aware of, which overlap somewhat) for spending large amounts of money on models then releasing them for free (i.e. the current Chinese approach) are soft power, marketing for a future paid model business (i.e. competing with the US models for customers and mindshare during the time you can't compete directly at the bleeding edge), and/or a geopolitical move to diminish the value of the US's frontier model companies.

Open models are probably also comparatively astronomically expensive to train - just less so than the frontier models because they’re somewhat smaller, +/- the creators are more incentivised to focus on getting more from less compute because they’re have to, +/- they rely on distillation of the frontier models and this is more efficient.

But efficiencies aside; creation of open models still requires a lot of money and compute from a large organisation which is willing to accept zero return for that spend. This largesse is unlikely to continue forever; so the question is which will crack first, the frontier models’ business model or the fast followers’ generosity?

Yes.

If we reduce back to the “LLMs are next word prediction algorithms” and they have a huge training corpus including positive and negative human interactions, it’s not crazy to think they’ll be influenced by the flow of those learned interactions and respond subtly better to positive interactions than negative.

Really nice visuals!

Could you give a little more information about the 'stack' you're using to display the clock?

From this:

This minicomputer will power the clock. The Pi 3B+ seems just enough to render some simple animations.

I also tried with a Raspberry Pi Zero 2, but 512 MB of RAM were too little to run a modern browser, sadly. I think a Pi 4 might be a good sweet spot between processing power and price, currently.

...are you essentially displaying a full-screen browser window in the default Raspbian window manager and then running local Javascript for the clock itself?

Would we do a trip like this again? It's certain a lot of travel. We weren't very spontaneous - most of the trip was planned out way in advance, along with hotels. Having 2-4 days in each place is like taking a series of minibreaks, which is delightful. But it can be exhausting. I don't want to complain that my diamond tiara is too tight, but there comes a point where there is such a thing a too much holiday.

This deserves a little more unpicking.

Something that I hear very few people discuss is that not all holidays are equal and you need to be aware of what you really need before choosing.

If you want to see new places, people, cultures, food, whatever - that's one thing. But if you're tired and need to recharge (sadly, this was often my experience in a corporate job - an endless sawblade cycle of work -> recharge -> work -> recharge) then don't go interrailing - go somewhere quiet and plan to sleep and lie around for a week or two. City breaks and events (e.g. festivals, sports, etc.) fall into this category for me - they are fun and make life better, but expect to come home hopefully happier but also tireder than when you went.

Agree. It's a fairly minimal list with very few extras added.

Current /context on a fresh session (compare to that above) is:

  Opus 4.8
  15.8k/1m tokens (2%)
  System prompt: 4.5k tokens (0.4%)
  System tools: 7.9k tokens (0.8%)
  Memory files: 441 tokens (0.0%)
  Skills: 3.1k tokens (0.3%)
  Messages: 8 tokens (0.0%)
  Free space: 984.2k (98.4%)

The main ones missed immediately were web access/search. Then the to-do list features (it was a nice surprise to try OpenCode and see this working immediately.). There were a couple of other niggles but it was a few months ago. Also, this may not be common, but it seemed to struggle to edit effectively (driven by Qwen 3.6 35b/27b) and often rewrote whole files instead.

Apologies, you're right - I used imprecise terminology. The entire initial JSON structure that was sent from Claude Claude to the LLM at the start of a session was 162k. This included the system prompt together with a list of tools (some with very extensive explanations), MCP server details, etc.

I was simply supporting the article's data - their reported 33k tokens is probably roughly 150-165k.

It's easy to add using plugins.

Sure, but you have to add almost everything, no? It deliberately only comes with read, write, edit, and bash. My point wasn't that you can't add stuff, but that I'd just rather use an harness that's a bit more full featured from the start.

(Pi is a bit like old 3D printing where fettling the printer to work is a central part of the hobby. I'd rather just buy a Prusa.)

Early on in experimenting with local models, I found that hooking them up to Claude Code worked very well, but it was also really slow.

I used mitmproxy (setup assisted by Claude, natch) to capture Claude Code's entire initial system prompt and the whole thing was (I just double-checked) 162k of JSON.

This led me to start experimenting with Pi, OpenCode, and Hermes...

I’ve been wondering about something similar - a system that enforces (or does the heavy lifting) of dividing a large task into smaller sub-tasks so that it’s easy to run/check/test each one independently - even on a fresh model instance if needed.

This is based on the observation that the medium-sized open weight models (~20-35b) are very able to one-shot smaller discrete tasks but seem to lose their way project managing themselves through larger tasks that have multiple steps.