HN user

bavell

1,655 karma

Solopreneur, full-stack web dev with a healthy helping of devops, happy k8s user, Arch enthusiast (btw), amateur guitarist.

Great men I admire: - Jimi Hendrix - Claude Shannon - Carl Sagan

Posts2
Comments983
View on HN

My rock collection causes concepts to enter my brain, but I don't think I'd say they're communicating with me, nor I with them.

after I realize what's happening, I tend to take a little longer for each reply so they figure out it's faster to just do the research on their own most times.

Agreed, and I do the same. They still get a courteous reply, but they also feel a little "pain" when they don't get a timely answer - an effective teacher.

I'm not a Windows user—I only use it for gaming—so I don't really know how to get around this issue.

5ish years back I used to have a PCI passthrough via OVMF [0] setup for my GPU and my windows VM (Arch host) so I could game on windows.

Then I realized Proton/wine had gotten good enough to play all my games (I don't play AAA competitive shooters) and I dropped the VM and never looked back.

I would encourage everyone to give Steam/Proton on Linux a shot if you haven't recently and see if you're able to drop windows for good. These days, I don't even look at compatibility - 95% of games work OOTB and the other 5% work by changing the proton version (i.e. proton-ge). YMMV of course but I've been much happier without windows on my system.

[0] https://wiki.archlinux.org/title/PCI_passthrough_via_OVMF

What a weird time for our industry. On one hand, small teams have never been able to move faster than right now.

On the other, the economy and market conditions are brutal for the little guys. Incumbent behemoths hoovering up value, talent and financing.

Instead of shaking things up as usual when a major paradigm shift hits, AI has mostly been a centralizing, consolidating force. Not that I was expecting it to be otherwise, but it's certainly dismaying to witness.

Or am I being too pessimistic / glorifying the past?

In a reductive sense, yeah it's a bit silly. But zooming out, I can understand. Sucks to have your hand forced. Sucks to be let down. Sucks to watch something that was great fall from grace.

Thanks for Ghostty, been my daily driver for awhile now. Hope the rest of your day/week goes much better!

The Prompt API 3 months ago

Perhaps you could generate a few tokens before the entire model is downloaded, but since every token takes a potentially different "path" through an MoE model, you'd still need to wait for the entire download before getting deeper than a handful of tokens... which is not really a UX improvement imo.

I would also expect to see it taking exponentially longer to process a prompt. I don't believe LLMs work like that.

Try this out using a local LLM. You'll see that as the conversation grows, your prompts take longer to execute. It's not exponential but it's significant. This is in fact how all autoregressive LLMs work.

Yesterday I was playing around with Gemma4 26B A4B with a 3 bit quant and sizing it for my 16GB 9070XT:

  Total VRAM: 16GB
  Model: ~12GB
  128k context size: ~3.9GB
At least I'm pretty sure I landed on 128k... might have been 64k. Regardless, you can see the massive weight (ha) of the meager context size (at least compared to frontier models).

As a user, I _expect_ the cost of resuming X hours/days later to be no different to resuming seconds or minutes later.

As an informed user who understands his tools, I of course expect large uncached conversations to massively eat into my token budget, since that's how all of the big LLM providers work. I also understand these providers are businesses trying to make money and they aren't going to hold every conversation in their caches indefinitely.

Definitely some problems in the current system, broad and creeping executive overreach extending back decades now.

Pretty sure stealing from stores is already illegal, not sure I understand your analogy... lots of case law / precedent there.

Nope, the original tariffs were under IEEPA, then Supreme Court ruled they didn't have authority to use IEEPA, so they had to drop those tariffs and start working on refunds. It'd only have been illegal if they kept the tariffs after the ruling.

Lot of propaganda & emotions around this straightforward chain of events.