HN user

fgonzag

2,061 karma
Posts0
Comments571
View on HN
No posts found.

The weights are open and when prices settle down again will be runnable with less than 10k of hardware.

I can easily run it in a 8 bit quant with the 4 x 48GB Radeon Pro W7900 GPUs I snagged for 2k each before the memory squeeze.

A 158B parameter model, especially in an architecture as efficient as DS4 is not that hard to drive currently if you got in before the craze, and will be relatively easy to drive with future hardware generations.

Qwen 3.7 Preview 2 months ago

In the china AI scene, there seem to be two separate types of companies.

Companies or labs like deepseek that produce less but larger and more innovative models, so seem to be more research oriented.

then there are companies like z.ai (GLM), Minimax, and Qwen which focus more on commercializing the AI and so produce far more versions, but with far less improvements between them (usually fine tunes)

Commercial providers like anthropic probably do the same thing, maybe even without labeling it like a different version if the model is similiar enough.

Nobody is serving models in BF16 precision, not even commercial providers. Especially with newer quant methods (like nv4)

The article states you can fit Q4 in 4 x 4090 and it works reasonably well.

I'd personally fo for deepseek V4 flash at Q8, hardware prices need to come down though. Once an NV4 version get released it'll be easier to run on commodity hardware.

I was on the same boat as you. I installed noctalia after having tried and not really getting a perfect arch / niri / waybar setup.

I think I'm completely done after like 6 hours which is insanely fast, and it really is everything I ever wanted. It is cohesive, easy to style, has good defaults, includes essential programs like polkit agent and notification daemon / osd, has a ton of plugins.

I should have tried it so much sooner.

There is some cross pollination. Women can play vs men, just usually don't. I'm fairly certain singles UTR is universal across players, it only distinguishes between doubles and singles UTR.

UTR can also include unranked games if one of the players submits a score and the other approves it.

[dead] 4 months ago

The whole text feels like a conversation with an LLM. Even the title.

And the content feels fluffed up, the core idea is: "I'm furious that Bitwarden doubled the yearly price and did not disclose it properly". Which is valid, but does not require an extensive article, much less one written by an LLM.

As an aside, I found out from this post about the increase, and I don't mind the increase, especially after 10 years, but I really don't like the fact that it seems they actively hid it. But I also don't think it's aa big of a deal as OP is making it out to be.

It doesn't.In tennis a 14 UTR whatever wins against a 13 UTR whatever. UTR is your effectiveness rating against every other player. Same in chess with ELO.

The issue is woman would disappear from profesional sports. Sinners 16.27 rating means that he double bagels Sabalenkas 13.29 essentially 100% of the time. The 500th ATP player has a UTR of 13.81, half a point is quite a bit stronger, do he's still very much stronger than Sabalenka. You probably have to start looking well into the thousand somethings for something that is consisently beaten by her.

Only the top 200 players make money, the top 100 good money, and the top 50 ridiculous money.

From an r/ArchLinux moderator responding to a censored user:

"I got a DM from much higher up the chain asking me to remove it. Whilst I technically don't answer to them, I do respect their wishes. They don't like someone they consider as part of the core dev teams being called out like that. What you did broke the Arch Linux CoC."

Wait so is this part not true?

You're right it was WordPad he tried.

A super basic text editor is available anywhere. He's comfortable with many of them when editing other things (like vim/nano/ed on Linux)

He just likes Notepad for his personal use. Now that you mention it I do believe he used Notepad++ at some point (only 3rd party editor he's tried). I don't remember at what point he dropped it but he didn't use it long.

Like you said, with Windows 11 replacing Notepad with a rewritten version, he's probably going to have to switch to Notepad++.

I was hoping he'd just switch to Linux. Even tried baiting him by giving a pre Linux installed fully configured laptop with a beautiful OLED panel and great keyboard. He loved the hardware and reinstalled Windows.

It's honestly short and pretty unique.

He wrote a programming language for his master thesis, so obviously he used it to write all his software. His first project was the POS/management system for his father's music store (Now famous as the Mexican company that acquired Sam Ash). I believe they didn't switch until around 2005 or so (so about 30 years maintaining it or training a software developer on it as a side thought)

He then started a large sized customs software company with i that ended up getting acquired.. Everytime the language required a new feature the devs would just ask him (like when he had to write a graphical toolkit for it because it started as a text only runtime). There is no record of this language anywhere as far as I know.

I believe around the 2000s as part of the sale of the company he rewrote the whole stack in C#. And he's been using it ever since, including the company we started together in 2013 (together doing a lot of work here). Still with good old Notepad and CSC.exe just like year 1.He curses everytime the ecosystem has big required changes (async, nuget) though I've managed to coerce him into keeping up with the times, dragging and screaming.

Change is hard.

My father is a 70-year-old software engineer who programs .NET Core in Notepad and builds using custom BAT files that build the project using csc (the outright compiler). He browses and copies files in the Windows Terminal. He is also accustomed to Linux since we deploy to it in our business, and he can do everything comfortably in the Linux terminal.

He trusts me almost blindly, yet I can’t convince him to swap to Linux even though every time he keeps fighting Windows. I'm actually fairly surprised since I'm certain he'd find himself at home almost immediately( he already is when managing servers)

I’m fairly sure it’s Notepad keeping him there, but I’ve told him there is also a Linux clone or Wine. I had been dabbling in Linux for 30 years, and it’s been about 7 or 8 years since I switched full-time and couldn’t be happier. But honestly, we're going to get there because it’s inevitable. It’s the only OS that's currently not wholly incentivized to "enshittify" itself and is actually improving at a pretty good pace due to Wayland's novelty fostering a plethora of alternative window managers.

It absolutely is a different and more insidious type of dynamic pricing.

First, you can use the airline's strategy to your advantage by planning early. It doesn't feel as unfair because everyone gets the same terms and the system is transparent and equal

WaPos daynamic pricing is simply maximizing value capture, without any way of a consumer benefitting. It's 100% lose-lose for the consumer. You always pay the maximum you are willing to pay. No discounts!

I was just answering OPs question about how airlines were transparent about their system and decided to answer it factually.

Mexican cartels absolutely use OF to launder massive quantities of money. They use OF because it's the one thats actually used by people. It's a lot easier and less suspicious to declare ridiculous high subscriber counts in OF that other platforms.

I see a ton of my peers driving around in 80k cars. I drive a 20k used one.

I'm planning a writing a ROCM inference engine anyways, or at least contributing to the rocm vllm or sglang implementations for my cards since I'm interested in the field. Funnily enough, I wouldn't consider myself bullish on AI, I just want to really learn the field so I can evaluate where it's heading.

I spent about 10k on the cards, though the upgrades were piece meal as I found them cheap. I still have to get custom water blocks for them since the original W7900s (which are cheap) are triple slot, so you can't fit 4 of them in any sort of workstation setup (I even looked at rack mount options).

Bought a used thread ripper pro wrx80 motherboard ($600), I bought the cheapest TR Pro CPU for the MB (3945wx, $150), I bought 3 128Gb DDR4-3200 sticks at 230 each before the craze, was planning on populating all 8 channels if prices went down a bit. Each stick is now 900, more than I paid for all 3 combined (730 with S&H and taxes). So the system is staying as is until prices come down a bit.

For AI assisted programming, the best value prop by far is Gemini (free) as the orchestrator + open code using either free models or grok / minimax / glm through their very cheap plans (for minimax or glm) or open router which is very cheap. You can also find some interest providers like Cerebras, who get silly fast token generation, which enables interesting cases.

My first switch was to open code + open router. I used it to try mixing models for different tasks and to try open weights models before committing to the hardware.

Even paying API pricing it was significantly cheaper than the nearly $500 I was paying monthly (I was spending about $100 month combined between Claude pro, chat gpt plus, and open router credits).

Only when I knew exactly the setup I wanted locally did I start looking at hardware. That part has been a PITA since I went with AMD for budget reasons and it looks like I'll be writing my own inference engine soon, but I could have gone with Nvidia and had much less issues (for double the cost, dual Blackwell's vs quad Radeon W7900s for 192GB of VRAM).

If you spend twice what I did and go Nvidia you should have nearly no issues running any models. But using open router is super easy, there are always free models (grok famously was free for a while), and there are very cheap and decent models.

All of this doesn't matter if you aren't paying for your AI usage out of pocket. I was so Anthropics and OpenAIs value proposition vs basically free Gemini + open router or local models is just not there for me.

Yeah, honestly this is a bad move on anthropic's part. I don't think their moat is as big as they think it is. They are competing against opencode + ACP + every other model out there, and there are quite a few good ones (even open weight ones).

Opus might be currently the best model out there, and CC might be the best tool out of the commercial alternatives, but once someone switches to open code + multiple model providers depending on the task, they are going to have difficulty winning them back considering pricing and their locked down ecosystem.

I went from max 20x and chatgpt pro to Claude pro and chat gpt plus + open router providers, and I have now cancelled Claude pro and gpt plus, keeping only Gemini pro (super cheap) and using open router models + a local ai workstation I built using minimax m2.1 and glam 4.7. I use Gemini as the planner and my local models as the churners. Works great, the local models might not be as good as opus 4.5 or sonnet 4.7, but they are consistent which is something I had been missing with all commercial providers.

OLED, Not for Me 6 months ago

I have a 49" QD-OLED panel. I have never been one to find visual artifacts distracting, but fonts were awfully jaggy in Linux to the point I spent a week tinkering with font config and almost switched panels to a larger miniled since code looked horrible. And I'm someone who was fine with horrible VA low res low quality screens back in the day.

The sub pixel geometry on samsung's qd-oled needs very specific font configuration to be correctly displayed, and even then it just stops looking bad.

Apple's M5 Max will probably be able to run it decently (as it will fix the biggest issue with the current lineup, prompt processing, in addition to a bandwidth bump).

That should easily run an 8 bit (~360GB) quant of the model. It's probably going to be the first actually portable machine that can run it. Strix Halo does not come with enough memory (or bandwidth) to run it (would need almost 180GB for weights + context even at 4 bits), and they don't have any laptops available with the top end (max 395+) chips, only mini PCs and a tablet.

Right now you only get the performance you want out of a multi GPU setup.