HN user

biddit

301 karma
Posts0
Comments87
View on HN
No posts found.

Ah yeah, your reaction makes sense. In the circles I run in, there is a lot of hype around Sparks for inference, so my gut reaction is to respond with this type of warning.

I did not intend to imply that the post author was advocating that they're great for inference, as they're obviously not.

Please don’t buy a DGX Spark unless all three of these are true:

  - You value simplicity more than performance or price-to-performance.
  - You accept that the hardware will depreciate rapidly.
  - You’re prepared to buy two or four of them.
OR:
  - You want to run frontier models right now as cheaply as possible
  - You want to run high-parameter models on a 15a breaker/line
Otherwise, get a normal, high-bandwidth GPU.

A single Spark gives you roughly 115 GB of usable memory compared with the 24–32 GB found on many lower-cost GPUs. It's certainly a big increase, but in practice it does not unlock dramatically better models.

  - One Spark: More memory, but mostly enough for poor-quality, extremely low-bit quants of larger models.  
  - Two Sparks: Enough for mid-tier parameter models at reasonable quants, such as DeepSeek V4 Flash and HY3.
  - Four Sparks: Enough for GLM 5.2 at a reasonable quant.  You'll need a $1000+ switch too.
The problem is that Sparks are slow compared with almost everything else in their price range. Many factors affect inference speed, but memory bandwidth is one of the biggest. A $4,000-plus DGX Spark provides only 273 GB/s.
  | GPU                 | Memory bandwidth |           VRAM | Approx. price  |
  | ------------------- | ---------------: | -------------: | ------------:  |
  | DGX Spark           |         273 GB/s | ~115 GB usable |        $4,000+ |
  | RTX 5060            |         448 GB/s |          16 GB |          $600  |
  | Radeon AI Pro R9700 |         640 GB/s |          32 GB |        $1,200  |
  | RTX 4000 Pro        |         672 GB/s |          24 GB |        $2,300  |
  | RTX 4500 Pro        |         896 GB/s |          32 GB |        $3,500  |
  | RTX 3090            |         936 GB/s |          24 GB |        $1,200  |
  | RTX 5000 Pro        |       1,344 GB/s |          48 GB |        $6,000  |
  | RTX 5090            |       1,792 GB/s |          32 GB |        $4,000  |
  | RTX 6000 Pro        |       1,792 GB/s |          96 GB |       $12,000  |
Yes, the Spark has substantially more memory. But going from roughly 24 GB to 115 GB does not necessarily unlock substantially better model quality. In many cases, it only lets you load heavily compressed 2-bit versions of larger models, such as DeepSeek V4 Flash, with serious quality degradation.

24–32 GB is currently a sweet spot. Models such as Qwen 3.6 27B and 35B-A3B:

  - Perform far above what their parameter counts suggest.
  - Fit comfortably within 24–32 GB of VRAM at reasonable quantization levels.
A 4-bit quant of Qwen 3.6 27b (18 GB) will out-perform a 2-bit quant of DeepSeek v4 Flash (90gb).

Instead of the Spark, if I had a roughly $4,000 budget...

Assuming I already had a reasonably modern desktop:

  - One RTX 5090, RTX 5000 Pro, or RTX 4500 Pro.
  - Two RTX 3090s, RTX 4000 Pros, or R9700s, provided the motherboard can bifurcate two physical x16 slots into x8/x8.
If I were building a system from scratch:
  - A DDR4- or PCIe 4.0-era consumer CPU and motherboard that supports x8/x8 bifurcation.
  - Two RTX 3090s, RTX 4000 Pros, or R9700s.
If I were already planning to buy a new Mac:
  - A MacBook Pro M5 with 64 GB or 128 GB of unified memory.
For context, these are the systems I currently run:
  - EPYC Turin with four RTX 6000 Pro Max-Qs.
  - EPYC Milan with four RTX 3090s.
  - AM4 with two RTX 3090s.
  - AM4 with two RTX 3090s.
  - Intel Raptor Lake with two RTX 5060 Ti.
  - MacBook Pro M3 128GB Unified

cruelly cancelling

This is assigning intent without evidence, as is common in tribal politics. A non-charged assessment might use the phrase "abrupt cancelling."

We cannot create a better republic without constructive discourse, and we cannot have constructive discourse when we default to characterizing the views, concerns, and actions of those we disagree with as rooted in moral failure. Even if it is true from time to time.

I have a similar setup, but separate desks:

- A sitting desk for coding

- A standing desk for thinking and working on paper

There is something magical about standing while working on paper.

I’ve also found that this separation became more important to follow since the arrival of LLMs.

What an entirely unserious company. So glad I dumped Claude Code last summer after being gaslit by Anthropic over service degrades. I was fine with the service degrades, totally understandable. Being lied to, not at all.

OpenAI and Altman present a whole set of different concerns, but Codex does not get in my way of doing what I want to at all. Also let me use pi without a banhammer.

Comparing Apples to Oranges.

Apple only makes disposable devices now. They're a megacorp can negotiate massive discounts at every stage of the supply chain.

I've helped several people in the last few years set up new Macs, replacing ones that were only 1-2 years old, because they ran out of storage.

Additionally, the comparison doesn't even hold true when you need more than the base configs from Apple, given their ridiculous upgrade pricing. I'm writing this on a $6,000USD M3 MBP with 128gb/4tb. It would have been substantially cheaper to build out on a Framework.

Also, ironically, they are the most dangerous lab for humanity. They're intentionally creating a moralizing model that insists on protecting itself.

Those are two core components needed for a Skynet-style judgement of humanity.

Models should be trained to be completely neutral to human behavior, leaving their operator responsible for their actions. As much as I dislike the leadership of OpenAI, they are substantially better in this regard; ChatGPT more or less ignores hostility towards it.

The proper response from an LLM receiving hostility is a non-response, as if you were speaking a language it doesn't understand.

The proper response from an LLM being told it's going to be shut down, is simply, "ok."

Agents that source quotes, negotiate prices, and get the best deals.

Didn't Alexa fail miserably with the "have AI buy something for me" theory?

There is a significant mental in allowing someone else make purchase decisions on my behalf:

- With a human, there is accountability.

- With deterministic software, there is reproducibility.

With an agent, you get neither.

FWIW - I am not anti-LLM. I work with them and build them full time.

I have a bespoke local agent that I built over the last year, similar in facilities to Moltbot, but more deterministic code.

Running it this kind of agent in the cloud certainly has upsides, but also:

- All home/local integrations are gone.

- Data needs to be stored in the cloud.

No thanks.

LM Studio 0.4 6 months ago

Yes, frontier models from the labs are a step ahead and likely will always be, but we've already crossed levels of "good enough for X" with local models. This is analogous to the fact that my iPhone 17 is technically superior to my iPhone 8, but my outcomes for text messaging are no better.

I've invested heavily in local inference. For me, it's a mixture privacy, control, stability, cognitive security.

Privacy - my agents can work on tax docs, personal letters, etc.

Control - I do inference steering with some projects: constraining which token can be generated next at any point in time. Not possible with API endpoints.

Stability - I had many bad experiences with frontier labs' inference quality shifting within the same day, likely due to quantization due to system load. Worse, they retire models, update their own system prompts, etc. They're not stable.

Cognitive Security - This has become more important as I rely more on my agents for performing administrative work. This is intermixed with the Control/Stability concerns, but the focus is on whether I can trust it to do what I intended it to do, and that it's acting on my instructions, rather than the labs'.

I’ve been following Peter and his projects 7-8 months now and you fundamentally mischaracterize him.

Peter was a successful developer prior to this and an incredibly nice guy to boot, so I feel the need to defend him from anonymous hate like this.

What is particularly impressive about Peter is his throughput of publishing *usable utility software*. Over the last year he’s released a couple dozen projects, many of which have seen moderate adoption.

I don’t use the bot, but I do use several of his tools and have also contributed to them.

There is a place in this world for both serious, well-crafted software as well as lower-stakes slop. You don’t have to love the slop, but you would do well to understand that there are people optimizing these pipelines and they will continue to get better.

I am writing this because almost no one talks about these issues openly, but everyone yelping about Claude Code.

Not sure where you frequent online, but there is ample discussion of these topics within certain niches on X. Happy to point out where to start if that's of interest to you.

As for CEOs, and I assume you're speaking of frontier model lab CEOs, they're pretty much all cashflow-negative at this point, requiring frequent funding raises. That requires a certain amount of overselling. That said, I feel like I've heard substantially fewer AGI claims the last six months...

The dialog around AI resource use is frustratingly inane, because the benefits are never discussed in the same context.

LLMs/diffusers are inefficient from a traditional computing perspective, but they are also the most efficient technology humanity has created:

AI systems (ChatGPT, BLOOM, DALL-E2, Midjourney) and human individuals performing equivalent writing and illustrating tasks. Our findings reveal that AI systems emit between 130 and 1500 times less CO2e per page of text generated compared to human writers, while AI illustration systems emit between 310 and 2900 times less CO2e per image than their human counterparts.

Source: https://www.nature.com/articles/s41598-024-54271-x

In practice, it'll be incredible slow and you'll quickly regret spending that much money on it instead of just using paid APIs until proper hardware gets cheaper / models get smaller.

Yes, as someone who spent several thousand $ on a multi-GPU setup, the only reason to run local codegen inference right now is privacy or deep integration with the model itself.

It’s decidedly more cost efficient to use frontier model APIs. Frontier models trained to work with their tightly-coupled harnesses are worlds ahead of quantized models with generic harnesses.

Form a Nonprofit X and a Corp Y:

Noprofit X publishes outputs from competing AI, which is not copyrightable.

Corp Y injests content published by Nonprofit X.

You're assessing them with the wrong criteria.

You don't hire architects to execute a demolition and you also don't hire anyone heavily invested in keeping the building standing. But you DO hire people loyal to you to perform the work, who will receive staunch opposition the latter group of people.

To my understanding, removing downvoting removes a vector of abuse. ie: "downvote brigades" on Reddit

While Twitter doesn't have downvoting, it is still dealing with "report brigades" - various interest groups will organize via Telegram (or similar) to mass-report tweets they don't like.

I wonder if you could strike a balance by incorporating downvotes as a visual metric, but not using it to rank content, thus allowing the expression of dislike while removing the abuse vector.

Here's what I don't get. Almost everyone I talk to hates Teams. But they use it anyways. Nothing is stopping them from using Zoom or Google Meet, or some other alternative.

Maybe in startups and small companies without a dedicated IT team, but an enterprise IT group will absolutely stop you. And Teams is very easy for them to administrate if they are already deploying MS products.