Many are investing in RTX6000pro hardware but hitting deployment walls. This guide covers everything from simple 4x builds to scaling up to 16 cards with PCIe switches.
HN user
chisleu
An open-source, code-first Go toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.
Total tangent, but I got to ride in some of these on a recent trip to India and I was really impressed with the build quality and utilitarian usefulness of the design.
Here is the demo video on it. The video w/ sound input -> sound output while doing translation from the video to another language was the most impressive display I've seen yet.
Because of the prompt processing speed, small models like Qwen 3 coder 30b a3b are the sweet spot for mac platform right now. Which means a 32 or 64GB mac is all you need to use Cline or your favorite agent locally.
I've been using GLM 4.5 and GLM 4.5 Air for a while now. The Air model is light enough to run on a macbook pro and is useful for Cline. I can run the full GLM model on my Mac Studio, but the TPS is so slow that it's only useful for chatting. So I hooked up with openrouter to try but didn't have the same success. Any of the open weight models I try with open router give sub standard results. I get better results from Qwen 3 coder 30b a3b locally than I get from Qwen 3 Coder 480b through open router.
I'm really concerned that some of the providers are using quantized versions of the models so they can run more models per card and larger batches of inference.
This is going to improve the quality of LLM responses for users. I'm for this.
but eventually they did it
They can do this with manual partitioning indeed. I've done it before, but it's not ideal because the auto partitioner will scale beyond almost anything AWS will give you with manual partitioning unless you have 24/7 workloads.
you can be throttled for more than a day before it kicks in
I expect that this would depends on your use case. If you are dropping content you need to scale out to tons of readers, that is absolutely the case. If you are dropping tons of content with well distributed reads, then the auto partitioner is The Way.
and indeed the bucket is not separate from the object key. the API separates it logically "for humans" but it's all one big string
You don’t have to randomize the first part of your object keys to ensure they get spread around and avoid hotspots.
As of when? According to internal support, this is still required as of 1.5 years ago.
/agree
We are in the infancy of LLM technology.
How was he doing "complex agentic coding" when the APIs have such extreme context and throughput limitations?
holy shit it does. The scene with him inventing the new compression algorithm basically foreshadowed the gooning to follow local LLM availability.
I use opus or gemini 2.5 pro for plan mode and sonnet for act mode in Cline. https://cline.bot
It's my experience that Opus is better at solving architectural challenges where sonnet struggles.
It looks like qwen3-coder is going to steal K2's thunder in terms of agentic coding use.
It's 480B params, not 480GB. The 4 bit version of this is 270GB. I believe it's trained at bf16, so you need over a TB of memory to operate the model at bf16. No one should be trying to replace claude with a quantized 8 bit or 4 bit model. It's simply not possible. Also, this model isn't going to be as versed as Claude at certain libraries and languages. I have something written entirely my claude which uses the Fyne library extensively in golang for UI. Claude knows it inside and out as it's all vibe coded, but the 4 bit Qwen3 coder just hallucinated functions and parameters that don't exist because it wasn't willing to admit it didn't know what it was doing. Definitely don't judge a model by it's quant is all I'm saying.
A Mac Studio 512GB can run it in 4bit quantization. I'm excited to see unsloth dynamic quants for this today.
I tried using the "fp8" model through hyperbolic but I question if it was even that model. It was basically useless through hyperbolic.
I downloaded the 4bit quant to my mac studio 512GB. 7-8 minutes until first tokens with a big Cline prompt for it to chew on. Performance is exceptional. It nailed all the tool calls, loaded my memory bank, and reasoned about a golang code base well enough to write a blog post on the topic: https://convergence.ninja/post/blogs/000016-ForeverFantasyFr...
Writing blog posts is one of the tests I use for these models. It is a very involved process including a Q&A phase, drafting phase, approval, and deployment. The filenames follow a certain pattern. The file has to be uploaded to s3 in a certain location to trigger the deployment. It's a complex custom task that I automated.
Even the 4bit model was capable of this, but was incapable of actually working on my code, prefering to halucinate methods that would be convenient rather than admitting it didn't know what it was doing. This is the 4 bit "lobotomized" model though. I'm excited to see how it performs at full power.
A mac studio can run it at 4bit. Maybe at 6 bit.
I love the interface. It makes it extremely easy to rewind time to undo code edits and rewinding the LLM context at the same time. It's prompting and toolset is great. It's got MCP which I have integrated into my workflow. It's got a solid marketplace of auto installing MCP services. I love it.
Like working with an incredibly talented and knowledgable junior engineer, but still a junior engineer.
If you want to try something better than claude code, try Cline.
It's infinitely useful for people who's workflows involve LLM agents.
It is indeed. I don't use Claude Code. I use Cline which is a VS Code extension (cline.bot).
This is a pretty killer feature that I would expect to find in all the coding agents soon.
Yup, slowing down the AI is a really hard thing to do. I've mostly accomplished it, but I use extensive auto prompting and a large memory bank. All of it is designed explicitly to slow down the AI. I've taught it how to do what I call "Baby Steps", which is defined as: "The smallest possible change that still effectively moves the technology forward." Some of my prompting is explicit about human review and approval of every change including manual testing of the application in question BEFORE the model moves on to the next step.
Yeah except it's not hearsay so much as repeating a credible source.
Dr Drew and Adam Corola had a show on MTV where they discussed it at length.
I don't know brother. I'm not making that claim. I'm repeating it.
Same. After enough use, I just kind of forgot how to have fun without it. Sleep issues are the worst.
Oh that’s absolutely part of it.
Heard it from the doctor on mtv. Remember that show?