Scanning mailbox, reading and classifying emails. Scanning a knowledge base, reading and improving individual articles, reading support interactions and creating summaries, checking and researching new leads that signed up. So much that is possible.
HN user
iagooar
I build my own computers and tools. Because I can.
Founder https://www.podigee.com/en
On paper the M4 should be roughly 1/3 of the M5, in practice it is only 1/2. With the right, optimized model like qwen3.6 35B MoE MLX you can get over 40 tok / sec on it. I run dozens of background jobs that are not time-critical on it.
I am not going to flag you, I am much OK with having good arguments.
I just purchased a Mac Mini M4 Pro 64GB for $3k - 2nd hand of course.
I am not a hater of Nvidia and I am planning on building a workstation based on RTX cards. You clearly do not seem to understand how convenient the MacMini actually IS - the form factor, how quiet it is, how durable it is, how well it integrates with other Macs, how well it works as a bridge to a personal agent like Hermes (integration with iMessage, Calendar, Reminders, iCloud, etc).
I am pretty sure I know a thing or two about computing, I have been in the trenches for many, many years and I have had machines of all kinds, shapes and colors. It just so happens that Macs are very capable, very convenient machines that happen to work great in the era of LLMs, too.
But you do you.
Buy a refurished or 2nd hand one.
I disagree LAN connection is the bottleneck. I do even work with it remotely via Tailscale on shaky hotel WIFI and it works fine (or as fine as any other API-based model).
Just buy a Mac Mini really is good advice if you want to get into real, always-on convenient agentic work.
Soon it is going to be good even for coding using local LLMs. Until then, just run API models on it for coding, local LLMs for "knowledge" work or daily driver agent like Hermes.
I thought they might ship an M5 Max version, but you are probably right.
qwen3.6 27B MLX 8bit -> 15 tok / sec. A bit slow but it is a delightful model to use, and smart too.
qwen3.6 35B A3B MLX 8bit -> 85-90 tok / sec! It is impressively fast and roughly 90% as good as 27B (in my opinion).
My problem is I won't accept anything lower than the 96GB the RTX Pro 6000 Blackwell has. My dream is a workstation with 2x Pro 6000 to run DeepSeek v4 Flash comfortably, possibly qwen 3.6 / ornith on turbo speed.
But man, I have never purchased a computer which is more expensive than a decent family car.
Ballpark 25-30 tok / sec on the Mac Mini Pro M4 + qwen3.6 35B. The generation itself is good, prefill is known to be slow on any Apple M-chip architecture. It is really decent.
M5 Max. But I also have a MacMini M4 Pro 64GB. Qwen3.6 runs on the M4 just fine - sure the M5 is at least 2x the speed. If Apple launches a MacMini with an M5, I will be the 1st one to get it.
Get a 2nd hand one. I was lucky enough to get a new one first, last week I get a 2nd hand one in order to run one of my Hermes minions at work.
I love my MacBook Pro M5 128GB RAM and I love qwen3.6.
BUT DO NOT buy this MacBook if you plan on doing serious coding using local LLMs with it. The reason is simple: your fingers will burn and your head will explode from the noise.
Running any kind of sophisticated job on the very laptop you are using is just not viable. Sure you can use it in clamshell mode, but forget touching it while working with AI coding or agents.
If you want to run Qwen3.6 27B / 35B at its best, get a MacMini M4 with 64GB of RAM and put it in the basement - or at least a few meters from your desk. Connect to it over LAN or Tailscale. The MacMini will also cost you almost 1/3 of the MacBook Pro.
Thank me later.
Then it probably was wishful thinking on my side...
I guess people are tired of each instance of an Electron-based app using 1GB+ of RAM.
I have noticed that Opus and GPT 5.5 are very good at adjusting their thinking / reasoning intensity depending on the task at hand, something the open weights models are still not as good at.
In addition to that, some of the open weights models like GLM 5.2 or DeepSeek v4 Pro tend to be MUCH slower when generating tokens, which contributes to the perceived slowness. Although I wouldn't call models like GLM 5.2 slow by any means, e.g. it is currently one of the fastest models inside Notion today.
It is a combination of Hermes agent as the orchestrator and a custom extractor script (that uses qwen or any other LLM) that runs every 2 mins (on Mac via launchd). I had the code + skill written by Hermes.
The beauty of it is that Hermes itself has a cron too - every 4h Hermes will wake up and check if the email ingestion is working fine. If not, it will fix it. Funnily this is one of the most robust setups I've seen in a while, it is like having your own little DevOps waking up at night, fixing the infra when needed.
Nice. Are you working on it for "fun" or as a lab? I am looking into building a semi-professional cluster + setting up a lab, but the investment needed is beyond what I could justify as a hobby.
Yes, Brave search is one of these services I highly recommend paying for, the search they provide (similar to Exa, Tavily) is what makes an "OK LLM" become super smart.
10 years worth of Claude Max today. Also - Anthropic recently removed a model I relied on and isn't giving it back. As a non-US citizen, I would rather pay in advance but be sure, I will keep having access to inference on my own terms.
Also, it will just be faster - and more fun too.
I love running two models locally: qwen3.6 27B 8bit (dense) and qwen3.6 35B 4bit (MoE).
The 27B is the smarter, more reliable one - but it is slower. The 35B is faster, still very smart but below 27B, a bit less reliable. The reason is the MoE - Mixture of Experts architecture, which only activates a subset of parameters, making the model much much faster.
I run the 27B on a MacBook Pro M5 Max + 40 GPU cores + 128GB RAM (well, on this beast I can have 27B + 35B in memory at the same time with headroom for all the other stuff). But because this is a laptop, it is not possible to run local LLMs all the time - it just gets too hot and too loud.
What excites me more: I run the 35B model on a MacMini M4 with 64GB RAM. It is fast, it gets a lot of work done (e.g. it scans, extracts and classifies my emails, it watches the mailbox all the time and does work). I also use it as my private Hermes assistant ("when is the next Starship launch?", "who is playing today at the World Cup? Give me some trivia").
Next step I am planning is a RTX Pro 6000 Blackwell workstation I can put in my basement. I want to run qwen really fast, with multiple threads / prompts / agents at once. And MAYBE if the budget allows, a 2x RTX Pro 6000 setup in order to run DeepSeek v4 flash on it (to run research on it).
I run the exact same model, on the exact same hardware - amazing results. Pair it with good search skills (Tavily, Brave, Exa) and you have a near-SOTA model on your desk.
I recommend MacBook M5 Max with 128 GB of RAM to run it comfortably and fast. If you have something like a regular M4, go with qwen3.6-35b-a3d - the Mixture of Expert architecture makes it run 2-3x faster than the 27b version.
Out of curiosity, what are you training with these cards?
I have created this in 3 different ways, also with e2b.dev (great service). And yeah, that is my problem - I spend 99% of time hacking cool stuff, 1% yapping about it. Should do more yapping.
From a European perspective, based on the assumption API inference cannot be trusted anymore -> it means investing in local inference + building harnesses that can squeeze out all the power from the best open weights models.
Besides the language being Latin vs local languages, there is one huge difference people don't know about. The Tridentine Mass has the priest facing toward the altar and the tabernacle, this is called "ad orientem". In "modern" day post-Second-Vatican-Council mass, the priest typically speaks the local language and faces the congregation.
Yes and no. The LLM that sees a JSON structure can decide to use tools to extract and format data as needed, whereas it cannot do the same with natural language.
The Unix philosophy of small, composable tools is still valid in the era of stochastic machines!
What if "the thing" is a human and another human validating the output. Is that its own output (= that of a human) or not? Doesn't this apply to LLMs - you do not review the code within the same session that you used to generate the code?
Polish has been written with Latin alphabet since the 13th century. And before it simply wasn't written.
Polish works with the Latin alphabet just fine.
"Do kraju tego, gdzie kruszynę chleba podnoszą z ziemi przez uszanowanie dla darów Nieba.... Tęskno mi, Panie..."
"Mimozami jesień się zaczyna, złotawa, krucha i miła. To ty, to ty jesteś ta dziewczyna, która do mnie na ulicę wychodziła."