HN user

dust42

709 karma
Posts0
Comments152
View on HN
No posts found.

With a M5 16c 48GB and Qwen 3.6 35B Q4 I get up to 1900 PP/s and 80 TG/s. With an Nvidia 5090 I get 7800 PP/s and 280 TG/s.

Together with pi mono I wouldn't want to go back to Claude & Co. Speed, quality of the answers, short answer times at any time of day - once you have eaten from the fruit your definition of SOTA will change...

For reference, I do software development since 30 years, I am not vibe coding the umpteenth todo list.

I just spent the last two weeks digging into workflow state engines and temporal was one of the candidates. It is a VC backed fork of Cadence. The got 0.3B funding and whatever positive I read about them on the net I take with a big spoon of salt. Just my 2 cents.

For many models the performance of llama.cpp on Mac is 20-40% lower than MLX. Did you try MLX? At least on HF there are MLX 2-bit quants. Unfortunately I have only 64GB, so I can't test it.

The output of any LLM is always 100% hallucination by principle. On top of that, most benchmarks are at best an approximation of LLM quality. Your use case decides which one to use. That said, I haven't tested v4 yet but the old 3.2 is still a decent model. And concerning use cases, I had coding problems that Opus couldn't solve but a local 35B model did.

All the talk about frontier and SOTA is do dig deeper and deeper into the pockets of VCs and finally do an IPO.

That website is so low effort that 2s is actually long to figure it out. Very sure that it is robot upvoted.

Edit: I live in the cheese triangle, France - Switzerland - Italy.

Ask on mikrocontroller.net. There will definitely be people who know who in Germany is doing it or maybe even offer to do it for you. One of the last old fashioned forums on the net. German speaking but nowadays that shouldn't be a problem.

I am using it with pi agent and I have stopped renting tokens. Much better for me than Claude Code, on M1 Max 64GB. This model with oMLX is at 16k context PP 919.9 tok/s and TG 54.7 tok/s. You have to manage the context but the better you manage context the more focused the output is. I use it without thinking.

I once worked for a company that bought a spot in the evening news (french TF1). It worked that way: a french minister was visiting a fair and coming to stop in front of the booth and getting a product demo. And that ended up in the evening news. Since that time I kind of watch the news with different eyes.

I really love it. The simplicity is key. The first play project I made with it was a public transport map with GTFS data - click on a stop and get the routes and the timetables for the stop and the surrounding ones. I used Qwen3.5-35B on Mac M1 Max with oMLX. It wrote 98% of the code with very little interaction from me. And very useful is the /tree feature to go back in history when the model is on a wrong track or my instructions where not good enough. I usually work in a two path approach: first let the model explore what it needs to fulfill the task and write it into CONTEXT.md (or any other name to your liking). Then restart the session with the CONTEXT.md. That way you are always nicely operating in 5-15k context, i.e. all is very fast. Create an account for pi (or docker) and make sure it can't walk into other directories - it has bash access. Add the browser-tools to the skills and load them when useful: https://github.com/badlogic/pi-skills

No need for database MCP, I use postgres and tell it to use psql.

Occasionally I use prettier to remove indentation - the LLM makes a lot less edit errors that way. Just add the indent back before you commit. Or tell pi to do it.

I've sold out 4 months ago

Context: Mario Zechner is the creator of the pi coding harness which powers OpenClaw. OpenClaw is made by Peter Steinberger, a friend of Mario Zechner. Armin is another friend who made public that OpenClaw is based on pi.

Pi itself is a minimalist coding harness with a tiny 1500 token system message and only read, edit and bash as tools. I only discovered it a few weeks ago and it is surprisingly powerful with a local Qwen3.5-35B - especially as it allows to keep the context low.

Mario's blog posts are not easily digestible (imho) until you have read a few of them but they have plenty of profound thinking. His blog is for me the first one in years where I have spent an hour to read several posts.

Mario is deeply rooted in the OSS system and basically that is what he is talking about here in this post. That said, I have no idea what earendil is doing, except that it is based on pi.

Edit: My personal take - "I've sold out" is very much Austrian style because actually it is the opposite. To quote one thing from the post:

"Then Miguel and Nat approached us. Long story short: we sold RoboVM to Xamarin. A short while later Xamarin closed-sourced our open-source RoboVM core, quickly followed by Xamarin selling to Microsoft. Then Microsoft shut down RoboVM immediately.

While there was some monetary gain, everything about this fucking sucked."

So Mario did a lot of vetting to hopefully avoid this from happening again.

Just to mention one thing, helium -which is a necessity for chip production- is a byproduct of LNG production. And 20% of that is just gone (Qatar) and the question is how long it will take to get that back. So not only a chip shortage because of AI buying chips in huge volumes but also because production will be hampered.

Tongue in cheek: we urgently need fusion power plants. For the AI and the helium.

You can spend every euro or dollar only once. If you consider CO2 emissions a critical problem, then you should spend every single dollar as efficiently as possible. Obviously independence of fossil fuels has a value too, as the current situation in the middle east shows.

It would make much more sense to import (renewable) electricity from Spain to Germany than strawberries.

I once read an article that in Berlin the sewage system is flushed with fresh water because too many people have installed water saving toilet flushers. So plenty of people bought these water savers and now the price of water has gone up because the water that is directly flushed needs to be paid too.

The 'balcony power stations' are the same thing. They get subsidised, and you even get a fixed kWh price when pushing into the grid.

The problem is that in the end it will become more expensive for everybody because at times you have a surplus driving the whole sale electricity prices into the negative while still paying fixed prices for injection into the grid.

To make this economically viable, you have to have everyone paying spot prices. Everything else is just green ideology driven inefficiency.

Just to make it clear, I think renewables are an important option for the future. But to make them a viable option of the electricity energy mix, supply and demand, storage and grid capacity need to be taken into account.

Last not least, there is plenty of low hanging fruit to drive CO2 emissions down: drive up the truck tolls. Currently you have potatoes farmed in Germany, driven to Poland to get washed, transported to Italy to be converted to french fries and transferred back to Germany into the super markets.

Same goes for home office, during Covid it was possible for many workers to continue with their work. Does an accountant need to drive to an office every day? Nope. How many business trips could be replaced by a video call?

If the CO2 emissions problem is to be solved rather sooner than later, the money has to be spend efficiently as there isn't enough of it.

I really think this is a security disaster waiting to happen, landing right in time for all the agentic terminal apps:

  printf '\e]8;;http://evil.com\e\\https://good.com\e]8;;\e\\\n'
The next step would be to embedd a full javascript VM in the terminal and a CSS engine.

I have to say correctly so. It is a case of "ROMANES EUNT DOMUS" [1]. What is "Lick eggs Merz" supposed to mean? In order to be a proper political message on a proper demonstration of proper school kids, it should at least say who's eggs Merz should lick and why. And clearly hint why that is either good or bad. Which would give Merz the opportunity to reflect and change his behaviour.

[1] From the Life of Brian

Sounds very plausible to me too. Because even if you refocus the business unit it makes no sense to lay off a highly capable team. Finding new people, integrating them into the team - all that costs a lot of time and money and there is no guarantee for success.

Definitely plenty of people further up the corporate ladder were not happy with the success, while the top is likely too far disconnected to understand.