My takeaway from this was that Claude was being completely useless as a helper Thanks to all the safety safeguards and that nonsense.
HN user
vardalab
Yeah, but maybe Kimi doesn't pretend to know better than I do. Like, I just had a task today where I had a silly form where I wanted to copy in a signature from one document to another document. And, you know, I could have just used some other dumb tools, but I decided, hey, let me us ask Codex to do it. But it started lecturing me on how we can't do this kind of stuff. And it was just pure silliness. So hopefully Kimi Work doesn't try to pretend that it knows better and lets me decide on ethics.
This blog post here summed it up the best. These are tools, so let them be tools.
https://geohot.github.io/blog/jekyll/update/2026/07/11/ai-20...
One thing I find that's missing a lot or at least I haven't come across other than commercial offerings like AquaVoice is a decent injected technical vocabulary so that the initial transcript requires minimum cleanup afterwards. Because I mostly use these tools to essentially ramble at the command line with coding agents. So there's a lot of technical terms that don't translate well. Like OpenBao comes out as open bowel sometimes,lol. That necessitates significant cleanup prompt or background text available to the cleanup llm, usually in the form of screenshot or something that gets converted to text but that in turn requires good hw for speed to be almost imperceptible. For example m5 max turns cleanup into a noticeable delay while 5090 is decent.
Only way I have found that's relatively easy to inject technical vocab is to use whisper, but limited, I think to about 220 or so tokens. Whisper has sort of like a priming prompt where one can put in a bunch of technical words and it will try to recognize those. But again, that's limited to small number tokens. And that limits one use a relatively slow, by today's standards, whisper.cpp.
I benchmarked it across a bunch of different hardware that I have available, and Whisper gives decent performance as far as speed goes only on a pretty top-end GPU, such as a 5090 or 4070, like for example on Strix Halo, it's still relatively slow for longer transcriptions because I prefer just a stream of consciousness ramblings for minutes and then that being transcribed and cleaned up versus short sentences. So in that scenario something like 5090 really is good because the cleanup prompt runs fast using usually Qwen 3.6 MOE model. Whisper on 4070 itself is about 0.7 seconds for two or three minute transcription. So the total wait time for a three-minute transcription is roughly a second, or a little bit more than a second, so totally acceptable. But it does take decent hardware, and it grows to be double that on if running totally local. Well, in my case, it's all local, but it's my own hardware all over the place, but truly running on laptop, it's much faster using Parakeet, but then the cleanup is the bottleneck.
Anyway, it's just my experience messing around with this for the last year. I did start using AquaVoice, but their speed was exceptional, and tech vocab was exceptional, but they would have some annoying delays occasionally, and I didn't like paying the money and sending sensitive topics and screenshots into the cloud, and I had hardware, so my local solution is basically almost as good as commercial one. But I think they train their own model. So what I'm doing is I collect all the samples of my transcriptions, and I am slowly building my own data set that hopefully at some point when I get energy I will find some way to fine tune something.
now it's back? i saw it was gone from the usage but then reappeared about 10 min later, lol
A lot of stuff that has to do with VLLM and troubleshooting and compiling and building VLLM, compiling kernel, or just dealing with setting up eval for local models, it punts to Opus 4.8 on a regular basis. To the point that I have given up on using it for that purpose.
I don't know what your evals are, but you need to reevaluate them.
It had a tough time updating today. Or this evening. It just wouldn't update. It actually just freaking disappeared from my MacBook. It took some googling and downloading and multiple tries to get it back and working. Because they also combine on a MacBook Codex with ChatGPT app. I guess codex became ChatGPT app or some silliness like that.
Fable was refusing to patch vllm for me when trying to get mtp to work on r9700 gpus. Kept on bumping down to opus. Tried to really sanitize my prompts and everything but it seemed intrinsically prohibited from doing this sort of work. I guess it’s useful for making inane one shot games and websites, lol.
Had it happen to my PBS backups running on ZFS without ECC just the other day. Turned out to be heat related, memory was fine but during long backups heat was causing bit flips.
That's what I did after my 15 year old Synology DS1010+ built in USB DOM failed. Put a network booted debian on it via netboot.xyz with zfs and now i get to reuse those 15 year old 2TB still chugging along disks. Fully open source, pretty nice way to keep old hardware going. It's my tertiary backup that wakes up once a week pulls in some open weights LLM models that I hoard just in case and goes back to sleep.
https://github.com/SeraphimSerapis/tool-eval-bench Trying to run this stuff really triggers it. Freaking frustrating. I have it set up some local inference and then I'm struggling to get the MTP working and it just refuses to work on evaluations.
Making hay was not an easy task. A lot of work and stress as well.
The stickers are because somebody got hurt. That's the only reason why they're there. And hopefully it helps somebody. It doesn't hurt anything to have the sticker, lol.
Scene is in Boston Public Garden
Which country do you live in that you have such a positive outlook about redistribution of wealth?
better prompt processing like 1.5x+ and more kv but tg most likely lower like 0.8x or so but I am just going by memory for Qwen3.5 without mtp.
Because it's a tax I think on second properties.
You forgot to include the actual living costs if you invest. You're not gonna be able to contribute $2,500 a month. You would be able to contribute not that much , around here rents are $2,500 a month.
Exactly. Here I am sitting talking to my freaking computer, arguing with it, whatnot. And people just dismiss it as if it's not a science fiction. We were not there two, three years ago. Now we are. It's amazing and scary, scary mostly because the society that we operate in. I bet it's less scary in Norway or elsewhere where govt is more biased towards people not corps.
Not all workers will fare equally. As an illustration, Reuters cited a union source estimating that someone in the memory chip unit earning an 80-million-won base salary could take home roughly 626 million won in total bonuses this year. By comparison, workers at SK Hynix stand to collect upward of 700 million won should their employer post annual profit of 250 trillion won, Reuters calculated. Unlike at Samsung, SK Hynix employees are not limited to stock payouts and may instead opt for cash, Reuters reported.
Almost 6x the base, not bad.
Exactly, the whole point of having a laptop is that you can close it at some point. Once you move to these newfangled workflows it is all gone. That is why I have been slow on uptake with all these new apps they are releasing like Codex GUI and Claude GUI and all that stuff. I like that work continues when I close my laptop, but also I like that it continues on my hardware, not somewhere in the cloud being a supply chain attacked by who knows what.
But I did.
Instead of discussing Kierkegaard or universality of human condition we are discussing em-dashes. Peak HN.
Yeah, this all sounds good in theory, but there's a lot of edge cases. For example, for me switching between Mac and Linux what often happens is that Linux just for some reason the it turns off the monitor port and it's black until I reboot and there was no easy way to get it back. This is while using fancy Dell monitors built-in KVM. Ultimately, I have settled on remote desktops as a more viable and quicker option.
It flakes out in less than 24 hrs. I tried leaving a session open on remote control mode in a VM but it inevitably stopped with some token auth error.
Why would you prefer less coherent article? If article has a utility, I will read it, no matter what the source is.
I disagree. The rigidity of YAML and stuff like that is what actually makes LLMs work better. I have strict linting rules and file size limits and it imposes discipline on LLMs. That's why it worked even last year. Even before Opus 4.0 it worked to some extent as long as you imposed discipline on these models Trust me, I do pretty complicated things with Ansible, key thing is to have decent established patterns and these models truly are getting better.
I like the forums of the old way better than what we have now, Discord and reddit suck I mean, even back button on reddit does not work 50% of the time, lol. Hacker News is the only thing that we have left that's similar to what we had from forums of the old. And even then, I think I like those forums better. Remember Deja News before Google bought them. Shit like that, that was good.
I think it's as simple as Citizens United and wealth inequality exploding.
Yeah, but I have Claude Code or Codex do this Ansible stuff and they do just fine with all this and then there's a gazillion of examples that they can lean on and once the patterns are established, it's pretty smooth. Opus 4.5 was when the big inflection was I was heavy into automation all summer. It was Opus 4.0. It was like pulling teeth. And then when 4.5 came out, it was just beautiful.