Yes, this is frustrating, but it doesn’t occur in CC. I run the conversation logs through an agent and opencode source, and it identified an issue in the reasoning implementation of opencode for Zai models. Consequently, I ceased my research and opted to use CC instead.
HN user
nkko
@nibzard on X
What happened in 2022?
This wqw an interesting postmortem for me. As a regular user, February definitely felt a bit shakier than usual.
I ran into the unicorn error page a few times and had intermittent issues with PR pages and bunch of weirdness. Nothing long-lived, but enough transient failures that my overall feel was it is degrading in general.
was listening to music while coming with groceries and simultaneously juggling stuff to open the doors and change the track with Siri (the only use for Siri I have)
FWIW I work at Steel (not the OP). While we’ve been iterating on the “right shape” for agent tooling, I’ve been building a benchmark harness to measure how different surfaces affect real web task completion: raw API context, CLI-only, opinionated “skills” (structured outputs + artifact capture), and combinations.
If you’ve run agents on the open web, I’d love suggestions for nasty-but-representative workflows to include in the benchmark.
This rings true, as I’ve noticed that with every new model update, I’m leaving behind full workflows I’ve built. The article is really great, and I do admire the system, even if it is overengineered in places, but it already reads like last quarter’s workflow. Now letting Codex 5.3 xhigh chug for 30 minutes on my super long dictated prompt seems to do the trick. And I’m hearing 5.4 is meaningfully better model. Also for fully autonomous scaffolding of new projects towards the first prototype I have my own version of a very simple Ralph loop that gets feed gpt-pro super spec file.
For no special reason, beside I could, I’ve slop coded this AI agents ephemeral VM orchestrator which I use inside any agent to manipulate and maintain my coding VMs on Proxmox. Probably it could make sense to simplify it further and move from Proxmox to something like this. Link: https://github.com/nibzard/agentlab
This is exciting. But I had to read and check everything twice to figure it out, as some already commented. Strong Feedback loop is an ultimate unlock for AI agents and having twins is exactly the right approach.
For sure! Just ask enough times "why" and you will find the root. The main issue here it is, how many people do that for real, and how this is becoming even more critical now.
That magic now moved to ESP32.
Reproducibility is a fascinating topic for me, and today with AI coding agents we could have automated reproducibility at least in some fields. The concept they touch on in the paper, of post publication verification could replace or add onto existing research valorization.
Annual full body MRI has become a trend. Not sure who first started promoting it, probably Peter Attia.
That’s a good idea, will have to think a bit on how to implement it.
Fast reading was always my Achilles heel ;)
Yep—many of these predate LLMs.
Strong point. I’m considering to tag patterns better and add stuff like “model/toolchain-specific,” and something like “last validated (month/year)” field. Things change fast and for example “Context anxiety” is likely less relevant and should be reframed that way (or retired).
Author here (nibzard). I started this back in May as a personal learning log. I agree with the skepticism about jargon and novelty. However, if something reads like overly complex common sense, that’s a bug, and I’d like to fix it. If you can point out 1–2 specific pages that feel sloppy or unactionable, I’ll rewrite them (or remove them). I’m also happy to add flags or improve the structure. Also, contributing new patterns would be grand. Of course, some or even all patterns are explicitly “emerging.”
At some point, we need to begin. My initial thought was that this is a growing and evolving resource, primarily for my own use. We are slowly but steadily learning what makes sense annd patterns emerge. Also, if others find it interesting and contribute, that would be even better.
I’m eager to tackle issues and PRs.
This type of things are time sink holes, as surprisingly it takes a lot of time to figure everything. I was hoping to dedicate a decent amount of time to review and structuring but sadly life got in the way. If you have a suggestion how to structure it better I am all ears.
Most of the patterns should link to external resources since they were derived from them. If there’s no link, it was probably obvious or I’ve derived it from my own project.
The flow was, me finding interesting pattern -> Claude ingesting the reference and putting it in a template -> Me figuring out if it makes sense -> push
Author here. Yes, CC is the maintainer. When I stumble on a decent idea, I would just feed it to CC to create a pattern out of it. This was my quick and dirty approach to a public learning log with an idea that I would get back to it at some point and clean it up. Which I did on a few occasions.
Yes!
See my comment above. The repository is from May when I was intensely exploring everything agentic. I used it as a public bookmarking tool and also in the hope of receiving contributions. Thanks to this HN share, I received four PRs.
Hi, author here. Honestly, I just used this as a bookmarking place for myself. Which you could infer if you go through some patterns. I’ve created a flow with CC where I would just dump a new source like a podcast, post, or whatever to have it for reference.
Got tangled up in human/AI worlds when writing, fixed it.
Treat it like a wind-tunnel test, not a wisdom oracle: the value isn’t profundity, it’s seeing what kinds of coordination and failure modes emerge when the only continuity is the record and the collaboration is the environment.
This is a very smart idea. Life orienteering, we are all running in life and some predefined checkpoints would be nice.
Beyond graduating students, I see model labs as “accelerators/incubators” bundling, launching, and productizing observed ideas that gain traction. The sheer strength of their platforms, the number of eyes watching them, near-zero marginal costs, and seemingly unlimited budgets mean that only slow decision-making can prevent them from becoming the next Amazons of everything.