I've been calling LLMs superhuman at writing `jq`. It's like you're talking directly with the JSON.
HN user
palcu
Member of Technical Staff, AI Reliability Engineering (AIRE) at Anthropic
Previously: SRE at Google Cloud (taking care of GCE and GCP) and Forward Deployed Engineer at Palantir.
https://www.palcu.net/
[ my public key: https://keybase.io/palcu; my proof: https://keybase.io/palcu/sigs/425bZ1Ip-RwE52Uws06rLF54qq21RU2pQQyi8cXICv8 ]
Yes, the general trend is the unprecedented growth that we've seen. Typically one would have some time in advance to re-engineer the systems to support the increased in traffic and users. But we're dealing with very compressed timelines and while most of the time we're able to fix the issues beforehand, sometimes we have to do them in production. Sorry for that.
Hey folks, I'm Alex from the reliability engineering team at Anthropic. We've just posted the retrospective for this incident:
On March 26–27, 2026, customers experienced elevated error rates when using Claude Opus 4.6 and Claude Sonnet 4.6. The issue was caused by a networking performance degradation within our cloud infrastructure that disrupted communication between components of our serving stack. We resolved the incident by migrating the affected workloads to healthy infrastructure, restoring normal service by 9:30 AM PT on March 27.
Hey folks, I'm Palcu from the reliability engineering team at Anthropic. I just posted a small retro on the status page:
Between 14:17 and 17:11 UTC, our primary application database experienced severely degraded I/O performance following a routine maintenance operation, causing slow or failed requests on Claude.ai and preventing new or refreshed sign-ins for Claude Code and the Console. API traffic via Claude Developer Platform was unaffected.
https://status.claude.com/incidents/jm3b4jjy2jrt
Sorry again, thanks for bearing with us as we're dealing with the influx of new users and scaling up all of our systems.
Hey folks, I’m Alex from the reliability team at Anthropic. We’re sorry for the downtime and we’ve posted a mini retrospective on our status page. We’re also be doing a more in depth retrospective in the following days.
Thank you! Opening an incident as soon as user impact begins is one of those instincts you develop after handling major incidents for years as an SRE at Google, and now at Anthropic.
I was also fortunate to be using Claude at that exact moment (for personal reasons), which meant I could immediately see the severity of the outage.
Hello, I'm one of the engineers who worked on the incident. We have mitigated the incident as of 14:43 PT / 22:43 UTC. Sorry for the trouble.
There is an external incident now.
https://status.cloud.google.com/incidents/8cY8jdUpEGGbsSMSQk...
I confirm. I’m still doing reliability, but in a fun and exciting way for Claude and Anthropic.
One of the problems has been that most users have requested that quotas get updated as fast as possible and that they should be consistent across regions, even for global quotas. As such people have been prioritising user experience rather than availability.
I hope the pendulum swings the other way around now in the discussion.
[disclaimer that I worked as a GCP SRE for a long time, but not left recently]
seems to be OpenAI
And the queries flow through.
My adjacent teams in London who work in SRE on Google Cloud (GCE) got some well deserved doughnuts today for rolling out the patches on time.
Yeah yeah, they've broken the old TweetDeck. You need to wait for the pop-up to ask you to transition to the new TweetDeck. Or, search on the internet for the Javascript variable you have to change in your console.
The more important question is that they've removed the Activity feed, where you could see likes from other people. Which was like a realtime feed to what your friends were doing on the website. The website is way more boring now.
Life and experience, if you're looking for a short answer. For example, last year we had an outage in London[0] and the folks who worked on it learnt a lot. Now, they applied the learnings in this incident.
There's not much emotion as the core team working on the huge outages is more like an "SRE for SRE". They are all people who've been with the company for a long time and they've been in the secondary seat for at least one previous big rodeo. Not to mention that we're all running a checklist that has been exercised multiple times and there's always somebody on the call who could help if a step fails.
Personally, I wasn't part this time for the actual mitigation of the overall Paris DC recovery, as I was busy with an unfortunate[0] side effect of the outage. These generate more anxiety, as being woken up at 6am and being told that nobody understands exactly why the system is acting this way is not great. But then again, we're trained for this situation and there are always at least several ways of fixing the issue.
Finally, it's worth repeating that incident management is just a part of the SRE job and after several years I've understood that it is not the most important one. The best SREs I know are not great when it comes to a huge incident. But, they're work has avoided the other 99 outages that could have appeared on the front page of Hacker News.
[disclaimer: SRE @ Google, I was involved with the incident, obvious conflicts of interest]
Hey Dang, thanks for cleaning up the thread. One thing to note is that the title is not correct. The entire region is not currently down, as the regional impact was mitigated as of 06:39 PDT, per the support dashboard (though I think it was earlier). The impact is currently zonal (europe-west9-a), so having zone in the title as opposed to region would reflect reality closer.
Finally, there's lots of good feedback on this thread and on the previous one (https://news.ycombinator.com/item?id=35711349), so we obviously have a lot of lessons to learn.
I absolutely love the answer to any of the other obvious questions that one would have about the membership.
Not ironic, just nominative determinism.
Hijacking the article, but has somebody managed to find a good iPhone/iPad keyboard for coding? I’m still able to do Python with the default keyboard, but I press a lot of times the symbol key.
The Wolfram Alpha custom keyboard on Android is absolutely the best keyboard for coding.
I tried getting into the Crowdcube investment round, but after they got the allocations, the whole deal fell through and I was refunded.
Needless to say I’m pretty happy they didn’t take my money in the end.
Rest of world has a really good article about people from emerging markets that chose Luna because their home currency is unstable.
Took it to the office today and I must say, this is a beautiful piece of engineering. Cheers to all the people that made it possible and here's for hundreds of years of continuous service.
I was just imagining crowds of people in Trafalgar square singing "it's coming home" and it's not about football for once. Maybe next year...
For the nerds, there you can read about Euro English on Wikipedia.
I think the DVLA is pretty unique in how dysfunctional it is as a British agency. I think the move to Swansea was very poorly executed. But you cannot criticise it because Very Serious People have said that devolution is always good.
You are forbidden to trade derivatives of GOOG as an employee of Alphabet. Source is that I am one and I have read my insider trading policy.
This Lesswrong thread has a good discussion about the subject.
https://www.lesswrong.com/posts/niQ3heWwF6SydhS7R/making-vac...
In the same category, if you live in the UK then http://monevator.com is the go to source for passive investing advice. The comments section is a gold mine as well.
It’s highly improbable. Medical recovering I assume it also includes being discharged from the hospital.