Major lesson from when I worked on Google Search indexing is that queues have a lot of hidden complexity and can make your outages much longer than they need to be. We had a big project to get rid of a bunch of queues by just scaling up our synchronous backends and making them faster.
HN user
nosefrog
Would be nice if the timeline matched up with the text of the blog post (missing "HackerOne provides disclosure guidance").
Interesting. We run pgbouncer via kubernetes so it was straightforward to make multiple pgbouncer processes on one machine. Also straightforward to get them running on multiple machines, which helps because we run on Azure and they like to cause rolling outages across our fleet via VM maintenance...
Don't reboot the db during your next outage.
It's highly unlikely that GCP banning their account without telling them is true, but GCP is probably not going to go public with the real reason.
Yup...
We run 1000s of machines in Azure. It's garbage. Very few features work. Nodes are always having strange issues, especially on the networking side. And the worst part is that Azure support has 0 interest in actually debugging things. We just got out of an outage today caused by the insanely slow SSDs that they attach to their postgres dbs by default.
At some point, you will have many teams. And one of them _will not_ be able to validate and accept some upgrade. Maybe a regression causes something only they use to break. Now the entire org is held hostage by the version needs of one team. Yes, this happens at slightly larger orgs. I've seen it many times.
The alternative of every service being on their own version of libraries and never updating is worse.
Big shoutout to Amin and Sacha for keeping this going!
Front Door is not good.
What's the most difficult part of distributed stream processing?
It was low. I got a bump to 90k that year, then 130k when I jumped companies, which I thought was a mind boggling amount. Do entry level devs even get out of bed for $130k these days?
My first programming job in SF paid $60k/year 10 years ago. I'd like to thank big tech for driving salaries up.
On the flip side, my baby likes the small loud airplanes. He points at them and says "oooh!"
Hi Chad!
I was vegan for 7 years, one of my vegan friends had the opinion that human hospitals should be banned and only animal hospitals should be allowed.
Infra teams like it, app devs don't like it.
We did that at Dropbox in Python for a while. Though they switched to async after I left.
At Google, they found that engineers L5 and above got more work done with RTO, and engineers at L4 and below got significantly less work done. WFH is great but it doesn't work for fresh engineers (who are often the most gung-ho about it as well).
We use Python to generate these configs at my work. Ends up working out pretty well. I previously worked on the biggest deployment of gcl (the inspiration for KCL, Jsonnet) at Google and it was a giant nightmare and the cause of many outages.
Nobody thought Instagram and WhatsApp were good acquisitions at the time.
By health check, do you mean the kubernetes liveness check? Does that make kube try to kill or restart your container?
When using Cloudflare Workers as an API server, I have experienced requests that would “fail silently” and leave a “hanging connection”, with no error thrown, no log emitted, and a frontend that is just loading. Honestly, no idea what’s up with this.
Yikes, these sorts of errors are so hard to debug. Especially if you don't have a real server to log into to get pcaps.
Reminds me of gcl (yikes).
The idea that engineers at Google don't get rich is not based on reality.
Yea, I don't have a problem with the quota, more that the "out of quota" throttling lasts 2+ hours even after the traffic spike dies down.
I'm curious why people pick Azure, if anyone here has direct experience with making the decision.
I work at a startup that runs on Azure, and we're only here because of Microsoft's monopolistic behavior. We switched because Microsoft gives Office 365 discounts to our customers as long as all the the SaaS services they use are hosted on Azure, and so our customers demanded we use Azure. Part of the monopoly playbook: "using a monopoly in one area to create a monopoly in another".
I used to work at GCP, and I thought it was almost shameful that we were in 3rd place behind Azure. Now it just makes me mad (especially since I had to migrate our startup from GCP to Azure).
Yup, the hardest part about migrating to Azure was jumping between all the managed services that work everywhere else but are insanely buggy on Azure. We ended up with the most basic architecture you can imagine (other than AKS, which works great as long as you don't use any plugins) and we're still running into issues.
We have a very long list of Azure features and services that we've banned people from using.
Just got off a call with someone at Azure today who told us to setup our own NAT gateway instead of using Azure's because of an outage where we made too many requests and then got our NAT Gateway quota taken away for the next 2 hours.
Are there any good alternatives to gunicorn out there that don't require async? We're not ready to migrate everything to async, but gunicorn's forking model is blowing up on macOS 15.
Very cool! Finding the Kubernetes API docs is such a pain.