HN user

0x002A

79 karma
Posts9
Comments30
View on HN

https://www.theregister.com/2025/10/20/aws_outage_amazon_bra... quoting from this

"And so, a quiet suspicion starts to circulate: where have the senior AWS engineers who've been to this dance before gone? And the answer increasingly is that they've left the building — taking decades of hard-won institutional knowledge about how AWS's systems work at scale right along with them."

...

"AWS has given increasing levels of detail, as is their tradition, when outages strike, and as new information comes to light. Reading through it, one really gets the sense that it took them 75 minutes to go from "things are breaking" to "we've narrowed it down to a single service endpoint, but are still researching," which is something of a bitter pill to swallow. To be clear: I've seen zero signs that this stems from a lack of transparency, and every indication that they legitimately did not know what was breaking for a patently absurd length of time."

....

"This is a tipping point moment. Increasingly, it seems that the talent who understood the deep failure modes is gone. The new, leaner, presumably less expensive teams lack the institutional knowledge needed to, if not prevent these outages in the first place, significantly reduce the time to detection and recovery. "

...

"I want to be very clear on one last point. This isn't about the technology being old. It's about the people maintaining it being new. If I had to guess what happens next, the market will forgive AWS this time, but the pattern will continue."

I think business types vs technical types inherently have different perspectives especially for american companies. One has the "get it done at all costs" the other has "this can't be done since impossible/it will break this".

When a company moves from engineering/technical driven to sales/profit/stock price/shareholders satisfaction driven, once it was not possible to cut (technical) corners, now becomes the de facto. If you push the L7s/L8s out of the discussion room, who would definitely stop or veto circular dependencies, and replace with sir-yes-sir people, now you've successfully created short term KPI wins for the lofty chairs but with a burning fuse of catastrophic failures to come.

As Amazon moves from day-1 company as it claimed once, to be the sales company like Oracle focusing on raking money, expect more outages to come, and longer to be resolved.

Amazon is burning and driving away the technical talent and knowledge knowing the vendor lock-in will keep bringing the sweet money. You will see more sales people hoovering around your c-suites and executives, while you will face even worse technical support, that seem not knowing what they are talking about, yet alone to fix the support issue you expect to be fixed easily.

Mark my words, and if you are putting your eggs in one basket, that basket is now too complex and too interdependent, and the people who built and knew those intricacies are driven away with RTOs, move to hubs. Eventually those services; all others (and also aws services themselves) heavily dependent on, might be more fragile than the public knows.

Any company CEO who lays off people openly or covertly because they hire too many people in the recent years. If the management decided to hire too many people they are the ones who are responsible for bad decision or bad management. Not the people who are hired by the decisions.

Dropbox Backup 4 years ago

It's also not clear for extra storage, it seems you need to jump to the next plan to get a slightly larger capacity. This does not make sense since the biggest tier is business with 5TB cap. If you need more you fall under the enterprise "contact sales" tier.

Do I understand correctly; that humans ate first the bigger animals to extinction and moved to next smaller ? Isn't this counterintuitive? Bigger animals need better traps and (killing) tools while smaller ones can be catched by hand and can be choked ?

If I had to choose between a moose and a rabbit to hunt, I would go for the cute little bunny.

But don't take my choice as guidance, I grew up and lived in city.

Each time a developer does something on a cloud platform, that moment the platform might start to profit for two reasons: vendor lock-in and accrued costs in the long term regardless of the unit cost.

Anything limitless/easiest has a higher hidden cost attached.

Pip is to create document trail (so when they get sued they can show evidence ) and to force people into submission and resignation. So this way they can handle things quietly and dirty things kept undercover.

That's what majority of the people think and do, those who find themselves in the similar situations and that's how companies like Amazon can continue doing this shady practices.

The upper management only gives a blink when it becomes more known by the public. They don't want to hear (just not to be in the responsibility zone) and when it's widely known they would go whitewashing or PR stunts.

Why they had to invent "to be the best employer of the world" as a leadership principle?

The more people should stand up and make the stupid policies known, that's the only way this can stop. Otherwise this policies will be the norm for many companies since "it can work"

Yes that's what you get when you have one thread for all network communications. The network stack did not fail, only the sole network thread got stuck. From the write up I understand if there was another thread for communications firefox would only fail to communicate with telemetry service but firefox would be able to function as users needed.

Yes that's another issue. But as a design pattern shouldn't we design our products to do their core functionality as much as independent from any anomalies that can happen? This is almost akin to me if Tesla rolls out an update and the car decides to pull over to the curb to do the update, while you are driving to your job or worse to hospital with an emergency. My theory is there should be at least one health enterprise using firefox as their only browser for business functionality out in the wild.

"Because all network requests go through one socket thread, this loop blocked any further network communication and made Firefox unresponsive, unable to load web content." Why side functionality (telemetry) of a tool uses only one network thread and can block any network communication ?

Thousands of years ago felines decided to domesticate the hoomans for two reasons: continuous and scheduled supply of food and door opening functionality on demand.

The issue to me is companies are not perfect machines and letting machines chewing up human beings is an issue that humans beings can address and solve. I think more people should stand up when they face unfairness, regardless of they face it or somebody nearby, also communicate and help each other. I don't think when people have accept their fate and go to the next job, they are relieved and not bringing the emotional baggage. People (ideally) should be changing jobs for a change or a new challenge, or better pay or things appealing them more in a new job, not escape from another.

So yes this is the easy way, but wouldn't this also help the vicious circle to continue? Would this let bad managers to push out people whenever they feel like they are challenged to manage or just as they like it, successfully everytime ? Toxic culture thrives on that and I don't think it stays contained in the company(ies) people are pushed out from.

I think we are hearing more and more managers put people in PIP or other manage-out tools as a retaliation and everyone suggests to apply other jobs when there is a confrontation. Well, when all the companies share the toxic management culture it will be too late ask the culture to change, won't be? I believe more people should stand their ground when they face unfairness, hostility and/or bad management for some more time as they can afford. Making bad managers life easier does not solve the bigger or growing problem. It won't change as long as they go away silently and conveniently.