HN user

asuffield

1,633 karma
Posts0
Comments411
View on HN
No posts found.

Hi! I wrote this paragraph. I promise that I'm not an LLM, but I was in about hour 10 of my work day and I was asleep not long after writing this. Any failures in comprehensibility are from exhaustion.

(Other comments have explained the bug so I won't repeat them)

2. Participation in project PRISM

I'm an ex-Googler, and I know how PRISM worked. This did not happen. All the statements made in https://googleblog.blogspot.com/2013/06/what.html are true at the time they were written. The sentiment in the title of that post accurately reflects how everybody involved felt about it.

(I can't talk about whether they're still true because I've not been there in nearly two years, I wouldn't know)

(I'm an SRE at Google. My opinions are my own.)

WoMs are a training exercise, intended to build familiarity with systems and how to respond when oncall. A typical WoM format is a few SREs sat in a room, with a designated victim who is pretending to be oncall. The person running the WoM will open with an exchange a bit like this (massively simpified):

"You receive a page with this alert in it showing suddenly elevated rpc errors (link/paste)" "I'm going to look at this console to see if there was just a rollout" "Okay, you see a rollout happened about two minutes before the spike in rpc errors" "I'll roll that back in one location" "rpc errors go back to normal in that location" ...etc

(Depending on the team and quality of simulation available, some of this may be replaced with actual historical monitoring data or simulated broken systems)

The "chaos monkey" tool, as I understand it, is intended to maintain a minimum level of failure in order to make sure that failure cases are exercised. I've never been on a team which needed one of those: at sufficient scale and development velocity, the baseline rate of naturally occurring failures is already high enough. We do have some tools like that, but they're more commonly used by the dev teams during testing (where the test environment won't be big enough to naturally experience all the failures that happen in production).

(I'm a Google SRE. My opinions are my own.)

That's not what our 20% time is for, and 20% is way too small a number for that purpose. "20% time" (the way we use the term) is for personal/career growth/scratching itches.

Time spent on building systems that make our service better is my primary job. Manual remediation ("toil") is something to be tracked as a dangerous antipattern that must not be allowed to take over.

Toil and oncall response should be less than 20% of my time, together. At least half my time should go into engineering projects. If the level of toil is in excess of 50% of team activity then I would expect only percussive intervention to get the team out of this situation.

(I'm an SRE at Google. My opinions are my own.)

core subjects, like container management (this happens...everywhere nowadays) or cluster management

Curiously, these are subjects which most Google SREs won't know much about. One team deals with all that stuff as a service so the rest of us can get on with something else.

What would I pick out as our core skill sets? Ignoring technology-specific details that won't apply anywhere else: troubleshooting a system that you don't understand (reverse-engineering it as you go), and non-abstract large system design.

Since you're posting from a throwaway account (149 days old with no prior comments): please email me your old @google.com username. Mine is the obvious one. I'll happily add a confirmation to this thread when I get it, it only takes a moment to check.

Yes, there are many things I just can't share, and the only data point you can get in that area is my own opinion. How much value you place on that is up to you; I'm giving you the only thing I can. If that is of no value to you, you're free to discount it. I freely acknowledge that I can't prove you wrong. The only alternative I have is to say nothing at all, which is what I usually do. If you would prefer to have no input from people like me at all, by all means say so.

It's obvious that discrimination has taken place that has affected a fellow engineer and human being.

I would like to make it clear that I am responding only to the comment I responded to, which raised a very specific question that I could answer. I do not feel that I have any basis to comment on the original article; please do not associate what I am saying with that.

Part of the company culture is that we value always being able to expect that everybody you encounter is a strong engineer who will do sensible things when presented with data.

You never go into an encounter with a new person or team being unsure of whether they're going to be difficult. You never have to avoid dealing with "that guy". You get to trust everybody that you meet.

It doesn't just improve overall quality, it makes it a better place to work. Good engineers are happier in this environment, and there is pretty near universal agreement from people who have experienced the results that this is a thing worth preserving.

There are plenty of things about the hiring process which get enthusiastic internal debate, criticism, and data-driven analysis. This is not one of them. This is a thing which we really like.

(Full disclosure: I have gone through the interview process twice, failed the first time, passed the second.)

All I am free to say about this is that we look hard at diversity issues in hiring and put a lot of effort into eliminating them, and that I personally believe we do a better job of stamping this out than any other company I have worked for in my career. I suspect, but cannot prove, that our process gives better diversity results than most of the ideas which are "popular" on HN at present.

I do some work on this personally. I'm just not allowed to discuss the details at present.

(The reasons why I'm not allowed to discuss this are due to tedious bureaucracy, not anything interesting. I have tried to get that changed, but it would require more effort than I am willing to expend on a problem that will go away in time.)

I am not allowed to share our data with you about how well the process is performing.

But I would like to point out that of the two of us, only the one who doesn't know is suggesting that we have "terrible hiring figures" or a "terrible interview process".

(Tedious disclaimer: my opinion only, not speaking for anybody else. I'm an SRE at Google.)

We expect and accept a high false-negative rate. Our interview process is optimised for zero false-positives at the cost of many false-negatives. This is a deliberate choice. So yes, I would expect to see a significant rate of rejections of people who are clearly qualified.

The sort of people that we want to hire are likely to come back for another try anyway, and the long-term effect of this process seems to be doing what it was supposed to.

I also went through some of the source code. That "adds things to hosts file" code has some rather questionable entries in it.

m.hotmail.com

watson.microsoft.com

Assorted *.msn.com domains

apps.skype.com

msftncsi.com

"Add spying domains to hosts file" is dishonest, at best. This appears to be a determined effort to break random services for the user which happen to be run by Microsoft. Hotmail, Skype and the NCSI detection are particularly inexcusable things to block under the guise of "destroy spying".

Ah, I see. So what they're saying in this case is that 0.005% is the fraction of Apple's global non-US profits that was paid as tax in Ireland.

That's not the same thing as saying it was "a 0.005% tax rate". The issue is over what fraction of the non-US profits are taxable in Ireland, not over the tax rate.

They could also build a larger road for the most popular commuter route. Or even some form of public transport that's good enough for people to want to use. I realise that the idea of spending tax funds on improving infrastructure is considered radical in California, but I feel that it has some weight of historical evidence behind it.

If they had renters, they could pass those savings on. The article discusses how landlords are keeping units empty because it has become too risky/expensive to rent them.

When the law has become so restrictive that people would rather not engage in commerce at all, then it is broken. This isn't helping anybody.

(Tedious disclaimer: my opinion only, not speaking for anybody else. I'm an SRE at Google.)

Performance. gRPC is basically the most recent version of stubby, and at the kind of scale we use stubby, it achieves shockingly good rpc performance - call latency is orders of magnitude better than any form of http-rpc. This transforms the way you build applications, because you stop caring about the costs of rpcs, and start wanting to split your application into pieces separated by rpc boundaries so that you can run lots of copies of each piece.

I cannot sufficiently explain how critical this is to the way we build applications that scale.

I'll just point out that in the UK and most of Europe we don't see a reason why this should be reciprocal. You can quit at any time, but you also can't be fired without a reason if you've worked at a company for more than 1-2 years.

It doesn't stop people from firing bad coworkers, and it appears to have no negative effects on employment.

This is mostly to provide documentation in the event of a wrongful termination suit.

While it might serve that purpose, it's primarily to make sure that middle-management takes reasonable steps to let people correct and thinks the decision through before acting. Nobody wants to work in a place where people get surprisingly or randomly fired, or where a manager is firing all the people they don't like. Having processes (mostly) prevents that sort of thing from happening.

To fully answer questions like "how much traffic will go in this direction?" you need your analysis to include a simulation of what the entire internet is doing. That's hard.

I can't talk about the details, but you can assume that "static analysis" of the form being talked about here is something we've already done, and it's not enough to handle cases this complicated.

(Tedious disclaimer: my opinion only, not speaking for anybody else. I'm an SRE at Google. My team is oncall for this service and I know exactly what happened here; I probably can't answer most questions you might have.)

Let's go with "yes", as the most accurate answer. As soon as I or whoever is oncall has figured out what change was responsible, we can usually revert it quickly and easily. Usually, if I'm oncall and I have reason to even suspect a recent change might be the cause, I'll revert it and see if the problem goes away.

The difficulty becomes more apparent when you realise the sheer number of infrastructure changes being made every hour, some of which will be fixes to other outages, and some of which will be things you can't revert because they are of the form "that location has fallen offline; probably lost networking" or "we are now at peak time and there are more users online". So if your question is "can we just roll the whole world back one day" - no, too much has changed in that time.

(Tedious disclaimer: my opinion only, not speaking for anybody else. I'm an SRE at Google. My team is oncall for this service and I know exactly what happened here; I probably can't answer most questions you might have.)

Perhaps your architecture wouldn't "compile" if the network traffic will go the wrong place, or if a rate limit is above the capacity something is expected to handle, or if the change would impact too many servers at once.

So in the first instance, I tend to like this sort of idea. However: we are already substantially ahead of the sort of things that you're thinking of.

Full static simulation of a system as complicated as all the components involved here is... well, I can sort of see how it could be done, but it would be a herculean effort; I don't think it would ever be good enough to catch cases like this the first time they happen. There are systems where this sort of thing can be done, but all the ones I can think of are much smaller in scope.

The value is immeasurable, it's just that finance can't get a piece.

Plenty of people make money from this: your ISP, the utility company that pulls cables under the road, the construction company that dug up the road to put them in, the owner of the site where your ISP's network equipment is located, and many of the websites that you interact with.

I believe that the source of the confusion here comes from trying to slice the world into "VC monetization" and "human value", as if these were different or opposed things. I have a different way to look at this space which reveals useful insights:

We see lots of technologies go past that people seem to be excited about, but which then fail in the market. It is currently popular to imply that this means "the market" is some alien thing which is not aligned with what people want. A more realistic view is that people have multiple levels of interest. We can order some of them from lowest to highest:

- willing to read an article - willing to write a comment on the article - willing to blog about the technology - willing to open their wallet - willing to pay the full cost of making it

What we see is that a lot of technologies can only reach levels 2 through 4: people are interested, but not interested enough to cover the cost of making it. By any reasonable standard, that means we shouldn't make the thing: its value to people is less than the value of the raw materials that went into it. "Failed in the market" is a way of summarising this decision, but it gets a lot of negative press because it hides all the details so people don't understand the value comparison being made here.

The neatest mnemonic to think about this is "money is the unit of caring: you can measure how much people care about a thing happening by measuring how much money they are willing to spend on it".

"VC monetization" fits neatly into this picture: VC want to know more or less immediately if people are going to reach interest level 4 or 5 on this scale. They do not want to burn time and money on things which can only reach level 3: those things never had a future. You cannot tell the difference without asking people to open their wallets.

I do not believe this statement to be correct: it seems entirely possible for this to be done via reckless incompetence, rather than criminal fraud. All it takes is for somebody to calculate the risk incorrectly and everybody else to fail to check their calculations.

It may involve criminal fraud, but it is also possible that it does not.

You seem to be talking about some specific incident, and guessing about what happened. We're discussing the general case of what you do when you've got a large collection of data that isn't yours, and all you know is that some part of it can't be distributed.

how difficult would it be to let him access his data sans the illicit material?

It's not my service, and I'm talking about the general case rather than this specific one, but that seems to me like it would be pretty complicated - you'd need some sort of review procedure to determine what material can and cannot be distributed, you'd need to somehow do this while preserving user privacy, and you'd need some engineering work to make all this possible.

A project of that scope could take weeks or months to complete, depending on the amount of data involved.