HN user

achiang

442 karma

https://sfba.social/@chizang

Posts0
Comments72
View on HN
No posts found.

I was a sysadmin (at uni, in the early 2000s) and I am an SRE today (at Google).

The two jobs are nothing alike, at all, whatsoever.

Sysadmins are support roles. Their functional role is to provide a healthy substrate to run the application layer on top of.

SREs work at the application layer itself. If the system can't scale due to internal architecture, an SRE would be expected to propose a new, scalable design. That would be in addition to maintaining the substrate.

To be clear, there is also nothing inferior about performing a support role. No org can succeed without support.

But the two roles are not the same, and if a job's set of responsibilities don't include shared ownership over application layer architecture, then it can be a great job but it's not an SRE role.

It has been 10 years since I left Canonical (on good terms), but what popey describes (hi popey) about the intentional lack of human review in the Snap store sounds very Canonical to me.

I agree with all the recommendations - add human gates. Yes, it's expensive, but still far cheaper than the unbounded reputational damage that just occurred around the untrustworthiness of the store (hi Amazon).

Full body scans are a common preventative measure in Taiwan.

My parents (expats, living in the US for over 50 years) flew back and got routine scans (MRI, PET, CT) in February for about $1000 USD total.

Similar to this story, they found a tumor on my dad's pancreas. A biopsy confirmed it, and he had surgery in August. They caught it at stage I. We're very lucky.

The latency from February til August was entirely convincing the US medical system to take his Taiwanese images seriously. They finally gave up and went back to Taiwan to get the procedure done.

I'm getting older myself and will absolutely be paying for any sort of imaging available.

This should be more broadly available to everyone. I'd be happy for more of my tax dollars to go to preventative care rather than rear guard action.

Real world usage is you only get to use ~70% of the stated range on a road trip, so we're really talking about 350 miles of range, which is, as you say, what most people actually want.

Why 70%? You obviously don't run the battery to zero, 10% is a common amount of buffer to leave. And then when you DC fast charge, the rate of charging drops dramatically around 80%, so people don't charge to full.

These are for ideal conditions, add in any sort of weather and the range drops again as you run a heater, etc.

Living in the Bay Area, driving to Tahoe in the winter without a mandatory recharge should be the gold standard.

It's not an unusual use case, "only" about 180 miles, and yet there aren't any EVs that can do it confidently because going uphill in the cold with aerodynamic-destroying ski rack is really hard.

A car with 500 miles of fair-weather range could probably do it?

You need to use the fully loaded cost of an employee when estimating opex savings, which includes health care costs, retirement funding, etc.

Rule of thumb is that fully loaded cost for US employees is approximately 2x yearly salary (although people who've actually run a company can correct my potentially stale or incorrect understanding).

Google SRE here.

Point of clarification on "losing skills" by not being oncall enough.

Google designs its SRE teams to scale sublinearly to the service, which means we're often responsible for entire portfolios, not just a single service.

It's common for individuals to be SMEs on a subset of the portfolio, as the focus of their coding projects.

Obviously the rest of the portfolio is also undergoing continuous change, and so to remain a broadly effective oncaller across the entire portfolio, you need regular production exposure to the parts you interact with less frequently.

Those are the specific bits that get rusty with disuse.

The semantic tagging is nice, I might start incorporating that into my notes.

On the overall topic of meeting notes, I picked this up as a new skill in the past year and it's been immensely valuable for myself and the people I meet with.

Specifically, I learned how to take realtime notes during the meeting while also listening and paying attention. It took practice but was achievable.

One key to success here is to explicitly ask for a helper to take notes while I speak. I've found this helps make the note taking seem like a whole team effort.

Colleagues have noticed and valued the notes and they do seem to lead to better meetings.

The politics is exactly the point of my comment.

The professional way to write a blog post like this is from your own perspective. Identify the proximate cause (the peer), name names if you must, talk about how awesome your own systems are, show some of your monitoring if you like, and talk about what you'll do in the future to be even more resilient to this class of problems.

That's all to the good and much of Cloudflare's blog was exactly that. Would've been fine if they left it like that.

Acknowledging there is no postmortem (yet) but then pointlessly speculating about what it might contain is what I have a problem with.

I don't speak for Google but if I found out we had written a post like this, I would speak up and advocate to change it.

Google networking SRE here (my team runs ns[1-4].google.com among other services).

Regardless of original intent, the blog doesn't land well with me. It could have provided the background on flowspec, using their own past outage as a case study, without any of the speculation or blameyness that came across here. The #hugops at the end reads quite disingenuously.

We see other networks break all the time and we often have pretty good guesses as to why. But I personally would never sign off on a public blog speculating on a WAG of why someone else's network went down. That's uncouth.

The design you propose is stateful, and if you read the chapter closely, you can see we spend a lot of effort to make things stateless.

The main thing I wanted to respond to in this thread about a single bad server destroying your yearly SLO is described in the first paragraph in the section on load balancing at the virtual IP address.

(You will notice that people like Google and Cloudflare skillfully respond with only one record with a 5 minute TTL. That is so the behavior of the browser is well defined, but it also eats their entire year of 99.999% uptime with one bad reply. Your systems had better be very reliable if DNS issues can eat a year's worth of error budget.)

This chapter in the Google SRE book explains how our load balancing DNS works:

https://landing.google.com/sre/sre-book/chapters/load-balanc...

Source: my team runs this service

So here's a nuanced view I'm sure will get downvoted into the ground: both FB and the employee were right, but along different dimensions, and this outcome was not only inevitable, but desirable.

The employee, as a white male in tech, is absolutely morally right to use his privilege to call out other powerful white males for their silence.

And make no mistake, silence is complicity. Many smart philosophers have written about this, see MLK Jr. or Maya Angelou for more.

This is the core of being an ally. Use your privilege to make the hard ask from your peers that a less privileged person, who is decidedly not a peer, cannot.

FB, on the other hand, is also right in a different sense, to maintain internal expectations that singling out colleagues with your political opinion in public is ineffective at best and toxic harassment at worst. FB are signalling to the rest of their employees what behavior they will not tolerate.

In the end, this employee leveraged awareness several orders of magnitude more than had he not been fired (and will likely easily find a new job) and FB protected whatever they believe their culture to be (and whatever other HR lawsuits they believed themselves to be at risk for).

Accurate headline but incorrect analysis.

Big Tech pays like sports, not because of average salary levels, but because of the spread between highest and lowest paid engineers.

Let's say an entry level role in big tech pays about $200k per year in total comp.

It would not be surprising for your top engineer (Jeff Dean ~= LeBron James, e.g.) to rate north of $10m in annual total comp, so 2 orders of magnitude difference.

Multiples in sports are higher, but the point I'm making is that just as LeBron makes multiples of what a bench warmer does, so do the Jeff Deans of the world make multiples of what new college grads do. This is a markedly different landscape vs say, the late 90s when spreads were much MUCH tighter. Unfairly so in my opinion.

Disclosure: I work for Google but have no special knowledge of Jeff Dean's (or any other superstar) comp. I simply claim I wouldn't be surprised if I ever learned the real numbers. :)

They are truly standard behavioral questions that you can't really prepare for, other than thinking about what you did in various scenarios.

"Tell me about a time you had to resolve conflict between two engineers on your team."

And then a bunch of follow-up questions. "What went well/poorly about that", "what would you do differently next time", etc.

This is why I think it'd be hard to go directly from an IC role to a manager as an external hire.

FWIW, I think that hesitation would apply at any company, not just Google. In my last company, I myself was in charge of hiring other engineering managers, and I can tell you that I didn't even consider any resumes unless they called out some sort of lead role.

If they were an actual manager, i could skip directly to the behavioral questions. If they were a tech or project lead of some sort, I did a lot more probing on the exact scope on how much they dealt with people, what they were and were not responsible for, etc. before even getting into the behavioral stuff.

Managers are hugely influential in any org and hiring is an inherently risky activity. You want to minimize risk, not increase it by hiring someone who's never done the actual job before.

Back to Google, I'm not sure the technical bar is lower at all. I got literally the same questions that any senior IC would get, just fewer of them in order to have time for the manager sessions.

No idea whether recruiters care about business school. From my own personal observation, b-school can prepare you to do some analytical stuff, like cash flow analysis or broaden your knowledge base by reading M&A case studies, but nothing in there prepares you to be in charge of running a team with actual humans on it.

I joined Google last year, hired directly as a manager. Germane to this thread, I'm in my early 40s.

At Google, the bar is that you are expected to be able to contribute as an equivalently senior IC, but will be expected to use those skills to inform how you do manager stuff.

So you will definitely get technical questions, the type of which will vary slightly depending on which level of management you're interviewing for. But they will be legit technical questions, like solving a graph theory problem or designing a distributed system.

In addition, you'll also get explicit sessions probing you on your leadership style and management fundamentals. Standard behavioral stuff.

Non CS background doesn't seem to matter as much as actual leadership experience afaict. Formal leadership roles seem to be weighted more strongly when considering your level, but you do seem to get some credit for informal roles too. I'm not super sure on this point.

I think it'd be hard to go directly as an external IC to a manager role. That's not a risk I'd personally want to take if I were the hiring manager in that situation.

If it helps you calibrate, I had about 6 years experience as a line manager at other companies, and I was considered for (and hired as) a line manager. I was never being considered as a 2nd level manager of managers.

hth.

All managers in Google (engineering and non-engineering) are encouraged to take a 2-day immersive course. There's a course for "experienced manager, but new to Google" and a course for "new manager" with content tailored to the two different populations. There are also many many many mandatory trainings that span the gamut from allyship and inclusiveness, to local laws, to how we do performance reviews and comp. In the first year alone, I would guess something O(~weeks) of these trainings.

There are also countless hours of opt-in training for pretty much any subject where you want to improve your skills.

In Google SRE, combo TLMs are considered to be an acceptable short term solution for team turnover, but not a long term best practice. TLMs are highly encouraged to find a different TL for the team.

In addition, all SRE managers (SRMs) are paired with an experienced manager as their mentor.

Ultimately though, apart from the mandatory trainings, no one can really force you to be a better manager. The big feedback mechanism is that internal mobility is very high, to the point where managers have essentially no power to prevent anyone from leaving their team. So if you suck, you will get bad scores in your own performance review and everyone on your team will just transfer away.

I called out the SRE-specific bits, and a few common practices, but it could be that there are different practices in other engineering orgs.

Source: I'm an SRM, and speak only for myself.

Nice article, I'd be curious to know if/how their engine abstracts the complexity behind multi-currency transactions, or whether they rely on the accounting model to handle multi-currency.

We built our own double-entry accounting engine at my previous company, and while the engine was not as fancy as what Square describes, the real challenge was building out the accounting models that manipulated the engine's primitives.

To this day, I have yet to find another resource on multi-currency that is as solid as this one:

https://www.mathstat.dal.ca/~selinger/accounting/tutorial.ht...

Why Google+ Failed 7 years ago

Google+ failed, but it begat Google Photos which is not a failure.

obDisclosure: I'm a Googler, but joined way after G+ both launched and failed.

Google SRE doesn't have magical incident response beans that we hoard from the rest of the world. What makes Google SRE institutionally strong is that we have senior executive support to execute on all the best practices described in the book:

https://landing.google.com/sre/sre-book/toc/index.html

At my last job, I bought a copy of this book, but we only had the organizational bandwidth to do a few of the things mentioned. At Google, we do all of them.

The incident on Sunday basically played out as described in chapters 13 and 14. There is always the fog of war that exists during an incident, so no, it wasn't always people calmly typing into terminals, but having good structure in place keeps the madness manageable.

Disclosure: I work in Google NetInfra SRE, and while my department was/is heavily involved in this incident, I personally was not.

Also, we're [always] hiring:

https://careers.google.com/jobs/results/?company=Google&comp...

The problem that gRPC solves for you is versioning your messages between your services.

As your json payloads evolve, you're going to encounter pain trying to keep your services in sync, whether it comes in the form of writing parsing code to crack open payloads and do conditional error checking based on the version (and expected fields), or whether it comes operationally in how you actually deploy updates to running services.

I was an HP-UX kernel engineer from 2002 til 2005, a brief interlude writing IA64 CPU diagnostics, and then and a Linux kernel engineer from 2007 til 2010, all on Itanium systems.

In that time frame, it wasn't clear that horizontal scale out architecture (aka "the cloud") was going to dominate, and that scale up systems were going the way of the mainframe. The thinking was that there would always be a healthy balance of scale out vs scale up, and btw, HP alone did $30B+ revenue yearly on scale up with very slow decline, just like the mainframe market, which is still $10B+, even today.

To put that in today's terms, if you pitched a startup with a $30B TAM, VCs will definitely be returning your emails.

So no, it wasn't embarrassing to talk about working on IPF any moreso than it would be to talk about POWER today. It's just another CPU architecture with some interesting properties but ultimately failed in the market place. Just like Transmeta or Lisp Machines.

What should be embarrassing, but clearly is not, is to slag off entire industries not knowing shit about them.

Edit: I think working on B-52 parts would be an amazingly fun job.

Fossil vs Git 7 years ago

The comparison table in section 2 could just as easily live in the git docs, under a page called "why use git instead of fossil".

"we do not break userland, period"

That quote is more accurately read as, "we change things all the time, including user-visible features, and very occasionally, even in breaking ways -- but only because it's impossible to know every single consumer of every single quirk in behavior, and as soon as we learn that one of our changes did in fact break userspace, then we'll change it back".

It's how the kernel community attempts to continue cleaning up decades of tech debt while maintaining the contract with userspace. Honestly, sometimes you just don't know until you try.

It's an inefficient process, but it does sound like the right outcome occurred in this case.

signed, a former kernel developer

I was hired 2 years ago by Angaza to be one of their early engineers. Found the role when I was idly trawling through the massive hiring thread one evening, mildly dissatisfied with my job at the time, but not really actively looking for a new job.

The company's mission was exactly what I was looking for -- a chance to do good in the world with technology. I reached out via email, got a prompt response, and have been here since.

We are still alive, still trying to lift people out of energy poverty, and still hiring. Our entry in the most recent thread:

https://news.ycombinator.com/item?id=12847949