HN user

CrankyFool

39 karma
Posts0
Comments23
View on HN
No posts found.

Honestly, I think 40 direct reports is borderline insane. You can't effectively support that many people (as you note yourself). Speaking as the author of two of these READMEs, at Netflix I had at peak around 12 direct reports (one was a manager with his own reports) and at Slack I have 3 (all of whom manage their own teams).

(Hi, I'm Roy Rapoport -- the author of both slide decks).

Referring to the Netflix version of the slide deck as "the old version" of the Slack version of the slide deck is inaccurate. I'm in a somewhat unusual position amongst the people whose READMEs were quoted, because I got to write one from both sides of the fence: As the established person welcoming a new member to the team (Netflix), and as someone joining a new organization (Slack). READMEs, for me, then end up looking very different.

It's also worth noting, as geofft notes, that Slack's culture and Netflix's culture are very very different, so the concerns of people within them would be different (and should be addressed differently). The concern that you might be fired was pervasive for new people at Netflix, and I sought to alleviate it; it's not a commonly-held concern at Slack.

(I manage at Netflix)

We can't publish stats about vacation usage.

That's because I don't know of a single manager here who tracks vacation usage. There's a general allergy to doing that, because that can lead to trying to manage that number and then the 'unmetered vacation' is no longer so unmetered.

I suppose I could, for my people, try to find all the out of office notifications, but that'd be a silly level of effort.

(There's no consistent process across the board for dealing with vacations, but to the best of my knowledge the typical way this works -- and the way it works in my own group -- is that an engineer will at some point probably mention to me that they're taking days off casually. I try to make sure it's clear to them that they're not asking for permission, and then move on).

I'm a hiring manager at Netflix. If one of my employees told me "I'm going to take the year off, see you in a year," I'd basically go "OK, have a great time with your kids. Send us a picture every once in a while."

And then I'd backfill them. And when they got back, I'd have an extra engineer. Chances are by that point I'll be looking to expand the team anyway.

Speaking as a hiring manager in a tech company here for a moment, I do find there's a shortage of reasonable people to consider. I found this thread because someone had pointed me at unemployable.pen.io and wanted to comment on it specifically, because having read that particular blog post, I started off thinking "man, that's someone I'd like to talk to," and then realized a few things:

If you're trying to get hired, then maybe when you write a blog post on the fact, you also include a link to your resume. Or at least your name. Because otherwise, I'm sitting here thinking to myself "welp, no idea who this person is"

This person acknowledges that for more than a decade, she/he coasted in an office job. I tend to hire for strong initiative and drive, and I'd worry about this person. And that still wouldn't stop me from talking to them IF THEY MADE IT EASY FOR ME TO FIND THEM.

Tolerance of failure is different for ICs and Managers, partially because Managers tend to have such an ability to really screw things up in a way that isn't technical and has an impact on a bunch of other people in the company.

The shortest IC tenure I've seen here was the result of being a brilliant jerk (which interviewers did not catch during the interview cycle (and I speak here as one of the people who interviewed this person)). The shortest Manager tenure I've seen here was noticeably shorter than this, and was the result of pissing off your engineers.

I've not seen any IC here screw up on a technical level in such a way as to get fired for a first or even second screwup -- it really takes a pattern. My shortest IC termination was six months from hiring to termination.

You can't just implement "unmetered vacation" and have it all work out -- it has to be part of a set of practices.

For example, at Netflix as a manager I don't "veto" any proposed vacations, because my engineers do not propose vacations -- they tell me what time they're taking off. It's their job to make sure that's not going to be a surprise to anyone covering for them, and it's not my job to check that they've done that work.

As for being careful not to be the guy who takes too much vacation ... putting aside the fact I have people reporting to me who aren't guys, there are people on my team who take two days of vacation a year; there are people on my team who take ~6 weeks of vacation per year. Nobody here seems to care much.

Working at Netflix 12 years ago

I'm dmuino's manager.

The day after I took over the team, my director sat me down and said "I don't know if you've looked yet, but most of your engineers make more than you do. That's because they're incredibly valuable. Your number one job is to keep them happy."

In the last two years of managing this team, I've consistently gotten paid less than the top engineers on my team.

Totally OK with it.

Working at Netflix 12 years ago

I'm a hiring manager at Netflix and spend a bunch of time talking to other hiring managers at Netflix.

I don't know anyone who does trick questions (most people I know find them anathema, and even Google -- famous for them -- acknowledges they're useless). I definitely do know some people ask the 'cycle in a linked list' question (or others like it); honestly, I don't see a problem with that set of questions.

Working at Netflix 12 years ago

As someone responsible for observability, I've got a real problem with biasing hiring decisions toward false negatives; what you're basically saying is that you're biasing the system toward the failure mode that you have the least ability to actually measure.

It's a great way to feel good about yourself and your interview process; it's also a great way to not find fault in your interview process as it becomes arbitrarily ... arbitrary.

We CAN do <1m, but very rarely do (at least in terms of persisting and showing it). We have a feature called 'Critical Metrics' that is a separate publishing pipeline into Atlas that is shorter, simpler, hardier, and supports 10s granularity, though for a trivial minority of metrics -- our current limit is on the order of around 400K metrics per cluster, IIRC, which means that if you've got, say, 1000 nodes, we're going to limit you to 400 metrics per node that would be flowing through the Critical Metrics pipeline.

(We haven't opensourced much of the pipeline ... yet)

above, copperlight quotes about 3%-5% of metrics being system-level. A pretty small number would be process-level, I'm guessing, with the vast majority being app-level.

At Netflix-sized, the answer to pretty much any question is "a lot." :)

It's actually more like close to 1.2 billion different time series -- we report most metrics on one minute granularity, but they're not all reporting at the same second (thank God), so on average we're getting up to 20M time series per minute.

But this of course just makes the question more reasonable -- 1.2B different time series? Really?

Yup. We get a bunch of system telemetry, and a bunch of default application telemetry, without even getting traffic hitting the box, but that's a relatively small percentage of the overall volume. Developers LOVE metrics.

So imagine you want to measure requests to our API, and these are some tags you want to keep track of: request type: 5 different types result: 2 possible values (success, failure) originating country: 50 countries originating device type: 200 devices

And let's say you've got a 1000 instances reporting this data.

Suddenly you've got 5 * 2 * 50 * 200 * 1000

Oh look. Here's 100M different metrics.

And that's a relatively trivial example.

"Anomaly detection" is one of those vague terms that can mean anything from "it's gone above the pre-set limit, and that's anomalous" to "the system has studied the signal to learn what the accepted limits should be, and it's exceeded these limits." We mostly mean the latter for anomaly detection.

The Insight Engineering team at Netflix is largely composed of four kinds of engineers: Platform/back-end engineers, UI engineers, Site Reliability Engineers, and Real-Time Analytics (RTA) engineers; it's the latter group of engineers who are looking into ways to quickly (and efficienlty) detect anomalies in a truly-absurd amount of data.

The RTA group is now about 6 months old or thereabouts; I have high hopes that we'll see some public presentations from them soon that will be helpful to other people outside Netflix.

"It's complicated."

As the announcement notes, we have multiple tiers holding different data horizons. The most active, and large, tier is the one holding the last six hours of data. That tier, being the most critical one (we try to train our engineers to only need 6 hours of data to understand how their system is working in the worst case), is mirrored. Right now, each of those mirrors is about[0] 756 r3.2xl instances.

[0] For a very exact definition of "about," though that exact number could change in the next 5 minutes, or 5 hours, or 5 days, or not until we see another metrics increase.