HN user

luu

113,217 karma

danluu.com | mastodon.social/@danluu | twitter.com/danluu

Posts6,006
Comments182
View on HN
carnegieendowment.org 8d ago

The Future of American Power

luu
4pts0
jonlu.ca 13d ago

The Web's HTTP Header Junk Drawer

luu
3pts0
www.rntz.net 20d ago

Evaluation order and nontermination in query languages

luu
44pts6
cccg.ca 21d ago

Computational Balloon Twisting: The Theory of Balloon Polyhedra [pdf]

luu
52pts2
graphicore.github.io 26d ago

Libre Barcode Project

luu
288pts65
nerocam.com 28d ago

SCC Technical Assistance Program

luu
24pts1
www.everything.one 29d ago

Everything*: An interactive voyage through all orders of magnitude

luu
2pts0
fgiesen.wordpress.com 1mo ago

PivCo-Huffman “merge” operations

luu
43pts4
pmc.ncbi.nlm.nih.gov 1mo ago

Evaluating sugar-sweetened beverage tax effects

luu
3pts0
zeux.io 1mo ago

Zigzag Decoding with AVX-512

luu
129pts22
blog.plover.com 1mo ago

Egyptian Fractions (2006)

luu
113pts24
homefreesociology.com 1mo ago

The Trouble with Municipal-Level Population Projections

luu
3pts0
www.rhawa.org 1mo ago

The redistribution of housing wealth caused by rent control (2023) [pdf]

luu
82pts148
open.library.ubc.ca 1mo ago

The Rise of Housing Nationalism in Canada and Transnational Ownership Patterns

luu
7pts0
www.sfu.ca 1mo ago

A modular impact diverting mechanism for football helmets [pdf]

luu
16pts7
www.kb.cert.org 1mo ago

Missing IPsec Integrity Protection for IMS Sip Signaling in Verizon VoLTE

luu
3pts0
link.springer.com 1mo ago

Generating Random Factored Numbers, Easily [pdf] (2003)

luu
1pts0
ariadne.space 1mo ago

Writing Portable ARM64 Assembly (2023)

luu
62pts27
lwn.net 1mo ago

Re: [PATCH] OOM_pardon, a.k.a. don't kill my xlock (2004)

luu
91pts98
www.imec-int.com 1mo ago

Quantum dot qubit using High NA EUV lithography

luu
24pts5
phabricator.services.mozilla.com 2mo ago

Bug 1950764: Work Around Crash on Intel Raptor Lake CPU

luu
182pts55
www.rfc-editor.org 2mo ago

RFC 9959: Careful Resume: Convergence of Congestion Control from Retained State

luu
2pts0
dl.acm.org 2mo ago

Encouraging Autonomous Driving Companies to Share Safety-Critical Data

luu
3pts0
blog.beeper.com 2mo ago

Build a Beeper Bridge

luu
2pts0
afilina.com 2mo ago

Progressively Improving a Ball of Mud

luu
3pts0
github.com 2mo ago

LLVM: Add support for poison-generating/UB-implying annotations

luu
1pts0
afilina.com 2mo ago

Why I Created phpc.tv

luu
55pts13
marijnhaverbeke.nl 2mo ago

Collaborative Editing in CodeMirror (2020)

luu
60pts10
niri-wm.github.io 2mo ago

Niri Security Model

luu
2pts0
geolog.sgai.uk 2mo ago

Why Geolog? [pdf]

luu
1pts0

Which campus? That doesn't match my experience in Austin at all, where most people ate the cafeteria even though alternate options were available with a very short drive, and the food was pretty good. Maybe not as good as Google's food at the time, but probably as good as the food at Google the last time I visited. And the food was decently subsidized (I'd eat breakfast there for $2). In Austin, Moto and Centaur known for having really good food, but IBM's food wasn't bad in the early 2000s. On my team, I think one or two people packed their lunch and everyone else would eat at the cafeteria except on special occasions.

I've heard from people who stayed at IBM that the food declined to cafeteria food quality over the next ten years, which led to the cafeteria basically being abandoned because people ate out so much. But that's actually counter to the narrative in the post — IBM had decent food before Google, and then some time after Google's IPO, the food declined to became standard cafeteria food.

It would have better to say Google popularized these practices rather than pioneered them

This also seems incorrect. Before Google, it was common to have company-provided before Google. IBM and Motorola had cafeterias. I don't know when AMD installed their cafeterias, but if it was post-Google, it would've been inspired by IBM and Moto's cafeteria and not Google's. In Austin, the Moto cafeteria was known for having very good food and IBM was moderately subsidized and pretty good until the 2010s, which doesn't line up with Google being influential at all. And Centaur had great, free, food. This is an old idea that predates tech companies that a lot of tech companies have picked up that Google also happened to pick up.

As a term, dogfooding spread through Microsoft after Paul Maritz wrote an emailed titled "Eating our own Dogfood" in 1988. If the term was popularized by anyone, it was probably Joel Spolsky who took the practice from Microsoft and blogged about it when he was the most widely read programming blogger. But there are a lot of examples of people doing this before Martiz's email (they just called it something else) and before tech companies even existed; this is another practice that predates tech companies that tech companies picked up.

I don't know about the history of A/B testing in tech, but Capital One was doing A/B tests at scale before they would've been influenced by Google and that's another idea that was used outside of tech.

Scored a 6 or a 7, depending on the answer to one question that I'm not sure of. I think more likely 7 than 6.

This pretty obviously impacted my happiness when I lived with an abusive parent and I have some physical issues that could possibly be related to starving so much as a kid and some others that fall out of not getting basic dental care, but I don't think it's obvious that my background is a dominant factor in how happy or satisfied I am. I find it plausible that my resilience in certain situations is lower than it otherwise would be due to my background, but it's not clear that it's even lower than average in those situations. And, overall, I was probably happier and more satisfied with my life than most people starting a couple years after moving out on my own until a few years ago.

Maybe the version of me that wasn't abused so much as a kid would've been even happier or maybe that version would still have above average life satisfaction today, but maybe not.

Unlike other striker-fired guns, the P320 is "effectively fully cocked at rest", since its striker is under constant spring pressure, which is released when the trigger is pulled. Most such guns, including the P320's military variant, have external safeties (like thumb safeties), which the P320's civilian variant lacks. According to gunsmith James Tertin, this is a rare and "uniquely dangerous" configuration

(from https://en.wikipedia.org/wiki/SIG_Sauer_P320).

MCO 5500.6H (Arming of Law Enforcement and Security Personnel) requires Condition 1, which is magazine inserted, round in chamber, slide forward, safety on ... After reviewing the security camera video footage, the mishap investigator concluded that P1 did not mishandle the weapon at anytime while on duty at Gate 1 prior to the weapon discharging. From the evidence and statements from the persons involved, it is apparent that the weapon fired while on safe and secured in the holster.

(from https://npr.brightspotcdn.com/86/f5/d87374f8497a81db30e2463a..., linked to by the article)

Some of my most popular posts are throwaway quips and memes that went viral on social media. One of my life’s crowning achievements is this: [witty, throwaway, quip tweet].

In contrast, some of the work I put weeks or months into essentially lost the SEO game and gets nearly zero traffic ... Even though I don’t write for money, there is an immense pressure to produce clickbait — even if simply to add “hey, since you’re here, check out this serious thing”.

This will be different for different people, but I've noticed a moderately strong negative correlation between how much effort I put into something and how much engagement it gets (this seems likely to be different for people who apply their effort to generating engagement). The highest engagement content content of mine tends to be thoughtless social media comments I make without thinking. Something like https://danluu.com/ftc-google-antitrust/, which summarizes 300+ pages of FTC memos and is lucky to get 10% of the traffic of a throwaway comment and is more likely to get < 0.1% of the traffic of a high-engagement throwaway comment. Of course there's a direct effect, in that a thoughtless joke has appeal to a larger number of people than a deep dive into anything, but algorithmic feeds really magnify this effect because they'll cause the thoughtless joke to be shown to orders of magnitude more people so something with a 10x difference in appeal will end up with, say, a 1000x difference in traffic on average and even more in the tail.

I don't think this is unique to tech content either. For example, I see this with YouTube channels as well — in every genre or niche that I follow, the most informative content doesn't has fairly low reach and the highest engagement content leans heavily on entertainment value and isn't very informative.

Finally? If I look at the last N mastodon.social links that have >= 15 points on HN besides this one, they are:

  Porting games to Linux pays for itself
  People go to Stack Overflow because the docs and error messages are garbage 
  Hallucineted CVE against Curl: someone asked Bard to find a vulnerability
  Curl on track for non-experimental HTTP/3 support 
  Jeff Johnson: "Passkeys are a lie and contradiction." 
  Swift was not given the chance to prove itself as a language 
  The only planet where 100% of Linux systems have working audio is Mars 
  Bevy Needs Sponsorship
That's the entire first page of results. On page 2, we have:
  Things that didn't age well 
  Google Maps is a critical dependency for nutrition facts on mcdonalds.com 
  Google Search Is Over
  Downloading a video should be “fair use” as recording a song from the radio 
  A list of recent hostile moves by Google's Chrome team 
  One thing that has changed in the professional game industry is that 
  SunOS 4.1.4 says it can't possibly be the year 2023 
  Mozilla should call for the removal of Google from W3C because of WEI 
  How Stuff Works replaced writers with GPT-generated content and laid off editors 
  A fun new feature we are working on in systemd: userspace-only reboot 
  Mastodon's active user base has increased by 110K over the last day 
  Mastodon has reached 13M accounts
It looks like about 10% "not finally" content and 90% "finally" content. You can see the same thing by clicking "mastodon.social" at the top of this page.

If you click through the link in that sentence to https://danluu.com/empirical-pl/ or read the study itself, you'll see that the paper doesn't support the claims made in the abstract at all.

It used automatic classification that's obviously wrong. Table 1 gives a list of "top" projects for each language and many of them are simply misclassified.

... the "top three" TypeScript projects are bitcoin, litecoin, and qBittorrent). These are C++ projects. So the intermediate result appears to not be that TypeScript is reliable, but that projects mis-identified as TypeScript are reliable. Those projects are reliable because Qt translation files are identified as TypeScript and it turns out that, per line of code, giant dumps of config files from another project don't cause a lot of bugs. It's like saying that a project has few bugs per line of code because it has a giant README. This is the most blatant classification error, but it's far from the only one.

For example, of what they call the "top three" perl projects, one is showdown, a javascript project, and one is rails-dev-box, a shell script and a vagrant file used to launch a Rails dev environment. Without knowing anything about the latter project, one might expect it's not a perl project from its name, rails-dev-box, which correctly indicates that it's a rails related project.

There are other major problems with the study, but that one is sufficient to make the results invalid.

If you can suggest an edit that will fit within HN's title length limit that conveys the sentiment of the entire tweet, please feel free to do so. Appending enough of the missing text to convey anything useful violates the limit, but maybe the title could be compressed in another way.

I don't think I can edit the title anymore, but one of the mods can edit it.

Going mouseless 5 years ago

If you look at what Tog says was actually tested, every single test he describes is completely bogus. There's more detail on this in https://danluu.com/keyboard-v-mouse/, but briefly, a test he describes is:

the author typed a paragraph and then had to replace every “e” with a “|”, either using cursor keys or the mouse. The author found that the average time for using cursor keys was 99.43 seconds and the average time for the mouse was 50.22 seconds

Sure, keyboard-only users who do that exact task who literally use the arrow keys plus backspace are slower than mouse users, but since keyboard-only users don't do bulk search and replace by navigating to each relevant character with arrow keys, that's a meaningless benchmark.

Doesn't seem to be any discussion of Nuhfer's findings . . .

That's an interesting result, but I think it would be pretty surprising if your link was discussed since your link is from 2017 and the post is from 2010 and doesn't appear to have a recent update.

the numbers are pretty much business as usual . . . going purely by deaths the situation is mostly normal

The article notes that, through July 25, there were 235,610 excess U.S. deaths. The link you gave shows that excess deaths started March 28th, for a period of ~119 days.

Wikipedia says that the U.S. had 291,557 combat deaths during WWII, which was over a ~1366 day period of U.S. combat involvement if we count the period from Pearl Harbor until the end of the war. The U.S. is about 2.46 more populous now, so adjusting for that, the ratio of the rate of excess death over the period studied vs. U.S. combat deaths in WWII is (235610 / (119 * 2.46)) / (291557 / 1366) = 3.77.

Since you're calling 3.77x the population-adjusted rate of WWII combat deaths "mostly normal", I would be curious to know what rate of death you would consider abnormal.

On a non-population-adjusted basis, which is arguably the correct measure if we want to think about waging a war that's the equivalent of WWII, the death rate over the period studied would've been equivalent to the U.S. combat deaths from waging 9.28 simultaneous WWIIs. Personally, I wouldn't consider simultaneously engaging in nine wars the size of WWII to be business as usual.

This article makes the reasonable point that the web is likely getting faster for people with cutting edge devices. For example, at one point they say

Someone who used a Galaxy S4 in 2013 and now uses a Galaxy S10 will have seen their CPU processing power go up by a factor of 5. Let's assume that browsers have become 4x more efficient since then. If we naively multiply these numbers we get an overall 20x improvement.

Since 2013, JavaScript page weight has increased 3.7x from 107KB to 392KB. Maybe minification and compression have improved a bit as well, so more JavaScript code now fits into fewer bytes. Let's round the multiple up to 6x. Let's pretend that JavaScript page weight is proportional to JavaScript execution time.

We'd still end up with a 3.3x performance improvement.

But then the author concludes

he web is slowly getting faster

Which ignores a pretty large fraction of users. A part of the article acknowledges that this all depends on the device, etc., but this is ignored in the conclusion!

Let's say that, as a first approximation, the first set of quotes is correct. I think most developers who look at user experience with respect to latency or performance today (or even ten years ago) would agree that we should not only consider the average and that we should also look at the tail. If we do so, we see that device age is increasing at the median and the tail, quite drastically in the tail even if we "only" look at p75 device age: https://danluu.com/android-updates/.

If we consider a user who's still using a 2013 Galaxy S4 and ask "does your phone feel 5 times faster than it did in 2013?", based on some js benchmarks improving by 5x, I think they'll laugh in our face. I've used a couple of Android devices that I tried to keep up to date (to the extent that's possible on Android) and each one became unbearably slow after taking some big Android update. Those updates probably included improvements in the Android Runtime as well as V8, and yet, the net effect was not positive. I don't think I'm alone in this -- if you read any forum where people discuss taking updates for the phones, one of the most common complaints is that their previously usable phone became unusable due to performance degradations caused by the update.

Sure, my personal user experience on my daily-driver phone is ok on my phone because I have very fast phone and I'm often using it from fast wifi. But it's terrible if I take a road trip across the U.S. and the experience is terrible anywhere in the U.S. with an old phone. I don't think we should just write off the experiences of people with old phones or who live in places where they can't get high-speed internet even if life is good for people like me when I'm at home on my 1Gb connection. When I looked at this with respect to bandwidth and latency (inspired by a road trip where I found every website from a major tech company to be unusable, excluding a few Google properties), I found that, on a slow connection like you get in many places in the U.S., websites can easily take more than 1 minute to load in a controlled benchmark: https://danluu.com/web-bloat/. My experience in real life (where I probably had higher variance in both latency, packet loss, and effective bandwidth) was that many websites simply wouldn't load.

One thing this post looks at is the 75%-ile onLoad time. When I travel through the U.S. on the interstate (major throughfares which will, in general, have better connectivity than analogous places off of major throughfares) most pages are so slow that they don't even load at all, so those attempts aren't counted in the statistics! I don't dispute that things are getting faster for the median user or even the 75%-ile slow user if you measure that in a specific way, but there are plenty of users whose experiences are getting worse who won't even be counted in this stat that's in the post because their experience is too slow to even get counted in the stats.

It's possible to use a bloom filter variant of this for text search (for example, Bing does this, see https://danluu.com/bitfunnel-sigir.pdf for details).

If you wanted a very small bundle to use with a static site, I don't think it's obvious that a bloom filter variant is a bad approach.

I mean, yeah, this thing that the author said "made for a fun hour of Saturday night hacking" is probably not the optimal solution, but that would've been true whether or not the author chose to build something based on bloom filters or an inverted index.

The biggest problem is people want the benefits of freely downloadable software but mainly aren't prepared to give anything back. Go and assist with emacs or some other project.

I don't really buy this line of argument in general, but I think Josh is a particularly poor candidate to pull rank on because they don't contribute to open source. I don't know Josh, but I recognize his name from his open source contributions.

If you don't recognize Josh's name, you can get an idea of some of what he's done from his website, which is linked in his bio: https://joshtriplett.org/

I work on Linux, primarily on the RCU subsystem and on Sparse-related code. I maintain the rcutorture test module.

I co-maintain the X C Binding (XCB). I developed the XML-XCB format to describe the X Window System protocol. I also work on other Xorg projects on Freedesktop.org.

I maintained the Sparse semantic parser and static analysis tool for C for several years, before passing it on to Christopher Li.

I maintain several packages in the Debian project.

It's been 30 years since he wrote that. Has anyone in that time been able to produce data that refutes Apple's findings?

Yes, all of the specific experiments Tog describes are bogus. In versions of the experiments where the keyboard user isn't crippled by weird limitations that don't apply to actual computer users, the keyboard is often superior to the mouse (not always, just often): https://danluu.com/keyboard-v-mouse/

This study claims 28% mortality, 46% moderate-severe disability, 26% able to live independently after one year.

https://www.ncbi.nlm.nih.gov/pubmed/31810636

METHODS: Older adults aged ≥65 years were included if they had an isolated hip fracture, were admitted to hospital between July 2009 and June 2016, inclusive, and were registered to the Victorian Orthopaedic Trauma Outcomes Registry. Mortality up to 12 months (365 days) post-injury, and functional outcomes (Glasgow Outcome Scale-Extended; GOS-E) at 12 months post-injury were examined. Multivariable Cox proportional hazards regression was used to estimate adjusted hazard ratios (aHRs), and multivariable logistic regression was used to identify predictors of living independently compared with severe disability or death on the GOS-E.

RESULTS: 4,912 patients were included, of whom 28% died, 46% had moderate-severe disability, and 26% were living independently 12 months post-injury. Mortality rates were lower in women (aHR=0.56, 95%CI: 0.50, 0.63), and in people injured in a high fall vs low fall (aHR=0.47, 95%CI: 0.31, 0.72). Mortality rates were higher in people in the older age groups (75-84 years: aHR=1.53, 95%CI: 1.21, 1.93; 95+ years: aHR=3.58, 95%CI: 2.68, 4.77), living in areas with the highest level of socioeconomic disadvantage (aHR=1.25, 95%CI: 1.01, 1.55), with a Charlson Comorbidity Index weighting of one (aHR=1.60, 95%CI: 1.36, 1.88) or more than one (aHR=2.21, 95%CI: 1.94, 2.53), whose injury occurred in a residential institution versus at home (aHR=2.63, 95%CI: 1.97, 3.52), that resulted in intensive care unit admission (aHR=1.68, 95%CI: 1.21, 2.32), and in people who did not have surgery versus people who had internal fixation (aHR=1.65, 95%CI: 1.33, 2.04). Independent living was inversely associated with most of the same characteristics; however, people also had lower odds of living independently if they were from metropolitan residential areas versus rural areas (aOR=0.77, 95%CI: 0.62, 0.96), or had mild to moderate (aOR=0.33, 95%CI: 0.27, 0.39) or marked to severe (aOR=0.13, 95%CI: 0.09, 0.20) preinjury disability vs no preinjury disability.

Relatedly, this study from the same institution claims 5% mortality rate after a year for people under 65: https://www.ncbi.nlm.nih.gov/pubmed/27527378

I believe the paper is https://www.nature.com/articles/s41467-019-12808-z.

If I'm reading the paper correctly, the authors don't argue that sea levels will rise more than previously predicted. Rather, they argue that the classical technique for estimating impact by using elevation (SRTM) is overestimating elevation and therefore underestimating the impact of rising sea levels. The authors claim that their technique (CoastalDEM) also underestimates the impact of rising sea levels, but reduces the systemic bias present in previous estimates.

Strange, nowhere in the article does it say how did they know that the views where inflated by 150 to 900%. How did they even know, without Facebook admitting it?

I don't think they did

From the complaint at https://assets.documentcloud.org/documents/5004295/d5cb8373-...:

In September 2016, the Wall Street Journal revealed that, for the past two years, Facebook had been overstating the average time its users spent watching paid video advertisements. Based on information from advertising agencies who had spoken with Facebook, the Wall Street Journal reported that Facebook’s metrics had been overstated by between 60 and 80%. In response to the media attention, Facebook admitted it made a mistake, but emphasized that it had only discovered the mistake “about a month ago,” … Internal records recently produced in this litigation suggest, … Facebook did not discover its mistake one month before its public announcement. Facebook engineers knew for over a year, and multiple advertisers had reported aberrant results caused by the miscalculation (such as 100% average watch times for their video ads).

It seems that advertisers saw things that were obviously incorrect (e.g., 100% average watch times). Some kind of weirdness eventually led to a lawsuit. During the suit, the plaintiffs found evidence in internal documents that indicated that at least some engineers in Facebook knew there was a problem and maybe knew the magnitude of the problem. I wouldn't call that "Facebook admitting it". It sounds more like Facebook handed over some documents that they were legally obligated to hand over and then some lawyers read the documents.

If you're curious about the evidence, most of the court documents are public and available online, although the smoking gun on the percentage appears to be sealed.

> What's your stopping distance at that speed?

Faster than my car or motorcycle. And that's just using the regenerative braking of the motors. I've never tried a full-on-drop-anchor stop with both the mechanical disk and regen brakes.

Do you have a heavily modified scooter or something otherwise unusual? I looked up braking tests for cars, motorbikes, and e-scooters, and couldn't find a single e-scooter that's been tested that has the same stopping distance as a normal car let alone a sports or a motorcycle.

I don't think this is super surprising considering the relative size of the contact patch and what the "tire" is made of on most scooters.

In another comment, you said you have a Xiaomi scooter. In this brake test (https://www.zdnet.com/article/mi-electric-scooter-review-com...), it has a 13 foot stopping distance from 12.4mph even when using the disc brakes. This is similar to what other scooter tests show.

NACTO's suggestion for a conservative car stopping distance is 11 feet from 15mph, not including reaction time, which is the same as in the scooter test (https://nacto.org/docs/usdg/vehicle_stopping_distance_and_ti...).

How are you getting your scooter to stop more quickly than a motorcycle?

It seems pretty good. Twitter has some talks on this where they claim signifiant performance improvements: https://www.youtube.com/watch?v=PtgKmzgIh4c.

I've heard that the Twitter JVM team has a road show where they've talked to some other large Scala users about the performance improvements. Initially, people are highly skeptical of the claims, but after trying Graal on their internal workloads, they generally see similar results.

Here's a paper which has some explanations for why you might expect Graal to improve Scala performance: http://aleksandar-prokopec.com/resources/docs/graal-collecti...

I hope Twitter makes a ‘retro’ mode where I can just see the actual tweets of only the people I follow in semi-chronological order. With the option to turn off retweets for certain users.

I think both of these exist? There's a way to turn off retweets for particular users (it's in the same menu you'd use to mute or block someone) and there's an option to show a plain timeline (uncheck "show the best tweets first").

The plain timeline option doesn't seem to consistently work on the mobile app, but many (most? all?) alternative Twitter clients show you the plain timeline.

CloudHarmony used to track this at some level for free, but it looks like you now need to sign-up or pay to get more than 1 month of history?

The last time I looked at it (back when it showed more info for free, IIRC), AWS had the best uptime of the three big cloud providers, with Azure in 2nd and GCP in 3rd.

IIRC, the memorable thing was that, shortly afterwards, the head of Google Cloud made a big announcement that CloudHarmony showed that GCP had the best uptime when CloudHarmony showed that it actually had the worst. Google was calculating this by computing downtime = downtime per region * number of regions, but at the time, Azure had ~30 regions and AWS had ~15 vs. ~5 for Google and if you looked at average region downtime or global outage downtime, Google came out as the worst, not the best.

No, the graphs are highly misleading because they autoscale the y axis to the highest point and the they do very loose string matching to detect errors.

If you click through to something not on Google cloud, you see moderately elevated error rates (e.g., Instagram is up by 4x) but if you click through to something actually on Google, you see very highly elevated error rates (e.g., roughly 50000x for Snapchat).

If you read the "error reports", they actually report that Instagram isn't down (same for Twitter if you check Twitter). The error report detection seems to be just string matching. Here's an actual "error report" from downdetector that's the caused of allegedly elevated error rates:

my twitter timeline: why isn’t snapchat working? anyone’s snapchat not working? snapchat’s being dumb. rip snapchat.

The Twitter "error report" is literally a report that Twitter isn't down.

If you look at the video, you'll see that "Software Should Be Perfect" is the title slide of the talk (you don't even have to click play, it's there at the start). And then the first words out of the speaker's mouth (other than a sound check) are "I'm going to try to convince all you that software should be perfect".

The "original title" you're referring to is what the person who uploaded the talk to youtube titled the talk, not what the speaker titled the talk.

Microsoft is switching from offices to open office plans. Buildings with offices are slowly being remodeled to open plan.

My first team started off two-to-an-office (unless you had something like 5 or 6 years of seniority, in which case you'd get your own office), but they moved to open offices when their building got remodeled.

Am I missing part of the argument here? At the beginning of section IV, Scott argues that stereotypes can’t explain the gender gap in engineering because of the differential rate at which people major in math vs. engineering (45% women in math, 20% women in engineering). He says:

Might girls be worried not by stereotypes about computers themselves, but by stereotypes that girls are bad at math and so can’t succeed in the math-heavy world of computer science? No. About 45% of college math majors are women, compared to (again) only 20% of computer science majors. Undergraduate mathematics itself more-or-less shows gender parity. This can’t be an explanation for the computer results.

Later, he introduces the thing-people interest spectrum and makes a case that it’s inherent. He then breaks down the gender gap in various medical fields in way that’s suggestive that the differences are due to the thing-people idea.

But, how does that apply math vs engineering or math vs. programming? Or for that matter, programming vs. chemical engineering or electrical engineering vs. chemical engineering? During undergrad, my recollection was that there were proportionally fewer women in electrical engineering than in chemical engineering in the classes I saw, and some quick googling seems to bear this out. If math vs. engineering is a mystery that can't be explained by streotypes, it also appears to be a mystery that can't be explained by thing-people. Although I think it's a long stretch, maybe you can argue that computers are more "thing-like" than math, but I don't think you can really push that argument through to explain the relative ratios in CS, math, EE, CivE, AE, MSE, and ChemE? BTW, the reason I think it's a stretch is because you could also argue that computers are more "people-like" than math, so you could flip the arugment around if the ratios were reversed. For EE vs. ChemE, maybe EE rates are depressed because there's a lot of cross-over between EE and BME classwork and BME is arguably more people-like, so the would-be EEs go into BME, but if there's crossover, maybe that makes EE more BME-like and therefore more people-like. I don't think you can give an explanation that's much stronger than a just-so story for some other set of observed ratios.

Sure, you can pick a subset of fields where thing-people appears to explain the variance[1], but you can also pick a set of fields where it doesn’t appear to explain the variance. Scott seems to view a set of counter-examples as a knockdown argument against stereotypes. But then why doesn’t this other set of examples invalidate the thing-people explanation he argues for? Why can’t you apply the exact same line of reasoning he applied to stereotypes to thing-people?

This line of reasoning seems internally inconsistent to me. Am I missing something that would make this line of reasoning consistent?

One line of reasoning is that Scott is merely rebutting someone else's argument and that he therefore doesn't need to explain what's going on and he only needs to explain why the other explanation is wrong. But in that case, there isn't a need to bring up thing-people. It seems like it's been brought up because Scott believes thing-people has more explanatory power than sterotypes and Scott is making a positive argument about thing-people, not just knocking down someone else's argument.

[1] Even within the fields he picks, one example he gives is the rate at which women go pediatrics at a higher rate than any other specialization he lists other than ob/gyn. But why is the rate in pediatrics so different than in psychiatry? He argues that they're both "people" fields, which sounds reasonable. But there's one field is 25% male and the other is 43% male. That's almost the same as the math/engineering difference he cites earlier with the genders flipped. Is dealing with babies somehow more people-like or less thing-like than talking to adults? It's certainly a strong stereotype that women are more interested in babies than men, but this is cited as an example of the explanatory power of thing-people, not of the explanatory power of stereotypes.

And the other other half are undervalued then.

You must be responding to the current HN title and not anything from the actual paper? This comment doesn't make any sense otherwise.

The paper literally has an entire section titled "All unicorns are overvalued" (section 4.2) and figure 3 shows the distribution of overvaluations given their methodology.