HN user

Bartweiss

12,328 karma
Posts2
Comments3,472
View on HN

in both cases someone could accidentally optimize the test

I think this is what I disagree with.

The water heater story is about a viable-for-market design which also optimized for the test. The equivalent for a car emissions test might be optimizing the transmission to reduce emissions at the specific speeds which will be tested. Those speeds could be sweet spots of the engine curve by accident, or they could be planned that way. I don't think that's necessarily right, but it's within the bounds of "natural" design for the product.

Instead of doing that, VW submitted something for testing which was fundamentally different from what went to market. Rather than being misleading, the test results were fundamentally irrelevant. Creating two completely different modes of behavior isn't something you could do by chance, and it means there's no real limit on how badly they could cheat.

This is omnipresent even where regulators aren't involved: every graphics card benchmark out there is 'manipulated' relative to real world performance. At this point it's so universal that I don't think anyone is even fighting it - as long as everyone games benchmarks roughly the same amount, the relative scores stay usable.

Your point about fairness and passive design is the one that makes me view these cases differently also. In the anecdote, the product being tested was the same one being sold, and there's no sign the heater was worsened to improve test performance. The designers just picked the best-scoring option among some reasonable configurations. (Frankly, once they noticed that issue, what were they supposed to do? Pick the worst-scoring, or pick the spec out of a hat?)

In the VW story, the test-bench vehicle was fundamentally different from the market vehicle, and the road version was designed to behave worse on the metrics to get other gains. I happen to know someone who bought a diesel Jetta specifically because it was more eco-friendly than other options, and I think he'd draw a clear line between tuning for test metrics and VW consciously lying to their buyers.

I do think that manipulating a purely instructive measure is less extreme than manipulating a compliance test; consumers can seek alternate tests and reviews, but the state emissions test has special status even if a dozen other tests give a different result. That said, I believe Energy Star ratings affect tax rebates and electric bills, and they're required to be printed on products - so that's not really an arbitrary test.

There are other differences here too, I think. The water heater trick is passive manipulation that stays in place at all times, which limits how far from "real" performance it can get. And per the story, it seems more like "teaching to the test" than "cheating". That is, Volkswagen consciously moved away from the mandate outside of testing. The water heater was (potentially) as energy-efficient as they could design, with the test score manipulated on top of that.

None of that makes it harmless - if "as good as you can make" doesn't hit standards without manipulating them, that's still a problem. But I do find it less galling than "intentionally worsens emissions outside the test bench".

Crash test dummies have basically this problem also. They're designed for realism in certain very narrow ways, and then the very small number of approved dummies are used for testing car safety.

The industry has made a bit of progress, surprisingly unprompted by regulations - female and child dummies came into circulation before they were required in tests. But overall, testing is still run against a tiny handful of body types which move 'realistically' in only a few regulation-guided respects.

Greatest Java apps 6 years ago

Citing Maven also feels a bit circular. It's an important Java application, but being a build tool it's only because there's lots of Java out there to build.

Minecraft and a lot of the other apps are terminally impressive, so it's easier to justify the ecosystem that produced them.

Wait, which countries are we referencing outside of those three?

Thailand looks straightforwardly exponential so far and has fairly heavy mask use, agreed. But Singapore, Taiwan, and arguably Malaysia seem too early to call: they're still plausibly on either of a European curve or South Korea's ramp-then-flatline.

Vietnam, Cambodia, Laos, Mongolia, and Burma all seem to be below the line for meaningful data. And Hong Kong isn't broken out. So I guess my questions are: do Indonesia and the Philipines have "mask cultures" to a level comparable to South Korea and Japan, and are their testing regimes wide enough to rely on those curves?

I don't know the answer to that. And I agree that the "masks work" graph/meme circulating is questionable. But unless I'm missing something/somewhere, this data just looks like "too soon to call"?

you had to be logged in to the web interface already with another account

Obviously I don't know specifics, but if this applies to any router which has multiple tiers of login then it could be a pretty serious problem. I suspect that might be true for routers designed specifically broadcast multiple networks (e.g. school or shared apartment-building routers)?

I don't think this is uncharitable at all. I'm sure Kinsa has made a good effort at controlling for testing frequency, and I'm sure it's helped. But there's no reason to think the dynamics of COVID-motivated testing are the same as for flu-season, or new-buyer novelty, or anything else.

And more importantly, how could we know if it is? That's not just a Kinsa problem; we see this over and over again with peer-reviewed studies that "control for" certain factors like socioeconimics or health history. They're inherently limited to controlling for what they know about, and it's never perfect. Often, the entire effect is from an undiscovered variable. Take, say, the widely-promoted study finding that visiting a museum, opera, or concert just once a year is tied to a 14% decline in early death risk. The researchers tried to control for health and economic status, then concluded "over half the association is independent of all the factors we identified that could explain the link." [1]

Now, what seems more likely: that the unexplained half is from the profound, persistent social impact of dropping by a museum or concert once a year? Or that some of the explained factors like "civic engagement" can't be defined clearly, others are undercounted (e.g. mental health issues), and some were missed entirely?

I suspect Kinsa did much better than that, because they're not trying to control for such vague terms. But I think "even after controlling for" should basically never rule out asking "what if it's a confounder"?

[1] https://www.cnn.com/style/article/art-longevity-wellness/ind...

Good point.

The TARP bailout in 2008 involved buying a ton of stock from troubled companies, but it was sold back to them as soon as they could buy the money back. And this will be the second bailout for a bunch of airlines.

So one of the most interesting ideas I've heard is that we shouldn't nationalize things by fiat, but when TARP-style bailouts happen, the government should just keep the stock, at least for a while. If it really was a one-off crisis, the shares are a good investment. But if it's a failing business, or one paying dividends and then looking for handouts, it's not just a money sink.

There's also a fairly good argument for this in the line of trains and highways. Planes aren't physically trapped on one course, but pretty much every nation heavily regulates who can fly where, when. Airports are often state-controlled, and even private ones need state approval to add new runways or flights.

What we have now is one of the ridiculous "private non-market" arrangements. When airlines in Europe fly empty planes to stop the government from taking their flight slots away, that's not the fault of the companies, but it's also not a functional market we should expect efficiencies from.

I'm not a fan of "regulate markets into dysfunction then nationalize them", but if the fundamental restraints on travel are too severe to let the market function freely, privatization stops making much sense.

The third option is "because they don't want to be blamed for model error". Governments aren't necessarily competent, but you can try to get them to understand 5%/95% confidence intervals, at least in hindsight. If you publicly release a prediction, and then the real outcome is the 10% confidence line, you're probably going to be yelled at for being wrong regardless of the error bars.

Of course, if the center of the prediction is horrifying, "people don't understand confidence intervals" then becomes a case of avoiding societal breakdown.

It's been fascinating to see the rise, fall, and rise of digital watches among techies.

I remember 1990s Dilbert having an entire storyline about the engineers getting into a calculator-watch arms-race. In real life, it was pretty common to laugh about how a $50 digital Casio could do far more things than a Rolex.

By about 2010 (or perhaps even by the iPod Touch or Palm Pilot), I stopped hearing that. Watches had lost all of their unique functions to smartphones, so their raison d'etre was either "rugged and cheap" or "jewelry" and calculator watches almost vanished.

Circa 2015, we get Pebble gen 2, Apple Watch, and Fitbit Blaze: smart watches have phone integrations, fitness tracking, and don't look like hell anymore. Since then, they've increasingly aimed for design good enough to wear with a suit; the Galaxy Watch is always-on and analog.

These days, I see two splits among watch-wearing engineers: smartwatch vs not, and practical vs decorative. So the result is quadrants like:

Fitbit | Galaxy Watch

Casio | Longines

The CEO will be under tremendous pressure if he/she tries to optimize for a 1 year timeframe (for example) as opposed to quarter-by-quarter.

Its bizarre to talk to well-meaning execs (even below C-suite) at public companies and hear them overtly say this. "Well we know X and Y are sound investments for the company's success, but it's a question of finding a way to sell something that long-term without tanking our stock price."

I try not to cry market inefficiency without good evidence, but "shareholders promote good corporate governance" starts feeling pretty bizarre when the people running a company describe shareholders like corporate raiders encouraging them to destroy value for a quick payout.

I wish boards can come up with a compensation structure for execs which optimizes for long term.

For all the talk about "when founders should get out of the way" and "what makes a good founder doesn't always make a good CEO", it's interesting to see that research still finds companies with founder-CEOs performing substantially better. Higher share prices (which might stem from overconfidence), but also better long-term financials, more R&D spending, more influential patent filings, etc.

And that doesn't necessarily mean founders are super-geniuses, exceptional managers, or even unusually attuned to their market. They get less of their salaries in cash, hold options and stocks longer, and vary their behavior less in response to compensation structure. (Also, they often hold so much stock they can't sell in full without panicking the market.)

So it really does look like we just haven't found a good way to compensate non-founder CEOs: their behavior is extremely responsive to their compensation, but nobody has found a scheme that makes them act long-term to the degree a founder would.

For any business with decent size, absolutely. There are a thousand ways to claw back options, and the reason they don't get used is that doing it even once would make hiring practically impossible.

For a small enough company? It falls in the same category as "diluting out of one guy's shares" - bad morals and bad business, but it still happens.

Not only is an easily discovered public event poor leverage, it becomes much worse leverage if it comes up in an interview.

When companies (or governments) try to manipulate employees, they frequently rely on some kind of willful ignorance. Wells Fargo is a great example: they set impossible performance targets and turned a blind eye to fraud, then fired and blacklisted whistleblowers - ostensibly for knowing about that same fraud!

If a shady employer wants leverage, even public events can suffice as long as they can claim ignorance. For example, most stock option grants are immediately lost if you're fired, but even at-will employment can't be terminated specifically to deprive someone of their options. So an employer might give a generous options package, then "discover" the IG video and use it for dismissal at just the right time to prevent a profitable exercise. But if that video comes up during hiring, it's no longer a plausible reason for later dismissal, at least without committing perjury regarding the interview.

I can't even work out a scenario where "lots of people know about this including us" is an effective way to manipulate someone.

I notice this pattern all the time in guides to "polite" workplace communication. Their examples are hypothetical, so they look at how positive something sounds without considering the underlying content, or go even further and change content to improve tone. The advice looks good on paper, but using it when there's an actual task at hand might just sound sarcastic or disingenuous. The worst example I've ever seen was something like:

Instead of "I need that report by the end of the day", try saying "I really appreciate you working to get that report out soon, it's a big priority right now!"

That's absolutely insane, because those are two completely different statements. The second one sounds less demanding because it's not the same request. So the tip isn't positive communication advice, it's either a schedule rework or failing to convey a deadline.

As for this specific example:

By adding an emoji below, it's clear that the sender is embarrassed to make this last-second request, and isn't trying to come across as sarcastic, rude, or overbearing

That wasn't clear to me at all. If you type in "embarrassed", Slack will only suggest :flushed:, although I'd also have understood :sweat_smile:. I guess the monkey was meant as "I'm hiding my face with shame", but Slack calls that emoji ":see_no_evil:", and at first glance it seemed like "I'm trying to not to look over your shoulder, but is this done yet?". If the problem is "making a last second request", there's no particular reason that emoji are the best way to address it - one example simply has more content than the other. So I like your direct phrasing, and I might add:

Hi <name>, will you be able to have the report on X ready by <time>? I'm sorry it's such short notice, thank you!

Eh, it probably buffers against overreaction, especially when a correction in fundamentals is being mixed with a reaction to new pressure.

But this is still a good point: if the market really is overheated then short-term monetary policy won't change that, and we can expect a lasting hit regardless of how disease issues play out. And it's not necessarily going to be obvious what's market movement and what's disease-related; I wouldn't be surprised if some over-hyped companies seize this as a chance to lower guidance faster than they normally could without spooking investors.

This is why the whole idea of "in the public interest" exists.

If a reporter received these same recordings in the mail, they would quite likely publish them. If they received a recording of a random person discussing their medical concerns, publishing that would be an outrageous breach of ethics.

(Hence the Gawker/Thiel debacle also. When Ted Haggard was caught having gay extramarital affairs, it was considered fit for publication because he was an evangelical preacher fighting against gay marriage. When a random private individual is outed, its not a public interest matter and can be libelous even when accurate. Thiel fell somewhere in between under both legal and journalistic rules, so we got a debate.)

I'm pretty baffled to see the parent comment imply that private discussions between politicians should inherently be kept secret. We could discuss specific news stories, reporters who violate attribution rules, and whether Varoufakis was bound by privacy laws or Eurogroup confidentiality rules. We could even argue the publication is in the public interest, and yet makes Varoufakis unfit to serve by destroying his ability to function with trust.

But just as you say, treating "that's a private discussion" as the end of matter would excuse Watergate also.

This is a novel and important result in antibiotics. It's also a proof-of-concept for using ML to produce vital drugs with novel mechanisms, rather than incidental alterations or discoveries in noncompetitive spaces. It might be an incremental speedup or computing-power advance in ML drug discovery also, but it could equally just be the result of a lucky break or a particularly large lab-test budget. (In which case, "why didn't someone do it already?" is closer to asking why nobody else bothered to win the lottery.)

It's not a major theoretical advance in ML drug-discovery techniques or the first big step in ML drug discovery. It's certainly not the invention of ML drug discovery or neural nets as an ML technique, both things I've seen implied in news stories on this work.

This is attention-worthy, absolutely. (I'll leave "publication-worthy methodology" to experts.) But it's newsworthy on actual merits, as a drug breakthrough and a demonstration of an increasingly-important technique. So I share the frustration when lazy or confused reporting implies this is the same style of ML-theory breakthrough as CNNs, Transformers, or even neural nets themselves.

I think the criticism is that it's not obvious whether success here was a function of improved performance, expanded throughput, expanded testing, or sheer luck.

Chess engines have clearly improved in both design and computing power over the years; doubling an engine's resources or pitting a new engine against an old one produces straightforwardly better play. But the drug-discovery technique in use here may not be "playing better" in terms of producing higher-quality predictions.

To extend the chess metaphor:

- Deep Fritz is a stronger player Deep Blue even with 4% as much computing power. This story does not appear to be an algorithmic breakthrough of that source.

- Deep Blue lost to Kasparov in 1996, then beat him in 1997 with double the computing power. That's a clear improvement in play, but not an improvement in efficiency. This story might represent such a change, modelling more prospective drugs to test higher-confidence candidates.

- If an AI that can only win 2% of games against humans plays 10 games, it has an 18% chance of beating someone. But over 100 games, it has an 87% chance of a win. This result might be a team with a larger testing budget claiming the 'first win' without any AI-side improvement.

- If a dozen grandmaster-level chess AIs play GMs, one of them will have to get the first win against a human. Labeling this result a 'breakthrough' in AI terms might be outright publication bias among equivalent projects.

As far as the drug, none of that really matters, except that efficiency improvements would have more potential to increase drug discovery. The drug itself is still useful, and the discovery is a proof of concept; in 1980 no possible computer would have beaten Kasparov. But this is being hailed as a breakthrough in AI in seriously questionable ways. The BBC article, for example, managed to imply that this specific project was novel and important for using neutral nets to produce a significant result.

How am I confident that this is realistic if you literally say its generated?

This is a particularly good question since it's recently been shown that even neural nets trained on real data often pick up substantial, predictable dataset biases.

Practically every single-dataset-trained CNN seems to pick up stylistic quirks in the photos or labels it's trained on. The most visible result is that the CNNs perform better on same-dataset test examples than they do in the wild, sometimes vastly better. More startlingly, it's possible to work backwards from this: the training source of a "finished" CNN can be discerned by looking for certain types of error, and adversarial examples can be predictably constructed based on training source.

Tagged imagesets undoubtedly have stronger and harder-to-remove 'fingerprints' than text data like addresses, but I'd be shocked if the problem was nonexistent for text. My first reaction to "synthetic sensitive user data" for ML is to worry about winding up with systematic errors coming from the generation scheme.

A related interpretation: your debugging tools need to be at least as good as your programming tools. Ideally, better.

Debugging a K8 cluster with print statements is hopeless, but if the cleverest code you can write in a dumb editor is going through a good test suite and profiler, you might be fine. How many as-clever-as-possible optimization tricks have been made viable by Valgrind?

had you built a tractor for a company they would definitely want to know why you want to build a brand new tractor when the old one is working fine.

You're not wrong about the cost and the need to justify it, but I think you might be underestimating how often this happens to physical products. Lists like "the worst cars of all time" are full of needless reworks, frequently driven by some executive's desire to fully "own" a product. I suspect the biggest difference is just that as the cost of throwing away a working design becomes more expensive and visible, the decision to do so slides higher up the ranks.

Truly gnarled legacy code can be almost impossible to understand just by reading, so that truly grokking it requires writing something in the codebase. (Mind, that's necessary for understanding but not sufficient.) And so at a certain point the question becomes - if you can spare the time to go beyond one-off hacks, why isn't that dev time going into adding test coverage, simplifying, etc?

There are still times when knowledgeable legacy maintenance is the right answer, I think. In particular, the Catch-22 of systems built without the capacity for online updates which can't be taken down to add that capacity. I've heard horror stories of the deeply-understood hacks that keep telephone networks running, because they're ancient monoliths contracted to >99.9% uptime. But in general, the argument for not rewriting legacy code is that it takes too long, or it's going to be expired eventually, and so relying on the original author or hacking in changes is more efficient. It seems like "don't rewrite it, but repeatedly spend the time to bring new people up to full expertise" has a very narrow window where it's the right choice.

Why not use this power to secrete chemicals that help you gain more control over yourself?

If you haven't tried it yet, you might be a fan of Nancy Kress' Beggars in Spain. Instead of adult use of nootropics, it deals with prenatal gene editing taken to the extreme: the first generation of the "Sleepless", children who don't need to sleep and can function 24/7 at peak of their (genetically enhanced) intellect and mood.

And so of course the questions become economic and social: what place will the world have for 'Sleepers'? What do the Sleepless, who didn't choose their condition, owe to those without their advantages? And what are the ethics of choosing to have a child with or without those benefits, since the choice can't ever be made with their consent?

(David Brin's Kiln People is perhaps less society-focused, but also fascinating. That one centers on the ability to create and recover memories from short-lived 'clones' - how do you build an economy when the world's best heart surgeon can operate on every patient?)

And even the ones who do practice decent anonymization are generally contributing to the problem just by holding a lot of data.

Lots of companies are content to stop at "our data can't be linked back to a person's identity", which doesn't prevent building a uniquely-identifying user profile. (e.g. via browser fingerprinting, plus enough metadata to associate a user's computer and phone accounts.) Even if they do better than that, its typically "our data is not uniquely identifying in isolation", which still isn't enough. If your differential privacy model says that these four pieces of data have a specificity of 10,000 possible individuals, that's a good start. But if someone with an individual's PII and three of those keys comes looking, they can still narrow down information about the fourth value from your aggregates.

And even if no one screws up, what happens when someone queries a half dozen differential datasets for different subsets of a uniquely identifying key? It's something like the file-drawer problem, where one researcher hiding bad data is malicious, but a dozen studies failing to coordinate produces the same result innocently. If outright failures to anonymize become rarer, cross-dataset approaches become more rewarding.

Which also begs another question about the lines around 'fake'. If you put a quote next to a picture of someone who didn't say it, is that deceptive? "Well sure, the entire point of 4chan's stunt was to deceive people".

Alright, fine. What if the attribution is just plain mistaken in a misleading way? What if you put an inspirational quote next to someone it applies to, instead of the speaker? If the speaker couldn't possibly have said it, or the quote is famously from someone else, so the goal is comedy? Or the countless pictures misattributing that "...fire inside me..." quote which escaped from Fallout: New Vegas? (Is the deceptiveness different when it's shared as a mistake vs meme vs prank? Can your algorithm tell?)

Twitter's put at least some thought into that, thankfully. They talk about "significantly" altered content and whether it's "shared in a deceptive manner". But good lord is that a fuzzy thing to decide or automate.

It's not just the tone - there are plenty of religious people who flatly disagree with the moral claim above.

"The baby is going to die, so this isn't murder, and it's going to suffer, so this is mercy" is not necessarily a background assumption. For some people, those claims are the debate. Deontologists actually exist, and some of them sincerely believe in souls, and eternal damnation or paradise. To them, the doctor is committing a mortal sin, and for the infant any life long enough for baptism can mean the difference between limbo and heaven. Whatever Northam's intent, they would genuinely find his position morally similar to any other infanticide.

I don't intend to argue for that position, but I think it's important to realize how quickly "simple" rulings on misinformation or malice can become judgements on entire ideologies. If Bentham, Kant, and the Pope would all disagree on whether a choice is ethical and humane, perhaps it's not a matter to be settled by Twitter moderation rules?

My first thought was that scope creep from facts to implication is what broke down trust in Snopes.

Per their blog, it appears they have several tiers of action but what standard label. I imagine that the first time people see "false" next to a clip that's real (and flattering to their opinions, perhaps), they're going to discount the warning pretty sharply.