HN user

CountBayesie

344 karma

I blog about probability and statistics at http://www.countbayesie.com/

Posts11
Comments6
View on HN

I've long argued that the biggest problem with orthodox NHST for A/B testing is that you actually don't care about 'significance of effect' as much as you do 'magnitude of effect'. Furthermore, p-values tell you nothing about the range of possible improvements (or lack thereof) you're facing. Maybe you are willing to risk potential losses for potentially huge gains, or maybe you can't afford to lose a single customer and would rather exchange time for certainty.

My favored approach I've outlined here[0]. Where the problem is basically considered one of Bayesian parameter estimation. Benefits include:

1. Output is a range of possible improvements so you can reason about risk/reward for calling a test early.

2. Allows the use of prior information to prevent very early stopping, and provide better estimates early on.

3. Every piece of the testing setup is, imho, easy to understand (ignore this benefit if you can comfortably derive Student's T-distribution from first principles)

[0] https://www.countbayesie.com/blog/2015/4/25/bayesian-ab-test...

The type of analysis being banned is often called a frequentist analysis

I find that there is a trend of associating "bad statistics" with "Frequentists Statistics" which isn't really fair. If you found a statistician trained only in Frequentist methods and asked their opinion on experiment design in psychological research they would likely be just as appalled as any Bayesian.

I'm a big fan of Bayesian methods, but the solutions of "we'll solve the problem of misunderstanding p-values by removing them!" is still a problem of misunderstanding p-values! The misunderstanding is the issue, not the p-value.

Yes! Sorry for not being more clear in my comment.

And just to be clear Jaynes is not explicitly arguing for or against the idea of ESP, though he himself does not believe he points out that there are many prior beliefs widely held in science that change dramatically over time. What he is saying if that if your prior belief in ESP is dramatically lower than your belief that people trying to prove ESP to you would deceive you in some way, the ESP is real hypothesis will never gain enough evidence to overcome the "I'm being tricked" hypothesis.

For anyone interested in Bayesian Statistics it is worth noting that ET Jaynes (imho the arch-Bayesian) disagrees with Kahneman's idea that "people reason in a basically irrational way". Jaynes died long before "Thinking, Fast and Slow" was published, so his critique is based on Kahneman and Tversky's early work on the subject.

Kahneman and Tversky's critique of Bayesian analysis is basically: If more data should override a prior belief, then why is it as more data comes in people have increasingly divergent opinions? For example we have 24 hour news media throwing information at us and people only seem to be more divided politically. If we reasoned in a Bayesian manner then our opinions should converge, which they clearly do not.

Jaynes' answer is really fascinating and is covered in the chapter "Queer uses for probability theory" from 'Probability theory: The Logic of Science'. Basically Jaynes' argues that we are never really testing just one hypothesis. He gives an example of an experiment designed to prove ESP, and points out that no matter how low of a p-value the experiment reports, if you have a strong prior belief that ESP does not exist the evidence won't convince you. He argues this is because you actually have other hypotheses with other priors: The subject is tricking the experimenters, there is an error in experiment design, the people running the experiment are intentionally being deceptive etc.

He then shows that if your prior belief in ESP is sufficiently lower than your prior belief in these alternative hypothesis, not only will further evidence fail to convince you of ESP, but will actually increase your belief that you are being lied to in some way. So while Jaynes agrees that these priors may be irrational, our reasoning given new information is completely rational.

I'm a lover of awful puns, and was actually surprised how few people I knew got it! I'm very glad you appreciate it, thanks!