Apparently it is a scientific program that was made into a web app https://www.ufz.de/index.php?en=39156
HN user
aaronjg
PhD Candidate in Evolutionary Biology at Stanford
Former Lead Data Scientist Custora (YC W11)
Email is username @ stanford.edu
This sort of approach is also used in reinforcement learning https://arxiv.org/abs/1702.01182
There is some research on these ecosystems [1], and there dynamics [2,3]. The ecosystem is made up of a photosynthetic algae as a producer, a protozoa as a consumer, and a bacteria as a decomposer, and can be stable for over 1,000 days.
[1] http://ir.obihiro.ac.jp/dspace/bitstream/10322/221/1/Prot.Vo...
[2] http://www.cell.com/cell/pdfExtended/S0092-8674(12)00515-6
[3] http://journals.aps.org/prx/abstract/10.1103/PhysRevX.5.0410...
When is this from? The last paper in the reference is 2010.
I wrote about the problem with sequential testing in online experiment three years ago on the Custora blog [1]. And Evan Miller wrote about it two years before me on his blog [2]. I'm glad to see Optimizely finally getting on board. Communicating statistical significance to marketers is always challenging, and I'm sure this will lead to better decisions being made.
[1] http://blog.custora.com/2012/05/a-bayesian-approach-to-ab-te...
[2] http://www.evanmiller.org/how-not-to-run-an-ab-test.html
Google Cache Link
http://webcache.googleusercontent.com/search?q=cache:kvX1Y03...
Gayle Laakman has a good post about self-publishing, and a lot of the hidden downsides, and why going the traditional route is often better (and more profitable).
http://www.technologywoman.com/2012/07/09/the-dirty-truth-ab...
First of all, great work. It looks like you boosted your conversion rate from 0.19% to 0.43%. Which is a 125% improvement, or with confidence intervals, 55% - 179% improvement.
However, before everybody goes out and puts puppies on their homepages, they need to realize that there are a bunch of things being tested.
Image vs. no image: Is it possible that having any image at all improves the conversion. You should test with other pictures: perhaps some animals, people, nature, and see if the puppy is what makes it work.
Call to action: The 'puppy' version also features a more succinct call to action in "Sign up now" rather than "Start your 30 days free trial." Perhaps this also contributes some of the difference.
Button size: The button size in the 'puppy' version is smaller. Perhaps this has some effect as well.
Length of text: The 'puppy' version has more description of what is involved in the free trial. It says "Pick a plan & sign up in 60 seconds. Upgrade, downgrade cancel at any time." vs. the no puppy version that says "Start you 30 days free trial."
Vertical vs. Horizontal Layout: The 'puppy' version has a vertical layout of the text and button, where they are stacked on top of each other rather than left or right.
So there are at least five different changes made between these two designs. Clearly the second design wins on conversions, but it's not entirely clear to me why it wins.
That approach was discussed by Anscombe, and I wrote up a summary in the Custora Blog. However just because an approach is frequentist or 'ad-hoc' does not necessarily mean that there is anything wrong with it. The bayesian approach requires making assumptions about the number of visitors to your site after you stop the test, which isn't really any less adhoc than picking an error cut off.
http://blog.custora.com/2012/05/a-bayesian-approach-to-ab-te...
I have a problem when companies start claiming personalization to this extreme level. How can they claim that they know that an individual "Gets bored and checks email at 4pm."
They are able to look at their customers and see when they open emails, even report on the average time, but people are so much more noisy than they make it seem.
It's interesting that this attitude of pinning customers to a specific thing is so ingrained in their mentality that they bucket their customers: Johnson only ever drinks water, Aubrey rides his bike every day, rain or shine.
In reality people are complex and multifaceted, and it is important to acknowledge this when marketing to them.
I am mostly critical of claims like '20 lines of code that will beat A/B Testing Every Time.' Multi armed bandits are also not as useful for inference as the frequentist methods that Ben presents in his posts.
That seems to be a pretty sound approach, compared to some of the stuff about multiarmed bandits that shows up here some times. And I certainly expect Noel Welsh to chime in as well.
There are two schools of thought about the approaches to sequential testing, the Bayesian approach lead by Anscombe, and the frequentist by Armitage. I talked a bit about this and outlined Anscombe's approach here [1]. And it is great to see such a nice write up of the frequentist approach and the tables of the stopping criteria
[1] http://blog.custora.com/2012/05/a-bayesian-approach-to-ab-te...
I've spent a lot of time working with pipelining software, first for my last job doing bioinformatics research, and now for handling analytics workflows at Custora. We ultimately decided to write our own (which we are considering open sourcing, email me if you are interested in learning more).
The initial system that I used was pretty similar to Paul Butler's technique, with a whole bunch of hacks to inform Make as to the status of various MySQL tables, and to allow jobs to be parallelized across the cluster.
At Custora, we needed a system specifically designed for running our various machine learning algorithms. We are always making improvements to our models, and we need to be able to do versioning to see how the improvements change our final predictions about customer behavior, and how these stack up to reality. So in addition to versioning code, and rerunning analysis when the code is out of date we also need to keep track of different major versions of the code, and figure out exactly what needs to be recomputed.
We did a survey of a number of different workflow management systems such as JUG, Taverna, and Kepler. We ended up finding a reasonable model in an old configuration management program called VESTA. We took the concepts from VESTA and wrote a system in Ruby and R to handle all of our workflow needs. The general concepts are pretty similar to to Drake, but it is specialized for our ruby and R modeling.
Some more useful links for those interested:
JUG https://github.com/luispedro/jug
Taverna http://www.taverna.org.uk/
Kepler https://kepler-project.org/
This is an interesting example of why randomization in experiments is important. If you allow users to self select into the experiment and control group, and then naively look at the results, the results might come up opposite from what is expected. This is known as Simpson's Paradox. In this case, it was only the users for whom page load was already the slowest that picked the faster version of the page. So naively looking at page load times made the pages look like they loaded slower.
However once Chris controlled for geography, he was able to find that there was a significant improvement.
Moral of the story: run randomized A/B tests, or be very careful when you are analyzing the results.
The question of Best Buy settling has come up. Most settlements involve a clause that prohibits either party from disclosing the terms of the settlement. Since the primary objective of First Round taking on the law suit was to 'teach big businesses a lesson.' A settlement would have been counterproductive for them, even if it would have resulted in a substantially higher payout or significantly decreased legal costs. First Round wanted blood to be shed publicly.
For sources of the numbers, click on the link in the lower left of the page, and a methodologies window will appear.
Or spell the header text correctly
"We will be launching soon Sartup Paradise ! The place to be for remote workers"We use Perfect Audience at Custora, and are really happy with it. Couldn't be easier to set up ad retargeting. And they have excellent service and support. Congrats guys!
That's a really interesting approach. Basically mining a couple of key features from the data set, and plugging it into a supervised learning algorithm. The really cool thing is the simplicity of the features used, just daily activity and daily playtime.
The write up of the methodology is very clear, but I'd love to see some more description of the results. Ninety-five percent accuracy is a pretty bold claim, and I'd love to see some ROC curves to back it up!
Brooklyn, NY Full Time
Custora (YC W11) is a customer analytics tool that helps retailers earn more from happier customers. To be a little more specific, we can point to a single retail customer and paint a meaningful portrait with his data: How much he’ll spend, how often he'll make purchases, what types of products he's inclined to buy, his predicted likelihood of returning, and more. Custora also integrates with email marketing providers and customer support systems to fuel a seamless, iterative flow of insights to actions.
From Fab.com to Etsy, some of the fastest growing and respected names in retail are using Custora on a daily basis.
Who We’re Looking For
We’re looking for a developer to join our core team. Our web stack is Ruby on Rails, and our analytics are done in R. Experience with these technologies is a plus, but we’re open to sharp developers with experience building products for the web in general.
Where We Are
Location-wise, we’re in Brooklyn, NY. We love it. Progress-wise, we’re a YC company from Winter 2011. We’ve recently been featured in the New York Times, GigaOm and BetaKit, and in the last 2 months we’ve had more signups than in the previous 10.
Day to Day Here’s a taste of what happened last month:
Aaron implemented a Dirichlet Latent Class Multinomial to power customer archetype analysis based on customer purchasing behavior.
Martin made dramatic improvements the email marketing part of the product. He made it easier for our clients to launch multiple email tests in parallel, and added four new email providers to our growing list of integrated partners.
Jon and David worked together to completely redesign the interface of the application. We moved from an interface that focused on browsing through dashboards to one that delivers answers to specific questions.
Outside the office, Corey and Dave manned a booth at a big e-retailer conference and developed a Blackjack-style Custora game to play with prospective clients.
What We Offer
Our compensation is competitive with anyone on the market. Since you’ll be a core member of the team, meaningful equity is part of the package. We offer comprehensive health coverage, including a dental and vision package. Lunches are paid for and we usually eat as a team. We do happy hours at least twice a month and play bocce ball competitively (sort of). Our vacation policy is based on trust — take what’s needed and keep the rest of the team up to speed.
Let’s Chat
If you’re interested, apply online at http://www.custora.com/careers
Brooklyn, NY Full Time
Custora (YC W11) is a customer analytics tool that helps retailers earn more from happier customers.
To be a little more specific, we can point to a single retail customer and paint a meaningful portrait with his data: How much he’ll spend, how often he'll make purchases, what types of products he's inclined to buy, his predicted likelihood of returning, and more. Custora also integrates with email marketing providers and customer support systems to fuel a seamless, iterative flow of insights to actions.
From Fab.com to Etsy, some of the fastest growing and respected names in retail are using Custora on a daily basis.
Who We’re Looking For
We’re looking for a developer to join our core team. Our web stack is Ruby on Rails, and our analytics are done in R. Experience with these technologies is a plus, but we’re open to sharp developers with experience building products for the web in general.
Where We Are
Location-wise, we’re in Brooklyn, NY. We love it. Progress-wise, we’re a YC company from Winter 2011. We’ve recently been featured in the New York Times, GigaOm and BetaKit, and in the last 2 months we’ve had more signups than in the previous 10.
Day to Day
Here’s a taste of what happened last month:
Aaron implemented a Dirichlet Latent Class Multinomial to power customer archetype analysis based on customer purchasing behavior.
Martin made dramatic improvements the email marketing part of the product. He made it easier for our clients to launch multiple email tests in parallel, and added four new email providers to our growing list of integrated partners.
Jon and David worked together to completely redesign the interface of the application. We moved from an interface that focused on browsing through dashboards to one that delivers answers to specific questions.
Outside the office, Corey and Dave manned a booth at a big e-retailer conference and developed a Blackjack-style Custora game to play with prospective clients.
What We Offer
Our compensation is competitive with anyone on the market. Since you’ll be a core member of the team, meaningful equity is part of the package. We offer comprehensive health coverage, including a dental and vision package. Lunches are paid for and we usually eat as a team. We do happy hours at least twice a month and play bocce ball competitively (sort of). Our vacation policy is based on trust — take what’s needed and keep the rest of the team up to speed.
Let’s Chat
If you’re interested, apply online at http://www.custora.com/careers
Good point. I helped run the contest, and want add a bit more background.
Robert's answer was in many ways an educated guess. He took some of the numbers, looked at other projects and put it together to get a good guess.
However his motivation for making a guess was spot on. We got a lot of answers that were more thorough. But just as Robert conjectured, these were all relatively similar, and all under predicted the true number of backers.
Robert's prediction was the highest prediction that we got, mostly because he was able to correctly guess that there would be another wave of supporters before the backing period ended.
The problem is that 'me not wanting to ever see your ad again' is NOT quantifiable. Marketers can see the positive impact of QR codes but not the negative.
Thus by their metrics the more prominent the QR code, the more people who swipe it and the 'better' the ad performs. But it's very expensive to quantify the negative impact of the QR code. Advertisers would need to run focus groups, ask people if they saw the ad, what they remembered about it, what their impression was, etc.
So we find ourselves seeing more, bigger QR codes, even though people don't really seem to care for them.
Option 1 will NOT give you the correct answer. You CANNOT use confidence intervals as a stopping criteria. If you do this, you end up running many tests, and then you need to apply a multiple test correction to account for this. Otherwise you run a VERY HIGH risk of picking the wrong result.
I emphasize, because this is a common problem made by A/B test practitioners. For a fuller discussion of the problems, check out the papers by Armitage (frequentist) and Anscombe (Bayesian) on the topic. Or see my summary of the issue here:
http://blog.custora.com/2012/05/a-bayesian-approach-to-ab-te...
Rather than try to mine historical data, run an experiment to pit UCB against Neyman-Pearson inference. For some A/B tests, split the users into two groups. Treatment A is A/B testing, treatment B is UCB.
In A/B testing, follow appropriate A/B testing procedures: Pick a sample size prior to the experiment that gives you appropriate power, or use Armitage's rule for optimal test termination. (Email me if you're interested, I'm happy to send over papers/scan relevant pages from his book). However , it's probably best to use a fixed sample size, as that is what most real life A/B test practitioners use. Picking the sample size can be a bit tricky, but as a rule of thumb, pick something that is large in enough to dectect differences in treatments as small as 1%age point.
In the treatment group B, use the UCB1 procedure. Subject the users to whichever design UCB1 picks, and continue with the learning.
Do not share any information between treatment groups A and B.
Run these tests for a sufficient amount of time over a largish number of clients, and then use permutation tests to determine which treatment, UCB1 vs Neyman-Pearson, performs better.
In all the simulations I've seen, UCB performs simple A/B testing, but it would be great to see some empirical evidence as well.
There are appropriate solutions to the multi-armed bandit problem, and a wealth of literature out there, however this is not one of those solutions.
Here's a simple thought experiment to show that this will not 'beat A/B testing every time.' Imagine you have two designs, one has a 100% conversion rate, one has a 0% conversion rate. Simple A/B testing will allow you to pick the the winning example. Whereas this solution is still picking the 0% design 10% of the time.
For some other implementations check out the following links:
For Dynamic Resampling:
http://jmlr.csail.mit.edu/papers/volume3/auer02a/auer02a.pdf
For Optimal Termination Time:
http://blog.custora.com/2012/05/a-bayesian-approach-to-ab-te...
Y-axis need to start at zero when the data being presented is a bar graph and the bar is filled in. If it is a scatter plot with points being plotted, then it's generally not misleading to shift the axis.
Somewhat minor point, but I take issue with the CLV calculations. Using average retention rates can significantly undervalue your customers [1].
[1] http://blog.custora.com/2011/08/why-average-retention-rates-...
Just sent you a copy of the paper. If you plan to use the result 'forever' then theoretically you would be willing to sacrifice a huge (infinite) amount of suboptimal performance now, so that you get the correct answer in the for when you decide to pick the winning idea. It would be very important to have the correct winning idea, because it is going to run for eternity.
In practice, we don't actually ever run the winning idea for ever. We do website re-designs periodically, we test new ideas, business needs change. So we can pick a reasonable value for k based on these constraints.
Alternatively, you can get better performance by _not_ picking a stopping criteria, and dynamically choosing which homepage to show. As soon as one idea appears to be doing better, you start showing that to more users. By choosing the appropriate adaptive sampling strategy, you can reduce regret to be less than if you have a constant sampling strategy. However, for many people the adaptive strategy may be more trouble to implement than it is worth.
The most important takeaway is to _not_ use repeated significance tests to determine experiment termination time. Either use the Anscombe bound with an appropriate k, or fix the sample size before starting the experiment.
What libaries are you currently using where you would like to have things like this?