HN user

brockf

348 karma
Posts37
Comments96
View on HN
www.strong.io 6y ago

Bias in machine learning: We can do better

brockf
1pts0
www.strong.io 6y ago

Designing, building, and shipping machine learning solutions: Three principles

brockf
2pts0
news.ycombinator.com 8y ago

Ask HN: Who is building something great in Vancouver, BC?

brockf
4pts0
medium.com 9y ago

Leaving academia to start a data science company: Looking back at our first year

brockf
1pts0
www.optimail.io 9y ago

Email optimization 101: How frequently should you email your customers?

brockf
2pts0
www.optimail.io 9y ago

Why you should move beyond A/B testing for email campaign optimization

brockf
2pts0
www.strong.io 9y ago

Predictive Model ROI ($) Calculator

brockf
1pts0
www.strong.io 9y ago

Show HN: My friends and I started a data science consulting firm after our PhDs

brockf
8pts5
www.optimail.io 9y ago

Testing an AI algorithm from concept to production

brockf
7pts1
www.optimail.io 9y ago

Show HN: Automatically optimize your drip email marketing campaigns using AI

brockf
17pts3
www.strong.io 9y ago

Free data science audit: How does your organization's data strategy stack up?

brockf
3pts0
www.strong.io 9y ago

Introducing Optimail: Email marketing powered by artificial intelligence

brockf
5pts1
www.optimail.io 9y ago

Optimail – Email campaigns powered by artificial intelligence

brockf
3pts0
www.strong.io 10y ago

Customer Lifetime Value: How to Avoid Common Pitfalls and Build Smarter Metrics

brockf
3pts0
www.strong.io 10y ago

Predicting user behavior: Using machine learning to identify paying customers

brockf
5pts0
www.strong.io 10y ago

Landing page science: Optimizing the conversion rates of 10k+ landing pages

brockf
3pts0
news.ycombinator.com 10y ago

Ask HN: Do you use the money you have made to market yourself?

brockf
5pts9
www.strong.io 10y ago

Five things every startup should be doing with its data

brockf
4pts0
news.ycombinator.com 11y ago

Ask HN: Where to advertise available desk for hackers?

brockf
1pts0
www.brockferguson.com 13y ago

Why does neuroscience cause people to question the roots of human behaviour?

brockf
2pts1
www.brockferguson.com 13y ago

Facebook vs. the iPhone: How (In)definite Articles Play a Role in Branding

brockf
1pts0
news.ycombinator.com 14y ago

Anyone have any marketing or SEO expert recommendations?

brockf
1pts1
www.heroframework.com 14y ago

Avoid Recurly et al. integrations. Use Hero + the eCommerce add-on

brockf
1pts0
electricfunction.theresumator.com 14y ago

Job: Looking to hire a part-time PHP developer and support tech - hackers only

brockf
1pts0
www.heroframework.com 14y ago

Hero PHP Framework

brockf
3pts0
www.heroframework.com 14y ago

My company re-released our flagship PHP app as open-source

brockf
1pts1
news.ycombinator.com 14y ago

Ask HN: Simple multi-currency bookkeeping on a Mac - impossible?

brockf
1pts2
news.ycombinator.com 15y ago

Ask HN: Is cold calling ever an acceptable marketing practice?

brockf
7pts11
www.cariboucms.com 15y ago

Show HN: Caribou CMS - 14 months of development, one 22-yr old developer

brockf
4pts4
www.electricfunction.com 15y ago

Are you trying to sell a product? Or sell your business?

brockf
2pts0

Most implementations are actually moving in the opposite direction. Previously, there was a tendency to look to aggregate words into phrases to better capture the "context" of a word. Now, most approaches are splitting words into sub-word parts or even characters. With networks that capture temporal relationships across tokens (as opposed to older, "bag of words" models), multi-word patterns can effectively be captured by attending to the temporal order of sub-word parts.

Strong Analytics | Chicago, IL | Full-time Data Scientists, Data Engineers | https://www.strong.io

We help companies integrate state-of-the-art machine learning into their products, internal tools, and infrastructure. We've designed, built, and deployed products in the automotive space, pharma, gaming, retail, tech, and many other verticals.

Requires an advanced degree (M.S./Ph.D.) in a quantitative science and 1+ years applying machine learning to real-world problems.

To apply: https://careers.strong.io/

Survival modeling is exactly what's needed for these situations. It allows you to (a) consider censored data (i.e., active customers who you know stay for at least X months) and, (b) use flexible survival distributions beyond the standard exponential distribution assumed in the typical monthly churn rate calculations.

Source: Run a data science company and we work on a lot of customer lifecycle modeling projects with companies much younger than yours.

A quantile-based confidence interval from bootstrapping can yield a 100% confidence interval that does not contain 0, i.e., with 100% of cases positive/negative. But that does not (necessarily) mean that there is a 100% chance that the new version is better than the old one. Confidence intervals are not Bayesian credible intervals and cannot be treated as such. (That said, making some certain assumptions about the underlying model can in some times allow one to treat nonparametric bootstraps in such a way.)

Data security is hugely important. Here are a couple things we do to deal with it: (1) We cleanse the data of Personal Identifiable Information (PII) as quickly as possible (i.e., before it touches our servers), and (2) We host our databases behind secure networks and follow best practices with regards to authentication, encryption, etc.

We were all drawn to applying the statistical, experimental, and algorithmic approaches we learned in graduate school (and in our spare time) to a range of problems in industry. Every project has a big learning component that keeps things exciting and fresh.

Our first handful of clients all came from our professional network. I've been a developer and consultant for a long time (shockingly, over half my life!) and so, despite selling my companies and heading to graduate school, I had a bit of a network of other founders who knew me and were supportive of the new venture.

A few lessons learned, in brief: (1) Try really, really hard to be specific about what you offer (even when in reality you offer a lot of different things), (2) Write great proposals -- they become the project bible and really help streamline client conversations, and (3) Understanding a client's data and business always takes longer than you'd think.

Author, here. In this post, I review the various ways that we put our email marketing optimization algorithm to the test, starting from simple sim environments in R, to scrappy real-world tests, more complex simulations, and ultimately a private beta with a production app. I hope it helps those thinking of bringing their own algorithms to market, and would love any feedback!

Hey thanks! I'm Jacob's co-founder at Optimail.

We only just launched yesterday, so I think we are still working to find that exact product-market fit. At the moment, however, we're targeting medium- to large-sized businesses that use drip email marketing campaigns, such as onboarding campaigns, lead nurturing campaigns, and retention campaigns.

Thanks for the feedback on the copy! If you are feeling extra generous and have a second, shoot me a PM and let me know what parts you found confusing. We want it to be accessible to non-technical marketers who might be using older software like Mailchimp, etc. (After all, a main benefit of using AI here is that you don't need to get into confusing automation-building stuff like multi-branching decision trees.)

I'm excited to announce the public release of Strong Analytics' new product, Optimail.io. Optimail uses AI to send, manage, and optimize your drip email marketing campaigns. It's a replacement for complex decision trees, A/B split tests, and hours spent staring at your screen trying to make your email campaigns more effective.

After my first software business was acquired, I began my PhD in Cognitive Science at Northwestern University, where I was trained in statistics, research, machine learning, and experimental design. It was a blast! I met some amazing people, and I found the ideas I got to think about every day very exhilarating.

But what excited me even more was the possibility of integrating what I was learning about (machine learning, AI, optimization) with what I'd spend my life to that point building — software that helped people grow their businesses. So, I left academia and, together with a couple of fellow PhD friends from graduate school, have formed a data science development and consulting firm, Strong Analytics, and we are now launching Optimail as our first product.

If you do get a minute to check it out, I'd love to hear any feedback you have!

To your first point, I'm not sure how this helps the sufferers. What they have shown is that there is a different neural signature correlating with a different emotional response. There's no causal link here, meaning that I might still tell someone to just shrug it off. Maybe they can control their neural activity, just like I can control my neural activity related to thinking about elephants by not thinking about elephants.

To your second point, I would offer two responses. First, while being able to measure something is a win in itself, we need to be clear about what they're measuring. They have shown that suffering ailment X increases the probability that they find neural signature X. They do not know the reverse, meaning this isn't going to unlock early diagnostics or anything. It is unclear how discriminating this response is. Second, it's not clear that this is the product/effect of the way their "brain is wired". Perhaps changes in neural activity caused the observations of different neural connectivity. Perhaps some other factor of their experience or biology caused this sensitivity and the visible differences in connectivity, neural activity, etc. We just don't know.

(P.S. Listening to people eat drives me insane. I'm not going to self-diagnose, but I just want to be clear that I'm not criticizing the finding/report because of a lack of empathy for the sufferers.)

(P.P.S. I'm a recovering cognitive scientist who had to hear about a lot of neuropsych findings that all boiled down to, "This part of the brain lights up when we hear/see/do this! Give me another $5mm grant!").

How is this an "explanation" or "cracking" the problem? They showed that an emotional response was correlated with neural behavior... what else could it have been?

More broadly, it's frustrating that neuroscientists reframe genuinely interesting questions ("Why do I get angry when people eat apples beside me?") into boring questions ("What does my brain look like when I'm angry about someone eating an apple?").

Great points. It's definitely more challenging than learning to play a simple arcade game or something, where feedback is invariant and often instantaneous. To address these challenges, we use a combination of (1) heuristics tailoring our RL algorithms to the problem at hand, (2) many converging sources of feedback. Most importantly, as with any machine learning implementation, it works in practice — our AI-driven campaigns beat randomized, control conditions!

At our data science company, we're building a marketing automation platform that uses deep reinforcement learning to optimize email marketing campaigns.

Marketers create their messages and define their goals (e.g., purchasing a product, using an app) and it learns what and when to message customers to drive them towards those goals. Basically, it turns marketing drip campaigns into a game and learns how to win it :)

We're seeing some pretty get results so far in our private beta (e.g., more goals reached, fewer emails sent), and excited to launch into public beta later this month.

For more info, check out https://www.optimail.io or read our Strong blog post at http://www.strong.io/blog/optimail-email-marketing-artificia....

It's great to see more hacker-friendly introductions to reinforcement learning. Like most facets of machine learning, there are so many interesting applications of reinforcement learning (e.g., we're using RL to optimize email marketing campaigns at Optimail), and we'll only find more as more non-academic hackers discover it.

Sorry, I'm not 'looking for work' as in handing out resumés, looking for a 9-5. The acquisitions are in the 7-figures. But I am always looking for new opportunities -- new startups to work with, etc.

I do really like your point about focusing on the operational metrics and processes, instead of the acquisition value.

Grade inflation has really made these inferences even more difficult; for example, the median grade at Harvard is actually an A- (not the well-balanced distribution you hypothesized).

This post is almost entirely inaccurate re: actual statistics. Power, as you say, is the likelihood of a given experiment rejecting the null hypothesis given some a priori sample size and effect size. And one- vs. two-tailed tests do not change effect sizes, or even your estimate of the variability surrounding an effect, but a p-value related to the effect (should you choose to calculate one).

You would think, given their team of "analysts" and "statisticians", that they might have known these basic pieces of statistics.

While I agree that it is unreasonable to expect someone to deeply understand the current state of all of these technologies/problems, an actual "full stack engineer" needs only to be able to learn very quickly about these problems and, ultimately, make things work.

Re: power law distributions. I'm not sure what you mean about finite versus infinite variance. Are you referring to whether you are analyzing a bounded versus unbounded scale? Even if a scale is unbounded, that still wouldn't make a dataset any more likely to power law distributed. A power law distribution would be, however, somewhat more likely to be observed in cases where there is only an upper- or lower-bound (though again, not always... it really depends).

Re: appropriate averages. Your case (metered billing) is an interesting one. I don't see why the mode is necessarily wrong - the next customer is most likely to spend the amount that is currently your most popular amount). However, in order to calculate a mode, you likely would want to bin your amounts into ranges so that $10.11 isn't treated as distinct from $10.12, etc.

You're definitely right about one summary statistic not being sufficient. I would advocate for visualization and summary statistics, with some estimate of your confidence in the estimate displayed visually.

I should also add... you can also use median absolute deviations, standard deviations, or interquartile ranges to identify and remove outliers who you think don't reflect your business's true status. But it all depends on what you want your model to do!

I'm not sure I follow this, or buy into the suggestion of the post.

First, I don't see how it's true that data with a relatively large amount of variance will tend to be power law distributed. Defining what a "large amount" of variance is is tough (it depends on your intuition and choice of variance metric) but there are lots of distributions with considerable variance that are, for example, normally distributed (many more than are power law distributed, as far as I can tell).

Second, if you find that this is misleading your projections, why not just use a different kind of average? For example, if you just want to know, "How much is the next customer likely to spend?", you might use the mode. Or, if you want a more robust average (i.e., less likely to be seriously thrown off by outliers), why not use the median? You can even complement these with confidence intervals if you want to get a sense of their precision.

Like twic already said, you need some indicator to understand what's going on with your business. I think that in many cases, this will be the mean. But if you want something more robust or more practical, perhaps the median or mode might suit you better.

Shopify POS 13 years ago

> Currently it feels they're only focused on things that can bring them more revenue.

That makes a whole lot of sense to me, and would be the only acceptable path to take in the eyes of their shareholders.

But, by talking about the "other factors", you are just re-configuring the group from which one would assess your individual probability of having a heart attack. You're right that group statistics aren't meaningful when the individual is not part of that group but, if that's not the case, then group statistics are the best source of predictive power.