HN user

ddrum001

171 karma

VP of Engineering at Insight (YC W11)

Posts35
Comments44
View on HN
blog.insightdatascience.com 7y ago

No one knows where America’s helipads are, except this neural network

ddrum001
1pts0
blog.insightdatascience.com 7y ago

Using NLP to gain insights from employee review data

ddrum001
1pts0
blog.insightdatascience.com 7y ago

How to build your own CDN with Kubernetes

ddrum001
4pts0
blog.insightdatascience.com 7y ago

Show HN: Using machine learning to tackle bias in machine learning

ddrum001
1pts0
blog.insightdatascience.com 7y ago

Optimizing walking routes to keep you in the sun or shade

ddrum001
1pts0
blog.insightdatascience.com 7y ago

How to start automating your data pipelines with Airflow

ddrum001
2pts0
chrome.google.com 8y ago

Show HN: Chrome Extension to find useful topics for products on Amazon

ddrum001
2pts1
blog.insightdatascience.com 8y ago

Gefilter Fish: Finding concise topics from Amazon’s customer reviews

ddrum001
2pts0
github.com 8y ago

Show HN: ElastiKNN – Elasticsearch plugin for scalable image search

ddrum001
5pts1
blog.insightdatascience.com 8y ago

ElastiK Nearest Neighbors: LSH and Elasticsearch for Scalable Online KNN Search

ddrum001
1pts0
blog.insightdatascience.com 8y ago

How to Use Machine Learning and Quilt to Identify Buildings in Satellite Images

ddrum001
1pts0
blog.insightdatascience.com 8y ago

A Newbie’s Guide to Scala and Why It’s Used for Distributed Computing

ddrum001
1pts0
medium.com 8y ago

Personalization: Explain like I'm five

ddrum001
1pts0
blog.insightdatascience.com 8y ago

Insight Data Engineering Fellows Program Expands to Boston

ddrum001
3pts0
blog.insightdatascience.com 8y ago

Using neural networks to detect car crashes in dashcam footage

ddrum001
2pts0
medium.com 8y ago

Sure Ways to Fail at Personalization

ddrum001
2pts0
blog.insightdatascience.com 8y ago

Heart Disease Diagnosis with Deep Learning

ddrum001
2pts0
blog.insightdatascience.com 8y ago

A Newbie’s Guide to Cassandra

ddrum001
119pts24
xyz.insightdataengineering.com 9y ago

Show HN: Interactive map for architecting big data pipelines

ddrum001
144pts29
blog.insightdatascience.com 9y ago

Using machine learning to predict baseball injuries

ddrum001
1pts0
blog.insightdatascience.com 9y ago

Automated ETL with Kafka Connect

ddrum001
4pts0
blog.insightdatascience.com 9y ago

Engineering features to automatically assess image quality

ddrum001
1pts0
www.youtube.com 9y ago

The Trouble with the Electoral College

ddrum001
3pts0
blog.insightdatascience.com 9y ago

Graph-Based Machine Learning: Community Detection at Scale

ddrum001
3pts0
blog.insightdatascience.com 9y ago

A Day in the Life of a Data Engineer: Yelp

ddrum001
2pts0
blog.insightdatascience.com 9y ago

A Day in the Life of a Data Engineer at Yelp

ddrum001
1pts0
blog.insightdatascience.com 9y ago

Scheduling Spark jobs with Airflow

ddrum001
1pts0
blog.insightdatascience.com 9y ago

Making it scale at DoubleVerify

ddrum001
1pts0
blog.insightdatascience.com 9y ago

Monitoring Your Flask Application on Kubernetes with Prometheus

ddrum001
2pts0
www.insightdataengineering.com 9y ago

Anatomy of an Elasticsearch Cluster: Part III

ddrum001
3pts0

Fair, perhaps the cursive analogy is too strong, though I'm sure that cursive still has a role in research of historical documents. IPA still has an important role in specialized situations, but I don't quite think it's something that should be known for the general public (e.g. the average wikipedia reader).

I guess I'm an anomaly in that I read far more books than watching movies, but I agree that most consumers will want a variety of music and shows, unlike books. However, isn't the price point ($35/month), which is 4x more than Spotify, roughly in line with the price for a single book, which is roughly 4x an album.

I guess the main difference is not that the price ratio between subscription and individual items is off, but rather that most consumers don't want 1 book a month. It's a bummer that they don't give people an option, but I'm still a fan of Safari...and I hope this move will mean they drastically improve the Safari app, much like Netflix has doubled down on Streaming now that they don't do DVDs.

Fair point - the set of technologies is based off the teams we work closest with, which admittedly have a bias towards open source and Linux. So far, our map is far from comprehensive, so appreciate the suggestions (exactly what we're looking for by show HN).

To that point, just added CosmosDB, and plan to add others soon.

We don't have a fixed syllabus since we don't operate like a school per se. Instead, the Fellows choose a project and have to make engineering tradeoffs about which tools to use. With that said, most of the past participants focused on distributed systems like Hadoop, Spark, Flink, Kafka, and NoSQL DBs. You can check out past projects on the blog (http://insightdataengineering.com/blog/) and the past Fellows are here (http://www.insightdataengineering.com/fellows.html)

I'm from Insight - our program is kept free for our Fellows because the companies that we partner with pay for the Fellowship. Rather than classes, we believe that the best way for advanced engineers to learn the detailed nuances of these tools is to use them, so the concept is to learn them by building a data platform. You're right that it's really difficult to understand distributed systems in a few weeks, but our Fellows already have several years of programming, so they have been able to pick up a general understanding quite quickly.

Very exciting news if it holds! Can someone discuss how close the quasi-polynomial complexity is to P - that seems to be the really interesting detail that gets covered less.

An interesting part that the article didn't discuss as much was how to distill all mistakes into the finished product. It's one thing to say artists who make 50 lbs of ceramic pots have a few good ones, but how often are those pieces presented as the best by the artist, or iterated on.

In other words, deciding which projects to continue pursuing and which should be scrapped entirely is non-trivial.

A little passive aggressive (better than the legally aggressive approach of Axel), but pretty interesting. Not sure if I would qualify the eschewing of an AdBlock work-around as Big Brother, but the legal precedent is worth discussing.

Indeed, the script works quite well for Spark - but we also wanted to provide a guide for those who would like to further understand the config files, especially for those who want to 'tinker' with these down the road (e.g. adding nodes to the cluster). Also, we find it helpful for learning to set up Spark on other platforms, and setting up other systems that don't have ec2 scripts yet.

2 Pi or Not 2 Pi? 11 years ago

I don't think there is much to discuss here, tau=2pi is clearly the more sensible constant for almost every use case. The only question is whether we give in to conventions, or pay the one-time technical debt of switching.

We (in the US) may never convert to metric, but the much smaller community who works with pi regularly should know better.

That was a great story, and it speaks to the fact that tropes like prison escape, bank heist, and marooned on an island appeal to some human instinct