HN user

mrflip

31 karma

Building tools to organize, explore and visualize massive information streams.

Posts0
Comments11
View on HN
No posts found.

The foursquare venue comes in from the firehose, so we have information on venues that have seen recent checkins. Thus the coverage (though imperfect) is best for venues that are most popular.

If that doesn't meet your needs please get in touch, as we'd like to evolve something that does -- flip@infochimps.com

Much easier to configure and install.

* You get Pig, Wukong and Dumbo out of the box. Don't learn Java hadoop, use one of those. * Standardized setup which really helps if you have to seek IRC or other outside help * Reasonable default parameters tuned for each of the EC2 instance sizes. * Spinning up and tearing down a cluster is so much easier: experiment away. EC2-backed instances help this too.

Now of course you're trading the complexity of setting up Hadoop with the complexity of setting up chef and poolparty and EC2 -- but those are far less esoteric; and they either work or they don't.

DrawnToScale, Infochimps and Cloudera are having a [Hadoop on Chef on Cloud] hack day this Monday before the Hadoop Summit. Holler (@mrflip) if you're interested in the hack day, or if you'll be at Hadoop Summit and want a demo.

We've been using cluster_chef for a month+ at infochimps, and it's awesome. Want a throwaway 4-machine DB cluster to pound on? bam. done. Need to shut down two dozen m1.large instances, spin them back up as 30 c1.xlarge to process a massive CPU-intensive job, then put the whole thing away for the weekend? It spot-prices, instantiates, provisions and designates all the nodes, including EBS volume attachment and service discovery. The kind of thing you'd allot a platoon-day for is now more like an intern-morning, most of it waiting for the spot-price bid to come through.

There's largely no such thing as "closed source" data. Many of the restrictions people claim on publicly distributed data are bogus: you cannot claim copyright on a comprehensive collection of facts. http://www.iusmentis.com/databases/us/ http://blog.infochimps.org/2008/04/02/good-neighbors-and-ope... I don't think baseball is cracking down on people making money on this, unless they infringed their (quite reasonable) hot news claims to the real-time data.

Baseball is the leading example of why giving away most of your data is the best use of it. The sport of baseball -- the way it's played on the field, the way players are scouted and trained, and the way it's enjoyed as a fan (Moneyball? Fantasy Sports?) -- have been revolutionized by amateurs making use of free open data.

If you give out the great bulk of your data, people will be enhancing it with metadata, building tools on top of it, and most importantly connecting it to the rest of humanity's knowledge store and mining it for connections you'd have never conceived. Giving out "up to last month" or "daily intervals" will grow sharply the market for "real time" or "second-by-second". Baseball's mission statement concerns bats, bases, butts and seats -- not visualizing correlations among heterogeneous data stores. By releasing their data for free they let the smartest people in the world have the opportunity to perform that second task for free.

We're about to enter the age of ubiquitous information. Drawing these data stores into open formats, making them discoverable, and interconnecting them across knowledge domains presents explosive opportunities. But who will own this data and what access will they allow? If you want to help ensure that the answer is 'everyone' and 'all of it', come join the http://infochimps.org project, a free open community effort to build an Allmanac of everything.

Hey! That's me &co., yay. Our goal is to build the best free repository of data on the web, like a 'flickr for datasets' or the almanac to wikipedia's encyclopedia. There's some similar, excellent work being done by numbrary.com, swivel.com and Freebase, and we intend to complement those projects (as well as to fuel and inspire the projects you all are building). The difference between infochimps and those sites is we're 'messy' -- we'll take the data as it stands, and we'll give it to you with no restrictions and no sandbox, to play with on your machine with your tools. As people find various datasets interesting or useful and thus choose to enhance them (better formats, more metadata), we'll make it easy to share those contributions back.

If you're excited about this project and want to help it grow, please (http://help.infochimps.org/help/show/Contact) contact us. The site's still in rough shape, but there's a lot there to love. You can also follow along with the development at our blog -- http://blog.infochimps.org/