Thank you for the comment. The fact that it can run independently as soon as it's forked was exactly the point I was making.
Also, the cat | grep pipeline is illustrative. I remove it at the end.
HN user
Thank you for the comment. The fact that it can run independently as soon as it's forked was exactly the point I was making.
Also, the cat | grep pipeline is illustrative. I remove it at the end.
Thank you for the compliment. Point taken on cat, but that's the way I like to introduce the process to people. I took cat and grep out at the end of the article anyway.
Hi all, original author here.
Some have questioned why I would spend the time advocating against the use of Hadoop for such small data processing tasks as that's clearly not when it should be used anyway. Sadly, Big Data (tm) frameworks are often recommended, required, or used more often than they should be. I know to many of us it seems crazy, but it's true. The worst I've seen was Hadoop used for a processing task of less than 1MB. Seriously.
Also, much agreement with those saying there should be more education effort when it comes to teaching command line tools. O'Reilly even has a book out on the topic: http://shop.oreilly.com/product/0636920032823.do
Thank you for all the comments and support.