Also, the article does state what Hadley's take on the question is: "As Wickham defines data science as “the process by which data becomes understanding, knowledge, and insight”, he advocates using data science tools where value is gained from iteration, surprise, reproducibility, and scalability. In particular, he argues that being a data scientist and being programmer are not mutually exclusive and that using a programming language helps data scientists towards understanding the real signal within their data. "
HN user
LisaG
Science geek with a strong affinity for computer nerds. Director of Common Crawl www.commoncrawl.org Former Chief of Staff at Creative Commons www.creativecommons.org
Did you watch all of Hadley's video? You might get the title more if you saw/see the whole talk :)
As long as you obey robots.txt there is nothing wrong with crawling. Your code in GitHub doesn't give any indication of what sites you collect data from so there is no indication that you are scraping instead of using it to crawl in an acceptable manner. Though it wouldn't hurt to label your work as crawler scripts instead of scraping scripts ;)
Why use your own scripts and not Nutch?
Do you know about Common Crawl? https://aws.amazon.com/public-datasets/common-crawl/ It obeys robots.txt so it may not have everything you want, but it could save you part of the effort of crawling yourself.
San Francisco CA Full-time / Onsite
New, somewhat stealth startup, for profit company focused on social good.
We have a very talented team so far comprised of : full stack web dev, data architect, 2 junior software engineers, CTO, CEO (me), 2 marketing people, and a business operations person.
We are looking to add a designer and devops.
We work out of the top floor of my house for now - it is a comfortable space. We have funding. The team members we have so far are wonderful to work with, everyone gets along well, and we all feel like what we are building is work that matters.
Please email me if you want to hear more about the team, the stack, and the product.
So excited so see Common Crawl data be useful for such fascinating work!
I work at Common Crawl :)
Love this idea!!
Great question and great blog post! I am looking forward to reading the Homay King stuff that uses Queer Theory and will probably reread Computing Machinery and Intelligence more thoroughly.
I played around with Prismatic before but it just didn’t grab me and I found I didn’t use it that much.
This new version is a whole different animal. Not only is it much prettier (great design) but they seem to have seriously improved their relevance algorithms. I would be very interested to hear from their data team why the relevance is so much better now - anyone from Prismatic monitoring these comments?
There will be news about a subset sometime next month!
If you don't feel like reading the paper Sebastian wrote on the Common Crawl data, he gives a summary of his findings in this video.
Link to full paper: http://bit.ly/14dxSJq
We do think it is worth it to avoid duplicative efforts.
Suppose you crawl 3 million pages and you pay for the compute and storage costs. Then the next person who wants crawl data goes through the same effort and pays the same costs. Doesn't it make much more sense to have a common pool of open data that everyone can use? Even if the effort and costs are low, they are not zero.
For the smaller frequent crawl, we are working with Mozilla and we are will do the top pages (top according to Alexa).
Internet Archive (currently) doesn't want to put their data on any cloud service. We believe it is crucial that people can easily access and analyze the data so we put it on various cloud platforms. We are talking with a few organizations about getting data donations that we could put in our corpus and make available to everyone, but nothing is settled enough that I can publicly comment on those potential partnerships yet.
Limited resources are the only reason. We are working on a subset crawl of ~3 million pages that will be published weekly starting two weeks from now. But doing the full crawl takes a lot of time, effort and money.
That's interesting. We also could revive phage therapy which uses bacteriophages (viruses that replicate in bacteria).
If you are bored in the San Francisco Bay Area the problem is likely internal rather than where you live, so moving (even to somewhere awesome like Austin) will not resolve it.
I hope that some of you who use/play around with the Common Crawl data will try out using the JSON files from the URL Search and then share your code.
If you didn't see the details in the blog post, Common Crawl is giving out $100 in AWS credit to the first five people who share code that incorporates a JSON file from the URL Search.
"Done" is better than "perfect" should be on a sign hanging in every startup.
Thanks for the catch Djoerd! We will fix it now
From @djoerd Why does @CommonCrawl URL search (http://urlsearch.commoncrawl.org/ ) need 'tld.domain' format rather than 'domain.tld'? Read Google's BigTable paper.
Love this post! Only thing I disagree with is that you only need on to say yes. The first yes might not be the best match for you. Compatibility is not quite as important in business as it is in romantic and sexual relationships.
Strongly agree! blekko's Bill of Rights is a great expression of their values and of why we should all be using blekko.
blekko Bill of Rights
1. Search shall be open
2. Search results shall involve people
3. Ranking data shall not be kept secret
4. Web data shall be readily available
5. There is no one-size-fits-all for search
6. Advanced search shall be accessible
7. Search engine tools shall be open to all
8. Search & community go hand-in-hand
9. Spam does not belong in search results
10. Privacy of searchers shall not be violated
Graue I am from Common Crawl. We don't filter for porn. A corpus of web data needs to include porn or it wouldn't be a representative sample of the web ;) We do want to enrich our sample of the web with high-value sites and that is where the blekko data will be so incredibly valuable.
I really appreciate your mention of LGBT and sexual health sites being collateral damage - we need to draw more attention to that problem. I would love to see someone work with Common Crawl to improve methods of distinguishing. Lisa
I am part of Common Crawl and I just wanted to say that we are super excited about blekko's donation! This is yet another demonstration how much blekko values openness and transparency.
Oh and the girl who says that engineers are secondary to the success of a startup and that people with ideas are the key element. There are a lot of those people around here.
Another archetype is the guy who tries so hard to be a "brogrammer". Sadly those guys exist too.
So cool you are putting together a presentation of results!!
San Francisco: Data Scientist, Crawl Engineer
Do work that matters on big data! Common Crawl is an open repository of web crawl data with a corpus of over 100 TB.
We’re looking for someone enthusiastic about open source, net neutrality, open data and keeping the web truly open. Common Crawl is dedicated to building and maintaining an open repository of web crawl data in order to enable a new wave of innovation, education and researchWe’re set to do amazing things this year, and there is no better place to hone your big data skills than helping us manage and process our 100 TB corpus. Plus, you’ll be working within a passionate community and have the chance to interface with plenty of talented researchers, educators, startup folks, and an incredible advisory board.
If you’re looking to do work that matters, come join us!
http://commoncrawl.org/team/jobs/ Email lisa (at) commoncrawl.org
He did this in from idea to finished experiment in 4 days and used about 300 lines of Ruby. Strong demonstration of how low the barrier can be to working with big data.
"The key lesson I’ve learned from the exercise is that given the tools and data available today, either for free, or at very low cost, it’s possible for anyone to work with relatively Big Data without too much weeping and gnashing of teeth."
Site is back up. Thanks for your patience!
Hi
I am from Common Crawl. Apologies for the site being down! Too much traffic from HN :) We're working on getting it back up. The Google cache below has all the contents, so please refer to there for the moment. Here's the excerpted beginning..
Learn Hadoop and get a paper published
We’re looking for students who want to try out the Hadoop platform and get a technical report published. Hadoop’s version of MapReduce will undoubtedbly come in handy in your future research, and Hadoop is a fun platform to get to know. Common Crawl, a nonprofit organization with a mission to build and maintain an open crawl of the web that is accessible to everyone, has a huge repository of open data – about 5 billion web pages – and documentation to help you learn these too
http://webcache.googleusercontent.com/search?q=cache:http://...