HN user

LisaG

641 karma

Science geek with a strong affinity for computer nerds. Director of Common Crawl www.commoncrawl.org Former Chief of Staff at Creative Commons www.creativecommons.org

Posts17
Comments45
View on HN
blog.dominodatalab.com 8y ago

Domino for Good: Collaboration Reproducibility and Openness for Societal Benefit

LisaG
7pts0
commoncrawl.org 12y ago

Lexalytics Text Analysis Work with Common Crawl Data

LisaG
3pts2
blogs.loc.gov 12y ago

Machine Scale Analysis of Digital Collections

LisaG
6pts0
commoncrawl.org 12y ago

Winter 2013 Crawl Data Now Available

LisaG
4pts0
commoncrawl.org 12y ago

102TB of New Crawl Data Available

LisaG
237pts37
commoncrawl.org 12y ago

SwiftKey’s Head Data Scientist on the Value of Common Crawl’s Open Data [video]

LisaG
38pts2
commoncrawl.org 12y ago

A Look Inside Our 210TB 2012 Web Corpus

LisaG
102pts36
commoncrawl.org 13y ago

Share code that uses new URL Search tool and win AWS credit

LisaG
17pts16
commoncrawl.org 13y ago

The Winners of The Norvig Web Data Science Award

LisaG
7pts0
commoncrawl.org 13y ago

Triv.io donates URL index to Common Crawl

LisaG
51pts16
norvigaward.github.com 13y ago

Norvig Web Data Science Award

LisaG
2pts0
commoncrawl.org 13y ago

Spend the 3 day weekend hacking big data - win $1000 cash + other prizes

LisaG
31pts4
semanticweb.com 14y ago

2012 Common Crawl data now available (free and open web index)

LisaG
2pts0
www.slate.com 14y ago

Your Kinect Is Watching You - amazing, disturbing things it can learn about you

LisaG
11pts1
goodsharer.com 14y ago

"The Rick Santorum Diet": hacking your health using campaign finance data

LisaG
3pts0
semanticweb.com 14y ago

Common Crawl dataset to move to AWS Public Data Sets

LisaG
7pts1
blog.stephenwolfram.com 14y ago

Stephen Wolfram on a .data TLD

LisaG
81pts35

Also, the article does state what Hadley's take on the question is: "As Wickham defines data science as “the process by which data becomes understanding, knowledge, and insight”, he advocates using data science tools where value is gained from iteration, surprise, reproducibility, and scalability. In particular, he argues that being a data scientist and being programmer are not mutually exclusive and that using a programming language helps data scientists towards understanding the real signal within their data. "

As long as you obey robots.txt there is nothing wrong with crawling. Your code in GitHub doesn't give any indication of what sites you collect data from so there is no indication that you are scraping instead of using it to crawl in an acceptable manner. Though it wouldn't hurt to label your work as crawler scripts instead of scraping scripts ;)

Why use your own scripts and not Nutch?

Do you know about Common Crawl? https://aws.amazon.com/public-datasets/common-crawl/ It obeys robots.txt so it may not have everything you want, but it could save you part of the effort of crawling yourself.

San Francisco CA Full-time / Onsite

New, somewhat stealth startup, for profit company focused on social good.

We have a very talented team so far comprised of : full stack web dev, data architect, 2 junior software engineers, CTO, CEO (me), 2 marketing people, and a business operations person.

We are looking to add a designer and devops.

We work out of the top floor of my house for now - it is a comfortable space. We have funding. The team members we have so far are wonderful to work with, everyone gets along well, and we all feel like what we are building is work that matters.

Please email me if you want to hear more about the team, the stack, and the product.

Great question and great blog post! I am looking forward to reading the Homay King stuff that uses Queer Theory and will probably reread Computing Machinery and Intelligence more thoroughly.

I played around with Prismatic before but it just didn’t grab me and I found I didn’t use it that much.

This new version is a whole different animal. Not only is it much prettier (great design) but they seem to have seriously improved their relevance algorithms. I would be very interested to hear from their data team why the relevance is so much better now - anyone from Prismatic monitoring these comments?

We do think it is worth it to avoid duplicative efforts.

Suppose you crawl 3 million pages and you pay for the compute and storage costs. Then the next person who wants crawl data goes through the same effort and pays the same costs. Doesn't it make much more sense to have a common pool of open data that everyone can use? Even if the effort and costs are low, they are not zero.

For the smaller frequent crawl, we are working with Mozilla and we are will do the top pages (top according to Alexa).

Internet Archive (currently) doesn't want to put their data on any cloud service. We believe it is crucial that people can easily access and analyze the data so we put it on various cloud platforms. We are talking with a few organizations about getting data donations that we could put in our corpus and make available to everyone, but nothing is settled enough that I can publicly comment on those potential partnerships yet.

Limited resources are the only reason. We are working on a subset crawl of ~3 million pages that will be published weekly starting two weeks from now. But doing the full crawl takes a lot of time, effort and money.

If you are bored in the San Francisco Bay Area the problem is likely internal rather than where you live, so moving (even to somewhere awesome like Austin) will not resolve it.

I hope that some of you who use/play around with the Common Crawl data will try out using the JSON files from the URL Search and then share your code.

If you didn't see the details in the blog post, Common Crawl is giving out $100 in AWS credit to the first five people who share code that incorporates a JSON file from the URL Search.

Love this post! Only thing I disagree with is that you only need on to say yes. The first yes might not be the best match for you. Compatibility is not quite as important in business as it is in romantic and sexual relationships.

Strongly agree! blekko's Bill of Rights is a great expression of their values and of why we should all be using blekko.

blekko Bill of Rights

1. Search shall be open

2. Search results shall involve people

3. Ranking data shall not be kept secret

4. Web data shall be readily available

5. There is no one-size-fits-all for search

6. Advanced search shall be accessible

7. Search engine tools shall be open to all

8. Search & community go hand-in-hand

9. Spam does not belong in search results

10. Privacy of searchers shall not be violated

Graue I am from Common Crawl. We don't filter for porn. A corpus of web data needs to include porn or it wouldn't be a representative sample of the web ;) We do want to enrich our sample of the web with high-value sites and that is where the blekko data will be so incredibly valuable.

I really appreciate your mention of LGBT and sexual health sites being collateral damage - we need to draw more attention to that problem. I would love to see someone work with Common Crawl to improve methods of distinguishing. Lisa

I am part of Common Crawl and I just wanted to say that we are super excited about blekko's donation! This is yet another demonstration how much blekko values openness and transparency.

San Francisco: Data Scientist, Crawl Engineer

Do work that matters on big data! Common Crawl is an open repository of web crawl data with a corpus of over 100 TB.

We’re looking for someone enthusiastic about open source, net neutrality, open data and keeping the web truly open. Common Crawl is dedicated to building and maintaining an open repository of web crawl data in order to enable a new wave of innovation, education and researchWe’re set to do amazing things this year, and there is no better place to hone your big data skills than helping us manage and process our 100 TB corpus. Plus, you’ll be working within a passionate community and have the chance to interface with plenty of talented researchers, educators, startup folks, and an incredible advisory board.

If you’re looking to do work that matters, come join us!

http://commoncrawl.org/team/jobs/ Email lisa (at) commoncrawl.org

He did this in from idea to finished experiment in 4 days and used about 300 lines of Ruby. Strong demonstration of how low the barrier can be to working with big data.

"The key lesson I’ve learned from the exercise is that given the tools and data available today, either for free, or at very low cost, it’s possible for anyone to work with relatively Big Data without too much weeping and gnashing of teeth."

Hi

I am from Common Crawl. Apologies for the site being down! Too much traffic from HN :) We're working on getting it back up. The Google cache below has all the contents, so please refer to there for the moment. Here's the excerpted beginning..

Learn Hadoop and get a paper published

We’re looking for students who want to try out the Hadoop platform and get a technical report published. Hadoop’s version of MapReduce will undoubtedbly come in handy in your future research, and Hadoop is a fun platform to get to know. Common Crawl, a nonprofit organization with a mission to build and maintain an open crawl of the web that is accessible to everyone, has a huge repository of open data – about 5 billion web pages – and documentation to help you learn these too

http://webcache.googleusercontent.com/search?q=cache:http://...