It's kind of interesting that you contradict much of what the article concludes, even though the article gives a lot of examples. Maybe your prediction will be true.
HN user
ccgreg
CTO at the Common Crawl Foundation
We do. First off we have a public parquet-format index of all of the urls we crawl every month. And then that also lives in a HDFS table that determines when we want to recrawl a page we've crawled before.
Appreciate your kind words! Many people have worked at Common Crawl over the years, and it's been a labor of love fueled by positive comments like yours and the large list of PhD theses helped by our public web dataset.
If you're referring to Common Crawl, which has existed since 2008, indeed your predictions are somewhat accurate. It's easy to opt out or limit what is collected. The crawling itself is inexpensive to us and the hosting is from the AWS Open Dataset Sponsorship Program. And there's no charge for downloading it.
We aren't sure if that really made a significant difference in Common Crawl's data quality. It does hurt our dataset from a humanities point of view, alas.
Common Crawl's archive has metadata that says when each record (html file) was crawled.
Common Crawl's dataset was downloaded in full 100 times in 2025.
We agree that it would be great if it was even more widely used.
A lot of websites want "bot defense" due to high volume scrapers, and that "bot defense" often also ends up blocking low-volume wget/curl and polite crawlers like Common Crawl's CCBot.
Good timing, I'm about to release that dataset.
Common Crawl is working hard to improve diversity in our crawl.
I don't know of anyone who uses Common Crawl as pre-training data without filtering it. We have an annotation system that lets people pick and choose which subsets they'd like to use.
Common Crawl is a sample of the web, so it's not that directly helpful for someone wanting to make a product price dataset.
I'm a life-long hacker, and my crawler crawls with consent.
The largest index we had was 4 billion, which is tiny. Our crawl frontier was much larger.
and the data that I’ve experimented with from 2014 seemed high quality
That's because it's from the blekko search engine.
That's already been happening for more than a year now.
Common Crawl is switching to reporting dataset sizes in nibbles. As an organisation dedicated to data preservation, we feel it would be remiss to allow this underrepresented unit to fall out of use. Our latest crawl now exceeds 689 tebibbles. Common Crawl Foundation
The complete list hides in the web graph:
https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main...
and the specific file that's every host we've seen in the latest 3 crawls is:
https://data.commoncrawl.org/projects/hyperlinkgraph/cc-main...
Common Crawl, with over one billion, nine hundred and seventy thousand web pages in their archive: 345TB.
Common Crawl is 300 billion webpages and 10 petabytes. I suppose your number is 1 of our 122 crawls.
Common Crawl has been running a low-resource language project for 1.5 years now -- it's a hard problem.
The guts on the inside changed several times during that timespan.
Well, yes, it is a bit distressing that ill behaved crawlers are causing a lot of damage -- and collateral damage, too, when well-behaved bots get blocked.
Please read our email reply. I have no idea if we received your request —- your HN username doesn’t match any request we have received.
Oh, and thanks for letting me know that I need to add our reply to Wikipedia.
Did you see our reply? Edit: by which I mean, we sent you an email that explains what we did and how to verify it. Did you not receive an email reply? If not, please contact us again.
Also, if your site has CC-BY-NC-SA markings, we have preserved them.
Did you see our reply? https://commoncrawl.org/blog/setting-the-record-straight-com...
Also, if your site has CC-BY-NC-SA markings, we have preserved them.
That 20% number is for a limited list of relatively large news websites. If you include the long tail of news, the % of blocking is much smaller.
Many AI projects in academia or research get all of their web data from Common Crawl -- in addition to many not-AI usages of our dataset.
The folks who crawl more appear to mostly be folks who are doing grounding or RAG, and also AI companies who think that they can build a better foundational model by going big. We recommend that all of these folks respect robots.txt and rate limits.
Thanks for the mention of Common Crawl. We do respect robots.txt and we publish an opt-out list, due to the large number of publishers asking to opt out recently.
There's a bit of discussion of Common Crawl in Jeff Jarvis's testimony before Congress: https://www.youtube.com/watch?v=tX26ijBQs2k
Prof. Jeff Jarvis speaking about copyright for news in front of Congress: