Since the books are available on the site as text and HTML the search engines index them already for you. Try searching for the below; it should take you to the book you expect as the first result:
site:gutenberg.org "it was the best of times"
HN user
Since the books are available on the site as text and HTML the search engines index them already for you. Try searching for the below; it should take you to the book you expect as the first result:
site:gutenberg.org "it was the best of times"
One author remains blocked in Germany (but only for a couple more years)...
This is covered in the FAQ - https://www.gutenberg.org/help/faq.html#why-is-project-guten...
And as another person noted, the vast majority of books have HTML, EPUB, Mobi formats. We are also looking at both KEPUB (Kobo) and PDF which will probably come in the future.
As another commenter said PG is almost all books from 95+ years in the past due to copyright law in the US. We partner with a sister organization, the World Library Foundation, who have a self-publishing portal for modern works by authors who wish to put their own work in the public domain. You might want to look there for more modern material. https://self.gutenberg.org
Cloudflare sells that as a product, they call it Labyrinth IIRC.
See what I wrote above (and let me say I am talking about Project Gutenberg and Distributed Proofreaders here, I am one of the admins on both). A large amount of the hassle traffic we've seen is as I wrote above, the IPs come from everywhere and in many cases, each IP makes a single request and doesn't come back. They change user-agent dynamically, etc, to masquerade as regular traffic. They come from residential, cloud/hyperscale, corporate, educational, government, all the networks, on every continent. This is many thousands of "open a ticket with someone" events per hour territory. It's as difficult to fight as DDoS itself for the same reasons (presumably the harvesting parties know that and that's exactly why this approach is used).
Others online have been writing about their own experience with the same stuff; it's not unique to PG at all, it's everywhere. Talk to anyone that runs a web server and they'll have these stories...
One could argue that this falls into the previous poster's thought about "the little differences to modern English are part of the charm" ...
OCR has improved a lot since then, but OCR is just step 1 of reading in text. They make a lot of errors (even now, especially on old worn out paper pages) and even if they didn't, one has to format the book, deal with footnotes, sidenotes, illustrations, etc. DP is very active, we will welcome you back with open arms :)
There are many books available as audio, some are human-read, some were automated. You can see lists here:
human-read: https://www.gutenberg.org/browse/categories/1
computer-generated: https://www.gutenberg.org/browse/categories/2
IIRC many of the human-generated ones come from LibriVox, many of the computer-generated ones came from a collaboration with Microsoft.
Worse than that - even if they would take action, you can't possibly orchestrate filing all of the complaints. It's a drown-in-quicksand problem, you can't fight quicksand one grain at a time.
The ebook editions are very good for this. Most of the e-reader software provides all the amenities (bookmarks, highlighting, notes, control of margins, etc).
The Alfred Döblin books are still blocked in Germany (for a couple more years).
wouldn't help, much of the traffic we've observed look closer to ddos patterns - IPs from all over the world, many different networks, each IP makes one request only, doesn't come back. highly distributed, no form of blocking would be effective except maybe captcha or proof of work.