+1. Also according to some swedish friends, aftonbladet is not a very high quality news service.
HN user
tyropita
Ask me about data science, data engineering, general sys admin stuff, music and the meaning of life.
Makes you wonder if they use LLMs and a chat interface to generate text-based role-playing games, renamed as “simulations”.
Documentation looks really neat and in-depth, always appreciated. Looks like you’re missing a .gitignore file. Folders like __pycache__ don’t need to be checked in.
It was indeed very unexpected. The first time you would try out a cli tool you’d expect just calling its name to return help info and maybe an error.
After a certain scale I think you can let clients do double-work and let the most common crawl data, among different clients, win.
And since you control what URLs need to be crawled, you protect yourself against rogue clients sending arbitrary URLs.
There certainly are a lot of elegant ways to reduce spam for this particular problem imo.
Quite a neat way to crawl websites using a browser extension. That by itself is a form of donation to the search engine. Maybe in the future you can have dedicated software for self-hosted clients that users can run to crawl and index websites for mwmbl? Kinda like folding@home.
How are the batches of URLs to be crawled generated/discovered and posted at your API?
How do you deal with duplicate crawls?
Have you found more open-source projects that follow a similar approach to spreading information about the Ukraine-Russia conflict?
Would be nice to compile a list of them!
What alternative solutions does the HN crowd recommend to Google Photos etc?
Just imagine regular Ukrainian folks commuting to work in their personal tanks. What a time to be alive!
What was the Great Scare?