Hello. Why it worked this way is because I can't find a list of 588 websites by thoae 16 companies (I've contacted the people in those articles to see if they have a list), I decided to find a list of top 1000 popular sites (which is much easier to find) and then try to clean those up. Obviously that's not a perfect method and leaves a lot of sites flagged by mistakes, but I didn't find any other way. I left it open source since I want people's help if they find mistakes in the list or want to supply their own list. There's nothing special about the code, I pretty much asked LLM to write it the whole way. I didn't set out to make money from this project, I just want more people to see what searches have become: a bunch of people with a lot of money pretty much own 90% of all searches.
HN user
petertsfn
I hear you. I wrote my reasoning on a comment above, but the gist is: It's an afternoon project, I fucked up the list because I didn't think much about it, and this project is probably better off being crowdsourced.
I think it's a proof of concept on how we can theoretically make searches better. Anyways, I'm pushing an update removing all the sites you mentioned. Do you have any other suggestions?
You guys made a valid point. I got a list of 1000 highest traffic sites in the us in 2023, and I did a skim to remove sites that might be actually useful, but apparently I missed a lot.
This was an afternoon project after I read the article about how 16 media companies own ~600 top media sites that's getting organic traffic.
I submitted an update with a more cleaned list (no more .gov and a few other misfires).
Truth is this project is probably better if it uses a crowdsource approach to the source list (like SponsorBlock). It's a bit out of scope for me since I'm not a developer. I just make things I think is fun. My hope with this thing is to raise more awareness to this problem of large brands publishing SEO garbage, and then maybe someone smarter will come along actually make searches usable again.