HN user

renegat0x0

1,163 karma
Posts19
Comments559
View on HN
github.com 6mo ago

Show HN: Open database of link metadata for large-scale analysis

renegat0x0
15pts1
github.com 1y ago

Show HN: BruteFeedParser – a way to brute parse RSS

renegat0x0
1pts1
github.com 1y ago

Show HN: Simple Crawling Server

renegat0x0
2pts1
github.com 1y ago

2024 Link Meta Data

renegat0x0
1pts1
rumca-js.github.io 1y ago

Show HN: Offline Search of Domains

renegat0x0
2pts0
github.com 2y ago

Why do we have GitHub curated lists?

renegat0x0
3pts4
news.ycombinator.com 2y ago

Is YouTube starting to protect channel RSS feeds?

renegat0x0
36pts12
news.ycombinator.com 2y ago

Ask HN: Is HTML meta field a hell incarnated?

renegat0x0
1pts3
www.404media.co 2y ago

Google Search Has Gotten Worse, Researchers Find

renegat0x0
7pts0
github.com 2y ago

Show HN: Circumventing YouTube adblock policy – Firefox extension

renegat0x0
1pts0
github.com 2y ago

Show HN: Link metadata. Complete year 2023

renegat0x0
1pts0
www.youtube.com 2y ago

Simple mobile tools suite is sold to Zippo apps [video]

renegat0x0
2pts1
www.forbes.com 2y ago

Web Browsing Data Is 'Serious Security Threat' to EU and US

renegat0x0
4pts4
news.ycombinator.com 2y ago

Tell HN: Personal page title is most probably incorrect

renegat0x0
1pts0
news.ycombinator.com 2y ago

Tell HN: You are looking at the Internet through a keyhole

renegat0x0
3pts1
github.com 3y ago

Link Archive – 03.2023 Update

renegat0x0
2pts2
news.ycombinator.com 3y ago

You should be saving RSS files in archive.org

renegat0x0
19pts5
github.com 3y ago

RSS link archive – update for year 2022

renegat0x0
2pts1
github.com 3y ago

RSS Link Archive

renegat0x0
1pts1

I tried recently to create dev. account. I have not yet been successful. It is a painstaking process.

I had to submit my ID, my phone number, email.

Then to verify I had to give my address. They rejected my ID twice, so I had to submit driving licence.

I am several weeks in, and could not even produce a single app.

Their algorithm already rejected me, for no obvious reason.

I maintain databases of links

- h ttps://github.com/rumca-js/Internet-Places-Database - Internet places / YouTube channels

- h ttps://github.com/rumca-js/awesome-database-feeds - feeds / RSS locations

- h ttps://github.com/rumca-js/awesome-database-top - smaller database from above

- h ttps://github.com/rumca-js/awesome-database-awesomelists - links from 'awesome lists'

- h ttps://github.com/rumca-js/RSS-Link-Database-2026 - 2026 year link metadata

- h ttps://github.com/rumca-js/RSS-Link-Database-2025 - 2025 year link meta data

- h ttps://github.com/rumca-js/crawler-buddy - crawler engine

Agree to disagree.

I think many platform did support RSS, or API, but at some time it was dropped. It is not that hard to provide RSS. It is not like it has to be implemented from scratch. It is just dropped by platforms. There must be reasons for dropping. One may argue it is not worth it, but even to drop functionality there needs to be a decision from management. The management always think about money. Is RSS or API monetized correctly? Free? No? Then we drop it, because data from algorithms serving user contents can be easily monetized. Just follow the money, the oldest truth.

How else to notify readers? RSS is quite good.

Corporations do not like it though, because they want to control spread of information. They prefer you re-enter page, be targeted with more ads, or spoonfed with their news algorithm on-site.

However RSS only works, if someone knows about your page. YouTube is quite good with RSS though. You can have channel RSS, be notified about new videos, while channels still are discoverable.

Also shameless plug to my own database of RSS feeds

h ttps://github.com/rumca-js/awesome-database-feeds

I already complained about post on reddit. It says that link to RSS is hidden, which is not true IMHO.

YouTube page contains HTML link to RSS feed in channel page, and most RSS clients should just pick it up just fine.

By the way I maintain a list of feeds, many of them are youtube in link below, so if you would like to find a channel you can use it

Links:

h ttps://github.com/rumca-js/awesome-database-feeds

Article asks what next. I know what's next.

It is similar with crypto wars. They try and try until they have backdoor everywhere.

About verification they will try to implement WEI on browsers, and verification on os.

It is a crusade to make you always identifiable. Companies and governments want it so much because it is so valuable to them, it adds so much power over people.

So what's next. They will move borders here, and there. Every year.

I follow awesome lists. These are curated lists of software. It reverts google indexing, because search is awful.

About personal blogs... I have many many personal blogs in my repository. Around 4k. Respository below. The real problem is to find quality stuff. You can have millions of them, but if they are not worth my time, then what is the point?

I cannot verify and decide what is good manually. Obviously.

I think we cannot also rely on Google to provide rating, nor any corporation.

So I have my own ratings, because at least I will be able to find what I found worth before.

Link to my repo:

https://github.com/rumca-js/Internet-Places-Database

I am talking about many things. Also about my programs, but also data. Programs are not that important for me. They are just vehicles to get the job done.

All my programs and data are open. It is something that anybody can pick up, and use as they wish

- https://github.com/rumca-js/Internet-Places-Database - domains I found

- https://github.com/rumca-js/Internet-feeds - feeds I found

- https://github.com/rumca-js/RSS-Link-Database-2026 - news from 2026

- https://github.com/rumca-js/RSS-Link-Database-2025 - news from 2025, etc.

Does make any change? I don't know. I run web crawlers. It is interesting for me to see what my crawlers pick up from the Internet. It did change my life, these project changed how I see the Internet. More pro-activly.

I think there are many projects which can be useful for niche groups. I suppose I have 390 stars on one repo. I hope at least my projects were useful for them. That is a hopeful thought.

I am nobody. I have little impact. I want my programs to be safe from government intrusions, from age checks, from encryption backdoors, from corporate surveillance. How do I win this battle with big tech?

I am deeply in self-host. For the self-host to succeed it needs to be better, unregulated, and free. It needs to be easily distributed. The data should be easily distributed. Import and export should be fast and easy.

That is why most of my programs use JSONs that are human readable, or use SQLite tables that are just copy-paste away.

I am from Poland. My ancestors were able to survive by hiding, and by fighting small partisan battles. My idea of software is "partisan". It battles big tech in small, distributed ways.

I am not sure, but I think what I said is similar to interoperability.

Why I forked httpx 4 months ago

There are many nice http clients:

- httpx

- curl cffi

- httpmorph

- httpcloak

- stealth crawler

I wrote a framework, link below, which uses them all. You can compare each to verify crawling speed. Some sites can be cleanly crawled with a one particular framework.

Having read the article I am in a pain. I do break things while development. I rewrite stuff. Maybe some day I will find a way to develop things "stable". One thing I try to keep in good shape is 'docker' image. I update it once everything seems to be quite stable.

https://github.com/rumca-js/crawler-buddy

Reminds me of "Website obesity crisis"

- https://www.youtube.com/watch?v=iYpl0QVCr6U

- https://idlewords.com/talks/website_obesity.htm

Some say that you should not use ad blocker, because that kills ad revenue, but I did not forced anybody to rely with their lives on ad revenue. Many of things were 'free' because we were all just using ad blocks, and then it all became commodified, simplified, so simpletons without ad blocks became a thing. Now they shame people for using ad blocks, even though it stops spreading malware and viruses.

I plan to use ad block, and use as many extensions that protect me. If there is some form of goods, be it streaming movies, audio, books I will happily pay for it. I will not accept a web with ads. I prefer touch grass. There is a clear line for me.

Also there is no line ad publisher will not cross. The goal posts are shifted, so you will never satisfy shareholder greed. The only pushback is trough ads and probably sometimes piracy. Not that I advocate it, but in reality if companies push too hard, there are consequences.

- I think I was upset when Google allowed fake ad for VLC to appear high in ranking

- I hate that Google returns content farms instead of product web pages

- I hate that Google provides a page of 10 useful links, later links are just pure garbage. I think that something in Google engine is profoundly broken

- I maintain my own search index, but it requires a lot of effort, and attention. I do insert links if I find them worthy. I think more people should have their personal search indexes. Mine is below. I am quite happy that problems like these do not affect me that much

https://github.com/rumca-js/Internet-Places-Database