HN user

denschub

315 karma
Posts7
Comments26
View on HN

Here's something for the next time you want to "expose" a phony: before linking me to your investigative source, ask for exact date-stamps when I made changes to the robots.txt and what I did, as well as when I blocked IPs. I could have told you those exactly, because all those changes are tracked in a git repo. If you asked me first, I could have answered you with the precise dates, and you would have realized that your whole theory makes absolutely no sense. Of course, that entire approach is mood now, because I'm not an idiot and I know when commoncrawl crawls, so I could easily adjust my response to their crawling dates, and you would of course claim I did.

So I'll just wear my "certified-phony-by-orangesite-user" badge with pride.

Take care, anonymous internet user.

just for you, I redeployed the old robots.txt (with an additional log-honeypot). I even manually submitted it to the web archive just now so you have something to look at: https://web.archive.org/web/20241231041718/https://wiki.dias...

they ingested it twice since I deployed it. they still crawl those URLs - and I'm sure they'll continue to do so - as others in that thread have confirmed exactly the same. I'll be traveling for the next couple of days, but I'll check the logs again when I'm back.

of course, I'll still see accessed from them, as most others in this thread do, too, even if they block them via robots.txt. but of course, that won't stop you from continuing to claim that "I lied". which, fine. you do you. luckily for me, there are enough responses from other people running medium-sized web stuffs with exactly the same observations, so I don't really care.

the robots.txt on the wiki is no longer what it was when the bot accessed it. primarily because I clean up my stuff afterwards, and the history is now completely inaccessible to non-authenticated users, so there's no need to maintain my custom robots.txt.

Is it all crawlers that switch to a non-bot UA

I've observed only one of them do this with high confidence.

how are they determining it's the same bot?

it's fairly easy to determine that it's the same bot, because as soon as I blocked the "official" one, a bunch of AWS IPs started crawling the same URL patterns - in this case, mediawiki's diff view (`/wiki/index.php?title=[page]&diff=[new-id]&oldid=[old-id]`), that absolutely no bot ever crawled before.

What non-bot UA do they claim?

Latest Chrome on Windows.

"I'm in Europe so I don't have to care because software patents are not enforceable here" isn't the solution. Yes, patent law doesn't apply - but copyright law does, and they very much can take down content that references the spec just based on copyright law alone.

The page you linked to has a FAQ section.

Q: Is membership in the Thread Group alone, at any membership level, sufficient to gain and receive royalty-free intellectual property rights (IPR) for Thread technology? A: No, membership at any level is not sufficient to gain and receive royalty-free intellectual property rights (IPR) for Thread technology.

and an Associate membership does not apply because I am not white-labeling or rebranding existing products.

It's baffling that some people here can read a sentence like

Membership in Thread Group is necessary to implement, practice, and ship Thread technology and Thread Group specifications.

And somehow think that the restrictions only apply when you "ship", but not when you "implement" or "practice".

I'm sure Thread Group would love to stop Google making OpenThread available

Google is a founding member of the Thread Group. OpenThread exists publicly because it's the only widely available implementation that's shipped in a lot of places. Nordic's SDK, for example, uses OpenThread.

OpenThread is built by and for members of the Thread Group, and used by them. It's fairly clear that Google doesn't care much about anyone else.

Lawyers have little issue defining blogs like mine as "with commercial interest". I have a side-business, so lawyers could make the argument that I use my blog as advertising. I have a Ko-Fi link in the bottom of one specific site, that's a commercial interest, too.

Unless your blog is "I'm sharing holiday photos and nothing else", there's a lot of instances where it could be define as an outlet with commercial interests.

And, ultimately, I have no desire to spend any time and money on fighting even completely invalid claims. I'd rather spend my time watching cat videos on YouTube instead.

You are 100% correct in your assumptions. This isn't my first incident, and it won't be my last. As much as some folks here want to call me an asshole for that, we have a very good understanding of how this ends if we don't lock comments.

That was a limited time experiment, which got shut down after it didn't prove very successful. But from that, it's probably a fair guess that it would not perform well if rolled out again today. There are a few very vocal people who want to support MDN, and that's great, but that's probably not true for the bigger audience. :(

That's amazing that you can do that.

Anonymous credit cards are ruled out by law basically everywhere in the European Union. Assuming that I live in the US, and that everyone on this planets is doing so, is - as you call it - incredibly stupid.

Tor, originally from "The Onion Router", works by routing your traffic through multiple Tor nodes. Like an onion, each node only peels off one layer and passes the packet on to whoever is addressed on that layer. Each node only knows the details about the next node. Eventually, the packet will hit an "Exit-Node", at which point it will be routed via the internet through the endpoint, but it's not a single route.

And while that does not change for every request (that would be highly unpractical), all Tor clients offer you a very quick "get a new route" with just one click.

We also have multiple documented cases of "no-log VPNs" submitting their logs to law enforcement. I even linked to one case in my post. What's your point here, exactly? Because my point was you have to trust either party.

Oh, and btw, here in Europe, it is actually illegal for ISPs to give connection data away for non-law-enforcement purposes. It's sad that there are some US-American ISPs that have a record of selling some information, but the world does not evolve around the USA.