You could well be right, though I think whatever the mix of Show HN spam vs. ham is the point that almost all it gets no attention remains the same.
HN user
jhpacker
Analytics architect at quantable.com
No worries & thanks! Yea I didn't expect to get them with the excerpts too, that was a surprise bonus.
Unlinked domains can definitely be found in a lot of ways, but like I show in the article there was literally no fetching of the page except for Googlebot. So even if the hostname was leaked somehow the contents of the page require fetching the page, which was only done by Google. Also like I show in the article the content that ChatGPT knows identically matches what's in a Google search snippet, down to where a word-break is.
What I am saying is that this was a glitch where the full prompt rather than a translated prompt was sent to Google Search. OpenAI says they fixed the glitch, so yes it was definitely an error on their part. My research doesn't show how to repro that error, just that it existed.
I definitely wouldn't! Beyond the "glitch" I'm reporting here where full prompts are seemingly sent to GSC, it may still be that searches are scraped... meaning that while it's less obviously personal than a raw prompt it still could leak user intent.
GSC does filter and threshold what shows, but that doesn't always work 100%. Also those filters are built to work against traditional keyword searches, not prompts. It's also supposed to threshold low volume queries which should have kept a lot of things prompts out of GSC, but for whatever reason that wasn't very effective.
I've worked in many GSC consoles over the years, and I've never seen anything like what I saw in this case. (I'm the original author)
Ok this is pretty funny. Just like any good internet commenter this bot didn't actually read the article... which doesn't actually say "AI is a tool, not a panacea" anywhere.
I agree with your take. I've definitely owned and played some excellent sub-$1000 guitars, but at the lower price points things it can be frustrating to deal with things like low-quality tuners, improperly shielded components, etc. I'd say 90% is about pickups, strings, and frets. Most of the 60s guitars I've played were not great tbh.
Yes, it's very much like HotJar, focused on session capture & heatmap.
Unlike Plausible and Fathom, it looks like Rybbit is NOT salting by default ( (but that it's an option to enable per site: https://www.rybbit.io/docs/enhanced-privacy). Which is why they can offer retention reporting.
This seems incompatible with ePD.
I don't trust tools that don't disclose precisely how they track you. They say:
Combining Inputs: We combine key session details (which shall not be named for security reasons) with a cryptographically secure secret value. SHA-512 Hashing: This combined input is hashed using SHA-512, producing a highly secure, anonymized session ID.
They know that we can see what they send in their tracking payload right? They send: hostname, language, referrer, screen resolution, page title, url, and a website id.
So I would presume their highly secretive & secure user session id is: hash(salt + website id + ip + HTTP user-agent + screen resolution? + language?)
I don't see that it says how frequently the salts are rotated, which is one of the key points on which the "no consent banner required" tools like this claim that consent isn't required.
I recently wrote an article on this topic, focused on the power law dynamics that leave such a small amount of room at the top of the industry: https://www.quantable.com/analytics/power-laws-why-our-new-a...
I don't personally think it's new vs. old as much as the power law distribution coupled with the fact that old music is more available and promoted than ever. Plus the algorithms are focused on giving us more of the same thing we have shown it we like rather than new music discovery.
Cloudflare radar, which presumably a much bigger and better sample, reports Bytespider as the #5 AI Crawler behind FB, Amazon, GPTBot, and Google: https://radar.cloudflare.com/explorer?dataSet=ai.bots And that's not including the most of highest volume spiders overall like Googlebot, Bingbot, Yandex, Ahrefs, etc.
Not to say it isn't an issue, but that Forture article they reference is pretty alarmist and thin on detail.
There's nothing in the law that says one-click.
It says, "A prominently located direct link or button which may be located within either a customer account or profile, or within either device or user settings."
I think where the interpretation that one-click sub == one-click unsub is from this passage:
"The ability to cancel or terminate an automatic renewal or continuous service pursuant to subdivision (c) or (d) shall be available to the consumer in the same medium that the consumer used in the transaction that resulted in the activation of the automatic renewal or continuous service, or the same medium in which the consumer is accustomed to interacting with the business, including, but not limited to, in person, by telephone, by mail, or by email."
The idea being that one-click is a medium, which doesn't seem to be the intent here.
With GA4, the tracker code is loaded from www.googletagmanager.com (even if the tag isn't loaded via a GTM container). The measurement requests can be sent to (region1|www).google-analytics.com or analytics.google.com (to share cookies with Google login better).
https://www.quantable.com/blog Analytics experiments and opinion, around 10 years of back deep dive articles -- mostly Google Analytics web performance, bots, SEO.
One of the sites (coop.se) in this decision did use a server-side GTM container to mask the IP before it was sent to Google, but they were still told to stop using GA, but they weren't fined. The DPA said that the _gads, _ga, and _gid cookies were enough to be identifiable. I don't follow the logic there, but that rules out using a proxy for compliance (at least done as coop did it).
Those looking for alternatives can take a look at my book which evaluates 15 different options: https://gaalternatives.guide
I also have a google sheet listing the basics of each of those tools: https://gaalternatives.guide/sheet
I do know Plausible, and their motivation is to make a sustainable business providing basic web analytics, which is why they charge for their service and Google doesn't. The data they provide to the users of their service is like an order of magnitude less detailed than what Google provides.
I get the cynicism about the industry in general since Google led this merger between web analytics and advertising, but there are plenty of providers in the analytics space that aren't following that path.
Cloudflare Web Analytics is extremely simplistic and does not allow for any persistent identification of users or storage of personal information. It uses HTTP Referrers to count visitors and that's it.
One could argue that since it's a US-based company it can't be Shrems II compliant, but you can make that argument about a lot of things.
My opinion is that this applies to GA4 as well.
The decisions don't explicitly mention a version, they say these particular sites: "...shall cease to use the version of the Google Analytics tool used on 14 August 2020". They don't say if that's UA or GA4. The original complaints from NOYB refer to UA, but the issues cited in this decision would apply to GA4 as well.
So when the DPA says "Companies must stop using Google Analytics", there's no reason to think they only mean the version that was already shut off when they published that post.
Most alternatives are not made by advertising companies, but they also frequently aren't free... Rolling your own from the ground up is not necessary or typically advisable when there are so many good options, including many self-hosted and open source options if you're wanting that level of control.
I usually describe the cost of GA as "subsidized by your customers' data".
Interesting, what kind of cookies? Like I say in the article Lou Montulli from Netscape is generally credited with creating the HTTP cookie, which they named cookie based upon magic cookies in unix, though its obviously quite a bit different.
This is a great way to unpack that phrase, thank you! I started the article with that phrase because of the insincere absurdity of it exactly as you describe.
I agree that most people would not mind that scenario, the issue is that that scenario requires the same consent box that the most invasive adtech would use and it's far too onerous on the user as things are now to discern the difference.
Hey, 3x more people clicked that then "I've reviewed the site's terms and they are acceptable to me." :)
Hi, the exact wording on the poll was "If given the option, I would prefer not to be tracked online."
The reason why I didn't define "tracked" for the people taking the poll is that I think that's the most representative way to replicate the question that consent boxes are theoretical asking, but in an abstract way out of the context of a specific website. When a user sees that consent box, they have little to no idea of what tracking is actually happening.
The scaling issues I'm talking about are on the reporting side, not the measurement side. With Clickhouse you can do complex analytics queries on live data, with MySQL as you say you may need to pre-gen with cronjobs, etc.
You've probably already discovered this, but the only way you drill down is with filters. You can ad-hoc define a page filter with multiple URLs, but you can't (to my knowledge) save that filter set, and content grouping UA-style doesn't exist. Their live demo is with their own site's data: https://plausible.io/plausible.io
Good luck!
For self-hosted, if you are doing into the tens of thousands of pageviews per day you may want to turn off real-time and switch to auto-archiving which pregens reports. That level of traffic depends a lot though on the report period you are running against and how big you've provisioned your MySQL server. YMMV and I haven't benchmarked at multiple traffic levels.