I replied in detail elsewhere:
HN user
pierrefar
Product management and technical consultant for anything related to websites. More https://deliberatedigital.com/
Previously:
Google Search engineer and PM in Google ads.
Built and sold Cligs ( http://cli.gs/ ), a URL shortener with analytics for professional marketers.
Built OpenCourseWare search engine: http://www.ocwsearch.com/ Donated to OCW Consortium.
Email: "hackernews" at my username dotcom.
My top color is #FFDC7A.
A search engine index is an economic exchange between the website and the publisher.
To massively (over)simplify the argument to its essence (and ignore other important points): the publisher goes through the trouble and expense of creating the content The publisher then allows its content to be copied by a search engine only because being shown in search results gets it traffic back. The traffic it gets in return has value, and the publisher is happy for this arrangement to continue as long as the value of the traffic is more than the cost of producing and serving the content.
Brave offering a "license", for its own financial benefit, to "allow" others to use the content for LLM training gives zero benefit to the original publisher. This is why I use words like "sleazy" to describe Brave's position.
This argument applies to Google and Microsoft. Right now both are failing at citing sources in their generative AI search results. That is terrible and I hope it's fixed soon, as otherwise they're being sleazy scrapers as much as Brave is.
Finally, I wholeheartedly disagree they what Brave is doing is for the "greater good". The fact they charge extra for the "license" to use the content for LLM training shows that.
There is a a difference between a human being able to access content vs a search engine indexing it (and in the case of Brave, "licensing" it on).
I share your concern about Google having this much power, and I'd add that Microsoft Bing is equally bad but gets away with it because they're smaller. Still, the final decision about which search engine indexes a website is purely the publisher's.
This is explained more in the article I referred to, but briefly: Brave delegates crawling to normal Brave browsers, so it's a huge IP addresses pool, not a single IP address or range.
Also, these search crawls by the browser do not identify themselves beyond the Brave standard UA header, namely a plain Chrome user-agent string.
That would be bad, and it is already bad that Google and Microsoft control so much of search queries, but the decision about which search engine indexes a website is purely the publisher's.
The major problem with Brave search is their position about indexing and licensing content against the wishes of the website publisher. Their robot does not identify itself, meaning the publisher cannot use the standard robots.txt to block its crawling if the publisher so wishes. Incidentally, the robots.txt file has been used in court cases litigating if a search engine is legal or not.
Even worse, they state that Brave search won't index a page only if other search engines are not allowed to index it. It is morally not their right to make that call. A publisher should have full control to discriminate which search engine indexes the website's content. That's the very heart of why the Robots Exclusion Protocol exists, and Brave is brazenly ignoring it.
Even worse than that, the Brave search API allows you (for an extra fee) to get the content with a "license" to use the content for AI training? Who allowed them the right to distribute the content that way?
I wrote about all this here:
https://searchengineland.com/crawlers-search-engines-generat...
and more references elsewhere in this thread:
https://news.ycombinator.com/item?id=36989129
Amusingly, while I was writing my article, this got posted to their forums, asking about how to block their crawler:
https://community.brave.com/t/stop-website-being-shown-in-br...
No reply so far.
This article is way out of date and wrong, although published recently (at least according to the timestamp).
Google's official position was published on 8 February here:
https://developers.google.com/search/blog/2023/02/google-sea...
It's a much more nuanced position that can be summarized as "make sure you create good content, however you create it". A focus on quality, not process, is reasonable.
Disclosure: ex-Googler in search.
In simplified terms, did you find everything you could have possibly found? Looking at the formula in the article, it includes the false negatives, that is, items you misclassified as negatives when you should have considered them positives. And because that happened, you didn't find them in the set, that is you "forgot them". The opposite of forgetting is... recall.
Another place this idea comes up is a search engine index. If the algo doesn't find, for a given query, documents in the index it should have (falsely classified as not matching the query), it will have lower recall.
Looks like the stages of mitosis (cell division):
https://www.nature.com/scitable/topicpage/mitosis-and-cell-d...
The faint lines are the cell walls and the bright spots in the middle would be the DNA. I can believe this is what they're going for with a bit of squinting.
Very neat. I built a virtually identical internal tool for Blockmetry. A couple of tips from experience:
1. Add other browse extensions, and you'll see a big difference between their effects. Defaults matter a lot in this space.
2. Compare mobile vs desktop. Getting mobile emulation to be good enough is a bit of work, but worth it IMO.
Based on internal usage, the typical web page will load 35-45% faster with uBlock Origin installed.
My email address is my profile if you want to compare notes or whatnot.
No that's not a solution. It's the tracking that counts, not the cookies. I commented elsewhere on this thread more details:
Before anyone thinks this (and similar) approaches are a way around the GDPR's cookie consent tracking crackdown: It's not.
The GDPR talks about online identifiers, of which cookies, IP address and fingerprints are examples. If you read any regulator's guidance carefully, you'll see they talk about "cookies and similar technologies", with just "cookies" being used alone for brevity.
To rephrase tracking of any kind is the issue, not cookies. Don't mistake the implementation for the activity.
Disclosure: Founder of a non-tracking web analytics service because of this exact issue.
Congrats on the launch.
The privacy policy is very not suited for this service. The most important point is that you're based in Germany based on the address in the policy, but there isn't a single mention of the GDPR. That and the ePrivacy Directive are what count for you the most. My recommendation is don't use a free policy generator and get proper advice. I appreciate this isn't something commonly seen as a launch blocker, but it's important to sort it out properly.
Find your German state data protection authority, and invariably you'll find they have great guidance.
Here is a write-up of the decision from EU’s highest court on this topic: https://www.whitecase.com/publications/alert/court-confirms-...
It’s easy to see why quote I gave says what it says with this context.
Also, if you’re worries, talk to your lawyer.
Yes, and also cookie IDs. Both are called out as examples in recital 30:
“Natural persons may be associated with online identifiers provided by their devices, applications, tools and protocols, such as internet protocol addresses, cookie identifiers or other identifiers such as radio frequency identification tags. This may leave traces which, in particular when combined with unique identifiers and other information received by the servers, may be used to create profiles of the natural persons and identify them.”
Source: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex%3A...
To add to the key point about privacy, this research from Princeton is really illuminating and scary: https://freedom-to-tinker.com/2017/11/15/no-boundaries-exfil...
Looks good! I'm the founder of a similar service (Blockmetry). Obviously non-tracking web analytics is the future!
I'm curious why you chose to host the data yourself instead of giving customers the data immediately at the point of collection. That's the path we chose for Blockmetry as it genuinely required to be a non-tracking web analytics service and makes it impossible to profile users. Any service that hosts its data would still be open to being untrusted on the "no tracking no profiling" argument.
Thanks, Pierre
PS - YC Startup School founders: ping me via the forums and get an extended-period free trial.
Speaking of commas, you're missing one after the end of the interrupting phrase in your last sentence (should say ", as well as the state of Maine,"). It's a pet peeve bigger than the lack of Oxford commas, and definitely affects readability and may affect meaning.
I don't have access to the raw log files from the customers, so can't give you a percentage. All I'll say confidently is that my service processes a lot of bot traffic that needs to be filtered out before reporting.
BTW, are you the same Peter Hartree on this Segment thread? https://community.segment.com/t/1889n1/how-common-is-client-... It would appear we've crossed paths before on this topic. Please do email me if you want to talk properly. That Segment thread has my email.
a lot of bots still don't execute JavaScript
I operate a service that measures this (see another comment on this discussion), and all I'll say is you'll be very surprised how many bots actually execute JS, especially stealth bots. You have to be careful either way.
I run a service called Blockmetry [0] that measures exactly that, directly from pageviews. Some numbers get published regularly [1]. The percentage from August (last public number) is 5.2% of non-bot JS-enabled pageviews did not fire the analytics tag.
The short answer is that it's significant on an aggregate level worldwide, but the reality is that it varies _massively_ by country, device, day of the week [2], and even different sections on the same site. Additionally, there is small percentage of pageviews that have JS disabled you have to account for. This analysis was on HN earlier today [4] saying 0.2% of pageview worldwide have JS disabled, but, again, with huge variation (notably, Tor, but elsewhere too).
Q4 numbers are not released yet, but the trend is generally up, with some notable drops. Get in touch if you want more info or to set it up on your site [5].
[0] https://blockmetry.com/ [1] https://blockmetry.com/weather [2] https://blockmetry.com/blog/weekday [4] https://blockmetry.com/blog/javascript-disabled [5] https://blockmetry.com/contact
Append hl=hi query string parameter to the page to get it in Hindi (hi is the language code for Hindi). Example:
(Former Googler worked on this topic and now help businesses with this kind of question)
The short answer is focus on what is best for your users and what they expect, but there are a couple of things you need to do for best indexing in some cases.
To answer your last question first, if your server can respond to a URL and browsers fetch it successfully, Googlebot and indexing will be fine. Want to use UTF-8 for fully localized URLs? Go for it. Want to stick to ASCII? Sure. What's best for your users is the answer.
Trends across languages is a very odd request. Take the string of characters "chat". In English it's used in many different ways, a verb, a noun, etc. Same string in French means cat. You really wouldn't want to look at the trends of this string by combining English and French.
Now your biggest question: a search engine tries to identify the language of the pages in its index, and also the language the searcher is using. To illustrate:
1. The "chat" example is perfect for this, but also any number of queries.
2. Take someone searching from Switzerland. Wouldn't it be better to show them pages in their preferred language, be it German, French, or Italian?
3. Take someone looking for a specific business. If that business has pages for its UK, Australia, and USA subsidiaries, it would be great to show searchers the right country page if they're searching from the UK, Australia, or USA, even if all are using English queries and these localized pages are also all in English. For that, you'll need hreflang annotations, which also works across languages in the same country (Switzerland) or across languages (global pages in English and Spanish).
FWIW, the FTC is getting really interested in influencer marketing. Also from Bloomberg: https://www.bloomberg.com/news/articles/2016-08-05/ftc-to-cr...
Or ask your favorite search engine about FTC influencer marketing.
Hi
This is a new service to measure ad blocking rates accurately by any combination of device type (mobile, desktop/tablet) and country.
Mention HN and I'll bump you up in the beta queue.
Ping me any questions.
Thanks!
I work at Google in web search, and I have a few comments about this discussion.
Firstly, I think this whole discussion about page speed is the wrong way to approach it. The primary motivation for page speed should be user happiness, which affects the key metrics you care about like user acquisition, conversion, and revenue. The fact it's a (small) ranking signal is a nice benefit, a cherry on top. Here is a nice case study about page speed and user metrics from Lonely Planet:
http://cdn.oreillystatic.com/en/assets/1/event/88/Performanc...
The conversion rate graph on slide 9 is what pretty much every study that looks at performance and user engagement finds. And here is one from Google search about the effect of page speed on searchers:
http://googleresearch.blogspot.co.uk/2009/06/speed-matters.h...
Secondly, there could be other issues with how the experiment was conducted on a technical level:
1. The experiment was about using JavaScript. Was Googlebot allowed to crawl the JS files? If robots.txt blocked crawling, that would have translated to less content visible to Googlebot, and so less content to index, which can easily result in a loss of ranking.
Note that the JS file itself may have been crawlable but it may have made an API call that was blocked. Same end result in terms of indexing.
2. Related to (1), we only started rendering documents as part of our indexing process a few months ago. When was this experiment conducted? If it was before full rendering was the norm, it's very likely we didn't index JS-inserted content that we could now, which, again, may have resulted in lower ranking.
For both of these, using the Fetch and Render feature in Webmaster Tools gives you the definitive view of how our indexing system sees your content. Before running any such experiement, it's worth running a few tests using Fetch and Render.
This suggests something else is going on. Please post in the forums with the site details.
Google treat the http and https versions of a domain as SEPARATE PROPERTIES.
That's not quite accurate. It's on a per-URL basis, not properties. Webmaster Tools asks you to verify the different _sites_ (HTTP/HTTPS, www/non-www) separately because they can be very different. And yes I've personally seen a few cases - one somewhat strange example bluntly chides their users when they visit the HTTP site and tells them to visit the site again as HTTPS.
This means that even if you 301 every http page to https when you transition, all of your current rankings and pagerank will be irrelevant.
That's not true. If you correctly redirect and do other details correctly (no mixed content, no inconsistent rel=canonical links, and everything else mentioned in the I/O video I referenced), then our algos will consolidate the indexing properties onto the HTTPS URLs. This is just another example of correctly setting up canonicalization.
By the way, if you're moving to HTTPS, following our site moves guidelines:
https://support.google.com/webmasters/topic/6033102?hl=en&re...
specifically, the site moves with URL changes:
https://support.google.com/webmasters/answer/6033049?hl=en&r...
But you did say you have a client with an issue. I suspect they either implemented the move to HTTPS incorrectly or something else is going on. Please ask for more help at our forums:
https://productforums.google.com/forum/#!categories/webmaste...
I was involved in this launch and I want to address a very common misconception I'm seeing here and elsewhere.
Some webmasters say they have "just a content site", like a blog, and that doesn't need to be secured. That misses out two immediate benefits you get as a site owner:
1. Data integrity: only by serving securely can you guarantee that someone is not altering how your content is received by your users. How many times have you accessed a site on an open network or from a hotel and got unexpected ads? This is a very visible manifestation of the issue, but it can be much more subtle.
2. Authentication: How can users trust that the site is really the one it says it is? Imagine you're a content site that gives financial or medical advice. If I operated such a site, I'd really want to tell my readers that the advice they're reading is genuinely mine and not someone else pretending to be me.
On top of these, your users get obvious (and not-so-obvious) benefits. Myself and fellow Googler and HNer Ilya Grigorik did a talk at Google I/O a few weeks ago that talks about these and a lot more in great detail:
Hi return0
This means that at some point Googlebot discovered https://siteB. It could simply have been a misconfigured CMS, or a bad sitemap, or an errant link one of your visitors shared on a forum, or something a previous owner of the domain did, or anything really. You may think there are no links to the site, and that may be true right now, but it's about something that Googlebot found in the past.
The correct fix is, as you say, to make sure the server doesn't respond to invalid certificate+site combinations.