HN user

__ka

420 karma
Posts28
Comments47
View on HN
psyche.co 5y ago

Introverts are excluded unfairly in an extraverts’ world

__ka
4pts0
dilbertblog.typepad.com 5y ago

The Day You Became a Better Writer

__ka
3pts0
www.economist.com 6y ago

Is there a role for options insurance in equity portfolios?

__ka
2pts0
en.wikipedia.org 6y ago

NASA Clean Air Study

__ka
2pts0
www.cloudave.com 6y ago

Software can be Resilient to Recession (2008)

__ka
4pts0
www.bizjournals.com 6y ago

Privacy study shows Google’s eyes are everywhere (2009)

__ka
5pts0
aeon.co 6y ago

Can digital books ever replace printed books?

__ka
5pts1
www.citylab.com 6y ago

Maps of Life Under Lockdown

__ka
1pts0
www.bbc.co.uk 6y ago

Covid-19 – a new regime of surveillance? [audio]

__ka
48pts15
www.economist.com 6y ago

Netflix will remain a blockbuster hit beyond the Covid-19 era

__ka
1pts0
www.ft.com 6y ago

US oil trades at negative prices for first time in history

__ka
42pts3
twitter.com 6y ago

Slack goes down as 42000 sign up for WirVsVirus Hackathon in Germany

__ka
2pts0
en.wikipedia.org 6y ago

Herd Immunity

__ka
3pts0
www.bbc.com 6y ago

WhatsApp to stop working on millions of phones

__ka
4pts0
bookofbadarguments.com 6y ago

An Illustrated Book of Bad Arguments

__ka
2pts1
news.ycombinator.com 6y ago

Ask HN: How would you transmit smell over the web?

__ka
4pts4
aeon.co 6y ago

Surviving a Climate Crisis – Lessons from History

__ka
5pts0
whotracks.me 6y ago

Show HN: Whotracks.me - Tracking the Trackers

__ka
12pts0
www.0x65.dev 6y ago

Privacy or Data, a Convenient False Dichotomy

__ka
77pts26
www.android.com 6y ago

Google to use fourth-price auction for places on search engine choice screen

__ka
16pts3
lectures.quantecon.org 7y ago

Quantitative Economics with Julia

__ka
1pts0
multiplier.gitlab.io 7y ago

Machiavelli on Mergers and Acquisitions

__ka
1pts0
privacore.github.io 7y ago

Goodbye – Findx is shutting down

__ka
5pts1
research.mozilla.org 7y ago

Firefox: The Effect of Ad Blocking on User Engagement with the Web [pdf]

__ka
181pts154
hbr.org 8y ago

Research: Learning a Little About Something Makes Us Overconfident

__ka
2pts0
news.ycombinator.com 8y ago

Ask HN: What are the steps to deal with Android's updated policy?

__ka
1pts0
mozilla.github.io 9y ago

Ask HN: Who is participating in Mozilla's global sprint?

__ka
4pts1
aeon.co 9y ago

Plato knew a lot about behavioural economics

__ka
144pts23
Brave Search beta 5 years ago

The chosen country is important. `uva` may be more commonly associated with the University of Virginia in the US. For Netherlands (same query) https://search.brave.com/search?q=uva&country=nl will correctly point to Universiteit van Amsterdam.

At present we default to country US. We're looking to implement better defaults soon.

We do hope you stick around!

When someone points out that someone did something bad, clarifying they only did it to 1% of one country's users isn't a super strong defense.

I don't see where I made this "clarification".

Let's be clear. Firefox was trying to test switching from Google to Cliqz (where it had a stake). Mozilla had a difficult time trying to break the golden cage they find themselves in. To their credit, they did try. Ultimately the Cliqz-Firefox integration was, unilaterally, cancelled. If your main source of revenue comes from your “competitor” you are slowly pushing yourself to irrelevance.

And also, the privacy issue again: I addressed a similar question in another comment in this thread [0]. If you want to spread FUD, please make a proper case.

[0] Another comment on this thread: https://news.ycombinator.com/item?id=23045099

This statement is misleading. Firstly, for context here's the paragraph you ought to be referring to:

This experiment also includes the data collection tool Cliqz uses to build its recommendation engine. Users who receive a version of Firefox with Cliqz will have their browsing activity sent to Cliqz servers, including the URLs of pages they visit. Cliqz uses several techniques to attempt to remove sensitive information from this browsing data before it is sent from Firefox. Cliqz does not build browsing profiles for individual users and discards the user's IP address once the data is collected. Cliqz's code is available for public review and a description of these techniques can be found here.

This section is a horrible write up of what happened. Over the years, among a lot of other privacy tech (e.g. [0][1][2]), we developed a privacy-preserving data collection framework we call Human Web [3]. The gist of it is simple: Users contribute data. There's no way to link any two messages with one another making it impossible to build profiles out of the data. In fact, most of the URLs are dropped thanks to these strict checks. Mozilla, Princeton University and Red Pen Team have audited the approach. The code is open-sourced [4]. Feel free to audit it, and also please feel free to use it in your projects. If you are genuinely interested in the approach, please read [3] and let's discuss details.

Here's the bigger issue. We have created an unhealthy, and wrong narrative around privacy vs data collection. It's a false dichotomy. We also wrote at length about this here [5]. Data from people, is not the same as personal data. Record linkage here is key, and we prevent it - even at a network level [6].

If you are interested to read more about the Firefox Integration (and the context in which it happened), read this: https://www.0x65.dev/blog/2019-12-11/the-pivot-that-excited-...

---

[0] Adblocker (fastest there is): https://www.0x65.dev/blog/2019-12-20/not-all-adblockers-are-...

[1] Algorithmic anti-tracking (first and only): https://www.0x65.dev/blog/2019-12-19/blocking-tracking-witho...

[2] Anti-Phishing: https://www.0x65.dev/blog/2019-12-21/anti-phishing-with-priv...

[3] Human Web: https://www.0x65.dev/blog/2019-12-03/human-web-collecting-da...

[4]: Human Web Code: https://github.com/cliqz-oss/browser-core/tree/master/module...

[5]: Is Data Collection Evil: https://www.0x65.dev/blog/2019-12-02/is-data-collection-evil...

[6]: HPN: https://www.0x65.dev/blog/2019-12-04/human-web-proxy-network...

Fundamentally, Google protects its search by owning all entry points to it (or paying massive fees for competition-prohibitive distribution deals.

The strategy can be seen at play with the Chrome browser, Android and its licensing model for hardware manufacturers, paying Firefox and Apple for placing search as default in their respective platforms. They know that once distribution barriers for search fall, their existence is threatened, explaining even moves like Google Fiber.

We wrote at length about this issue in our blog https://0x65.dev/blog/2019-12-22/google-competition-is-just-...

Reading big difficult books is offered as a way to teach reasoning from first principles. On how to go about it:

"We have our recommended ten-stage process for reading such big books:

1. Figure out beforehand what the author is trying to accomplish in the book.

2. Orient yourself by becoming the kind of reader the book is directed at—the kind of person with whom the arguments would resonate.

3. Read through the book actively, taking notes.

4. “Steelman” the argument, reworking it so that you find it as convincing and clear as you can possibly make it.

5. Find someone else—usually a roommate—and bore them to death by making them listen to you set out your “steelmanned” version of the argument.

6. Go back over the book again, giving it a sympathetic but not credulous reading

7. Then you will be in a good position to figure out what the weak points of this strongest-possible argument version might be.

8. Test the major assertions and interpretations against reality: do they actually make sense of and in the context of the world as it truly is?

9. Decide what you think of the whole.

10. Then comes the task of cementing your interpretation, your reading, into your mind so that it becomes part of your intellectual panoply for the future."

it's legal wise the same as a cookie

I am not sure that is the case. GDPR makes provisions for personal data that can uniquely identify users. Blanket statements like: Local storage is not allowed I think are misleading. The state is persisted in the client's machine. Unlike cookies, which get attached to all requests in the specified path, local storage items are not transmitted with the request. Furthermore, in the approach I recommended earlier, no unique identifiers are being sent with the request at all. I am pretty sure that is GDPR compliant, but would love to be pointed to legal provisions that would suggest otherwise.

it's a trade-off we are willing to accept.

Referrers, in my opinion, are not reliable enough to derive uniques, and I would assume (although I would not have any numbers to back it up), that the margin of error is very significant when you consider every condition under which referrers would not be sent (some very good cases when that happens are mentioned by other people in this very thread)

I think a more robust (and privacy-preserving) way to calculate uniques ( per time-frame as well) is to make use of the local storage / indexDB in the visitor's browser.

This is how it would work:

- Visitor X lands on your website.

- With some JavaScript you check local storage if `dailyUnique` for today's date is set.

- if yes, send `visit` signal

- if not set, send `unique visit` signal and set `dailyUnique` to local storage.

You can apply this to any analytics use case. Its is private and does not rely on referrers. We have been using this goal-attainment approach to do analytics at my company for quite some years.

Very interesting.

The iSmell failed to get the interest of the public. When looking at what went wrong for the iSmell, it is revealed that the missing link was a market survey. According to Startupover’s Andrea Dusi, the iSmell “was definitely a nice idea, but not a useful one”. DigiScents has shut down due to a lack of funding, although it still “continues to license its technology and is looking for funding for a relaunch. [1]

I wonder if with virtual reality, something like this will pick up.

[1] https://en.wikipedia.org/wiki/ISmell

Anolysis stands for Ano[nymized] [Ana]lysis. It's a new approach to do telemetry without sending unique identifiers (like most analytics / telemetry) systems do - but focus on goal attainment at a group level. This makes work harder of course, but it's a price we've been willing to pay. It is a pity you would take a domain name as evidence of malice. We should have a paper coming up at some point on the approach.

1) Trackers Stats.

This feature is powered by another project we run, where we measure the tracking landscape in the web (most popular domains): https://whotracks.me. Details on how that works can be found in our paper [0]. Also - we are flirting with the idea of providing a mode where the ranking is informed by the trackers in the destination site. Would love to hear your thoughts on whether you'd like smth like this.

2) Page previews (I'm not sure about whether I like that)

At the moment it's only a placeholder for a lengthier title and description (if available), but we are planning to use the space for rendering a short summary of the content/media in that site + similar sites in terms of content (query-relevant of course). This is more work in progress as we want to make sure content creators are on board. Again: would love to hear your thoughts on that.

Disclaimer: I work at Cliqz.

[0] WhoTracks .Me: Shedding light on the opaque world of online tracking - https://arxiv.org/abs/1804.08959

There was a reply to the parent comment (mine), that appears to be removed, but that we believe it still deserves to be addressed.

--------------

REMOVED COMMENT:

--------------

Is data collected by Cliqz Browser and your search engine used in any way to provide the MyOffrz marketing service? https://myoffrz.com/en/

Edit: according to the privacy policy the browser and the data collected through it is indeed used to deliver targeted ads.

https://cliqz.com/en/privacy-browser

This is the problem, we do not need any more ad companies that care about our privacy. We need browser vendors that are in the primary business of creating browsers, and possibly charging for advanced features.

We need search engines that do not collect our entire browsing history and how we interact with sites on the pretense that its needed to deliver better search results, while their business model actually revolves around using the same data to deliver targeted ads.

Cliqz could very well collect data and use it only to make search better, and release paid search products for businesses and customers, but they chose to make their users the products.

I had doubts about why people were so upset about your company, but now I see it. Cliqz is esentially capitalizing on the growing privacy movement to market itself, uses subpar technical solutions to ensure data privacy (routing sensitive data through FoxyProxy), and in the end delivers the same old service: our data is harvested, and we are delivered targeted ads.

-----

REPLY:

-----

You're cherry-picking and putting things out of context:

"the data collected is used to deliver targeted ads"

That's not said like this anywhere [e.g. in the linked privacy policy], because it would be false. We do targeted ads, but the targeting is done on the >>client<<, which is the only place that has your history information. It's not possible to do targeting outside the browser itself, because, unlike the rest, we do not have this information.

you collect browser history

Also incorrect. We collect URLs, but on isolation, we never have the full history. In the article we discussed at length why this is dangerous, of course, we would not do the same.

Uses subpar technical solutions to ensure data privacy (routing sensitive data through FoxyProxy)

There is a nice short reply here https://news.ycombinator.com/item?id=21696963, but just in case please have a look at the paper: https://arxiv.org/abs/1812.07927.

Cliqz could very well collect data and use it only to make search better, and release paid search products for businesses and customers, but they chose to make their users the products.

Everything we do is privacy preserving, the business model is no exception. We've been experimenting with paid products too (see: https://https://www.lumenbrowser.com/en/) - but it seems to me that if you had found that, you'd say "Yeah, but you want me to give credentials" - and suggest we live on donations instead. We need to be skeptical, and we thank you for that, but also be constructive and reasonable. Preventing anyone to improve the web, only helps companies that do not care about privacy.

Cliqz is esentially capitalizing on the growing privacy movement to market itself.

We do not do it for marketing. We started building products with privacy in mind since 2014, back when privacy was even more niche than it is today. (from Day 2: Is data collection evil? [0]

We understand that people whenever the see data collection they assume it's for the worse, and whenever they see ads, double on that. In a way, we are victims of the wrong-doings of people that has come before us. But we are trying to build good products maintaining the privacy of the people. If you don't trust, you can verify, we are transparent on what we send and in how we do it. I you still don't like us, well we cannot please everybody, but please do not accuse us of doing things that we are actively fighting against, unlike many, we do not claim that the world is wrong, we are trying to change it.

[0] https://0x65.dev/blog/2019-12-02/is-data-collection-evil.htm...

We could not agree more. We tried our best to make a case why the data or privacy dichotomy is a false one, especially in the context of building a search engine - one that is competitive and independent. Our rules are: no personally identifiable information (even minimize probabilistic attacks), minimize data collected to the bare minimum. This is what we use ourselves and what we want our families to use. We care.

It does come with challenges though.

1. It requires a change of mindset by the developers

2. Processing and mining data implies that code be deployed and run on the client-side.

3. The data collected might not be suitable to satisfy other use cases. Because data collected has been aggregated by users, it might not be reusable.

4. Aggregating past data might not be possible as the data to be aggregated may no longer be available on the client.

However, these drawbacks are a very small price to pay in return for the peace of mind of knowing that the data being collected cannot be transformed into sessions with uncontrollable privacy side-effects.

Disclaimer: I work at Cliqz (some of the comment comes from the article itself).

I agree. That data should stay with their rightful owners - the data subject.

I believe that we have fundamental issues with personal data ownership in the web for two reasons:

1. People do not believe web is real life. In the physical world, it is very easy to see how your rights are violated. If there's a person following you for days (when you shop, when you buy your medicine, when you talk to friends), you call the police. The majority of users have no idea of third party trackers. They have no idea what (or rather how much) information they are emitting at each point in time.

2. People have access to amazing products for "free", and they do not know the price of their data. Imagine if you were given a TV, but you would have to babysit a boring kid 2 hours a day for years. I guess not many would want that TV. Force each company to offer two plans (legally): one free + (ads / tracking), one premium (no ads / no tracking), then see how much people care about their data.