HN user

latk

285 karma
Posts0
Comments88
View on HN
No posts found.

Author here. Yep, that's close to my thinking. I don't actually believe that Cursor (or similar tools) are completely shit.

But I worry that the Cursor team perhaps doesn't care whether their product actually delivers value. That they just want to sell the appearance of productivity.

This, to me, is a much bigger concern than everyday performance of their tool. Tools can be improved, organizational culture usually not.

But this is wild speculation. I didn't want to write this as the conclusion of the actual article, which tried to be more factual and to take their marketing at face value.

Offering a service to European consumers?

Probably not a big issue. GDPR compliance can be challenging without a suitable mindset, but it's not impossible.

* Consider that the GDPR has an extremely broad concept of “personal data” – it's not just identifying info but anything that can be reasonably linked to a person!

* Data minimization – only collecting what is needed, and only using it as actually needed – is already a great step.

* Writing a GDPR-compliant privacy notice can be a good exercise to understand what data you're processing for which purposes. Art 12–15 GDPR are the closest it gets to a checklist.

* And you'll have to implement “appropriate” security measures, but what is appropriate is largely up to you.

The more challenging part is ensuring that you're only using data processors/vendors that are contractually bound to use the data as you instruct, and that you protect “international transfers” where the recipient (e.g. vendor) is outside Europe. If you're looking for server locations in North America, I recommend looking at Canada since they have an “adequacy decision” from Europe.

You will have to be GDPR-compliant if you “offer” your service to people who are in Europe, i.e. actively market to such people, or have testimonials from EU customers, offer French localization, accept payment in EUR, and so on. Mere availability of your service is not an offer.

Offering a B2B SaaS service to companies that need to be GDPR-compliant?

You're fucked. There is no legally safe way for a company to use an US-based data processor, i.e. to engage you as a vendor. However, and this is your “get out of jail” card, many customers don't care, and will be happy as long as they can sign “SCCs”.

I can't agree, but maybe this is semantics :)

For something to be personal data, it must be information that relates to an identifiable natural person. There are two criteria here: (1) it must relate to a natural person, and (2) that person must be identifiable.

Your “loop over IP all addresses”example does not involve personal data because the information doesn't relate to anyone – it is just a list of numbers. Even if it were to relate to individuals, no court would order an ISP to disclose information about corresponding subscribers for such generated IP addresses. Then, the identifiability argument in Breyer cannot work.

In contrast, an IP address that is part of an IP packet received by a server clearly relates to the person sending the packet, if there is such a person. And, with the help of third parties, the person on the other end of the connection is reasonably likely to be identifiable. This does not depend on the website operator having any additional information such as cookie identifiers, other than the date. To avoid confusion, let me quote the relevant part from Breyer:

49. Having regard to all the foregoing considerations, the answer to the first question is that [Art 4(1) of the GDPR] must be interpreted as meaning that a dynamic IP address registered […] when a person accesses a website […] constitutes personal data within the meaning of that provision, in relation to that [website] provider, where the latter has the legal means which enable it to identify the data subject with additional data which the internet service provider has about that person.

The only additional data involved here is that held by the ISP, not by the website. That the judgement scopes its conclusion to website providers must be understood not as a limiting factor (as in: IPs can be personal data only for website providers), but as a contrast to the uncontested observation that IPs clearly are personal data for ISPs.

An IP address that relates to an identifiable person is personal data by itself. Thus, its mere disclosure to a third party without a legal basis is a breach of the GDPR. The article you linked highlights the “absolute vs relative” identifiability discussion, but this reasoning holds even under the “relative” standpoint because Google too is a website operator who has the same reasonably likely means for identification as the original website operator, if not substantially better means due to its trove of other data it can correlate with the IP address.

In this LG München case, the court determined that sharing this data with Google was illegal, regardless of whether there is any additional data. It is, in a sense, a very formal argument, that doesn't consider it necessary to dive into specific fact patterns (that's the abstract vs concrete means part quoted in my previous comment). The court did consider the impact of Google's tracking abilities in calculating damages, though.

To summarize my disagreement with your comment: (1) I assert that an IP address by itself can be personal data for a website operator (such as the defendant or Google), per the Breyer argument. (2) The LG München judgement in this Google Fonts case is not concerned about additional data when considering the legality of processing. (3) Additional knowledge held by the website operator is irrelevant for both this case and the Breyer judgement. Since a negative is difficult to prove but a positive can be shown by a single example, could you please point out the paragraphs in the Google Fonts case[1] or the ECJ's Breyer judgement[2] where I'm mistaken for disagreements 2 or 3?

[1] https://rewis.io/urteile/urteil/lhm-20-01-2022-3-o-1749320/

[2] https://curia.europa.eu/juris/document/document.jsf?docid=18...

Identifiability for IP addresses uses an even lower standard. The GDPR says that for something to be truly anonymous, there must not be any “reasonably likely” means for identification, even with the help of third parties, even when relying on additional information. There has of course been litigation about this, in the form of the Breyer v Bundesrepublik Deutschland case. It was based on the GDPR's predecessor law, but it used virtually identical phrasing so the conclusion still holds.

The European Court of Justice constructed a hypothetical scenario to show that identification can reasonably be likely. Let's say the website was attacked by a hacker. In a logfile, you find the attacker's IP address and want to prosecute them. So you report the incident to whatever authority is responsible for such incidents, which then gets a court order so that the attacker's ISP discloses information about the IP address. As long as the ISP knows to whom that IP was allocated at the time, there is now a reasonably likely chain of events that leads to identification of the person behind the IP address.

In this case about Google Fonts, the court says that it's sufficient if the website operator or Google have the “abstract means” for identification, not whether they actually did this for this plaintiff's specific IP address.

A solution would be if the EU forbids ISPs from keeping such logs, but given repeated attempts at mass data retention laws for national security purposes and pressure from the IP industry^W^W film and music industry for copyright infringement prosecution purposes, that doesn't seem likely.

Do you, as the website operator, have the right to copy and serve these fonts to your visitors?

All the fonts on Google Fonts are open source. When GDPR came into force in 2018 I downloaded all the fonts I needed, checked their licenses, and uploaded them on my servers along with necessary notices as required by the licenses.

The matter could also be sidestepped if the CDN were to offer a GDPR data processing agreement (DPA) and would make guarantees about the locations of servers. The free public CDNs understandably don't do this, and it seems Google Fonts is not covered by the Google Cloud DPA.

The court judgement addresses this exact point. There are previous judgements (Breyer v Bundesrepublik Deutschland) that establish that dynamic IP addresses are personal data. There are reasonable means to identify the data subject with the help of third parties, such as the ISP. “For this it is sufficient that the defendant has the abstract means for identification of the person behind the IP address. Whether the defendant or Google have the concrete means for linking the IP address with the plaintiff is irrelevant.”

That there is correlating information like timestamps, useragent strings, or referer headers increases the likelihood of actual identification, but the mere reasonable possibility of identification is sufficient for IP addresses to be personal data.

I discussed that argument over here: https://news.ycombinator.com/item?id=30139489

Summary: A company did try the “it was the browser, not us” argument in the “Fashion ID” case. The court did not fall for it. Data controller and thus responsible for compliance is whoever determines the purposes and means of processing. Being able to control what the website does seems to be good evidence for being a data controller.

In this Google Fonts case, the website operator didn't even try this discredited argument.

The fact is that CDNs and similar third party services play an important role.

They no longer do, since browsers implemented cache isolation.

if I "host" my fonts in S3 do I have to get consent for sharing IP with Amazon?

No, you're supposed to contractually bind your vendors/service providers as data processors with a contract (“data processing agreement”) per Art 28 GDPR. There's some debate around whether US-based companies are legally able of entering into such an agreement (say hello to the Cloud Act from me), but the general consensus still is that non-US cloud regions might be OK, and that CDNs that let you sign a DPA (like Akamai, Cloudflare, Fastly, …) are also OK. In contrast, Google Fonts does not seem to be covered by the Google Cloud DPA.

with every router that goes through tracert?

No, such mere transmission doesn't count as processing, and/or the intermediaries are responsible for their own compliance. In any case the connection should be protected by TLS so that only the client IP address + your domain name is visible to intermediate routers.

websites will add more crap "opt in" CYA forms

Unfortunately, I agree, though the point of this judgement is that self-hosting some assets is a perfectly cromulent alternative. I think relying on “consent” would be difficult in a case like this, since it is not generally possible to make access to a service conditional on consent to unnecessary processing activities. Using a CDN for assets like files is unnecessary.

I just wish that websites wouldn't force us outside of the EU to the asinine UX required by the EU

For EU-based websites there is no choice, as the law doesn't care about where the users are.

There's also a bit of irony in here that there has been a lot of work in replacing the cursed cookie consent requirements that gave us most of these annoying consent banners – but the past few months revealed that the US tech giants have been successfully lobbying against the proposed ePrivacy Regulation. So please redirect your ire against Google. Without them this might have been fixed in 2018.

This argument was tried in the Fashion ID case. A company had inserted Facebook Like buttons on the web page, and argued that it was not responsible for the ensuing disclosure of personal data (such as IP addresses or possible tracking cookies) to Facebook. See, it was the browser and not the website operator that disclosed the data, and the website operator never had access to the data in the browser in the first place!

The European Court of Justice did not buy this argument. By coding the website in a particular way, the website operator was responsible for causing the user's browser to act in a particular way, so it was the “data controller” for the collection an transmission of personal data by the Facebook Like button, though Facebook is of course jointly responsible for what their code does.

The underlying argument is that someone is a data controller and thus responsible for GDPR compliance when they determine the “purposes and means” of processing, alone or jointly with others. Embedding the code for the button was an exercise of this power to determine purposes and means. In contrast, the website operator is not a data controller for whatever Facebook does with the collected data on its servers, because it cannot control what FB does.

The given case from Munich is a very straightforward extension from the Fashion ID judgement, though the website operator didn't even claim that they weren't responsible. Instead, they argued that they had a “legitimate interest”in loading fonts from Google servers, which the court rejected. While I consider it probable that Google does not use data from Fonts servers for tracking, the judgement correctly points out that Google is well-known for tracking – but this doesn't matter anyway, since already the disclosure of personal data without a legal basis is a problem.

Careful. That is an 100% unofficial site. It is not chartered or funded by the EU. The linked article is from “Richie Koch”an editor working on human rights stories who wrote the article on behalf of Proton VPN, which runs the GDPR.eu site as a content marketing scheme. The linked article is not the law and not official guidance, though it provides a reasonably good summary.

Everything sqrt2 says in the comments is entirely correct, as far as I can tell.

JSON lets you write numbers. They can have a sign, decimal part, and an exponent. The standard euphemistically describes this as:

JSON is agnostic about the semantics of numbers. […] JSON instead offers only the representation of numbers that humans use: a sequence of digits. […] That is enough to allow interchange.

But can you encode/decode an arbitrary integer or a float? Probably not!

* Float values like Infinity or NaN cannot be represented.

* JSON doesn't have separate representation for ints and floats. If an implementation decodes an integer value as a float, this might lose precision.

* JSON doesn't impose any size limits. A JSON number could validly describe a 1000-bit integer, but no reasonable implementation would be able to decode this.

The result is that sane programs – that don't want to be at the mercy of whatever JSON implementation processes the document – might encode large integers as strings. In particular, integers beyond JavaScript's Number.MAX_SAFE_INTEGER (2^53 - 1) should be considered unsafe in a JSON document.

Another result is that no real-world JSON representation can round-trip “correctly”: instead of treating numbers as “a sequence of digits” they might convert them to a float64, in which case a JSON → data model → JSON roundtrip might result in a different document. I would consider that to be a problem due to underspecification.

If a study is observing how human reacts to a certain situation, that's research with human subjects. The Linux study observed how maintainers react to bugs, this CCPA/GDPR request spam observed how data protection staff reacts to requests about their processes.

And the backlash is not hypocritical. You're of course right that FB has also done really questionable research, but that doesn't matter here. I've also seen significant uncertainty about this spam series in the data protection/privacy community, i.e. criticism by those people who get to deal with these emails.

There is like 15 years of official guidance and case law on ePrivacy, with relevant guidance from the Art 29 Working Party (precursor to the current EDPB) published around 2014. But I don't think regulators are in a hurry to get into arguments about the finer points when the ePrivacy Regulation could be passed any year now, which would allow a more nuanced approach to cookies (e.g. allowing legitimate interest instead of consent).

TTDSG is finally a correct implementation of the 2005 ePrivacy directive. § 25 TTDSG literally just rephrases the exact ePrivacy requirements. The pendant to the above quote is § 25 Abs 2 Nr 1:

Die Einwilligung nach Absatz 1 ist nicht erforderlich, wenn der alleinige Zweck [der Speicherung oder des Zugriffs] die Durchführung der Übertragung einer Nachricht über ein öffentliches Telekommunikationsnetz ist oder wenn [sie] unbedingt erforderlich ist, damit der Anbieter eines Telemediendienstes einen vom Nutzer ausdrücklich gewünschten Telemediendienst zur Verfügung stellen kann.

You're quoting something about “5G Ultra Wideband”, which seems to be a brand name for mmWave. Yes, mmWave has very short range. But 5G isn't just mmWave. It's in many ways an evolution of LTE/4G, supporting the same frequencies and offering the same range, i.e. multiple km/miles. But it's up to carriers how they allocate their frequencies. To quote Wikipedia:

5G can be implemented in low-band, mid-band or high-band millimeter-wave 24 GHz up to 54 GHz. Low-band 5G uses a similar frequency range to 4G cellphones, 600–900 MHz, giving download speeds a little higher than 4G: 30–250 megabits per second (Mbit/s). Low-band cell towers have a range and coverage area similar to 4G towers.

https://en.wikipedia.org/wiki/5G#Overview

5G is _perfect_ for providing coverage in rural areas, except for the problem that 4G devices are incompatible with 5G networks. Starting 5G rollout in urban areas makes more sense because (a) 5G provides most benefit when clients are close together, and (b) because denser cells make it reasonably economical to maintain 4G coverage in parallel to 5G coverage.

Transport encryption is table stakes. It's really no longer something that can be mentioned as if it were something special. When I browse to a random website I don't think “wow HTTPS, so secure”. The channel client <-> service is encrypted, but the service still gets all of the data in plaintext.

On a technical level, Telegram is are as secure as Facebook Messenger. Both offer transport encryption and optionally E2EE secret chats. Actually, I might trust Facebook more (on a technical level) because they don't have Telegram's disastrous history with home-brewed crypto protocols.

The site does explain its methodology.

By default, it shows the costs when using the cheapest post-paid plan with at least 500MB allowance for at least 30 days – cheapest in absolute terms, not per GB.

You can toggle the option to show prepaid plans with the same parameters (at least 500 MB for at least 30 days).

But since it takes data from the ITU and not from the real market, the numbers do indeed seem inflated. One of the cheapest (in absolute terms) prepaid plans for mobile internet in Germany is AldiTalk at 4€/1GB/4 weeks, so that the pageload should cost even less than 1ct (0.0056 EUR). Similarly, your Telekom plans are for 4 weeks. Maybe they've excluded these plan because it's not for at least 30 days.

Every EU company has been having these compliance problems since the Privacy Shield invalidation in last year's Schrems II judgement. It is only Facebook that had the lawyers (and the gall) to sue the Irish DPC to prevent them from enforcing this judgement. Other authorities have already started enforcing, for example a small Bavarian company got a slap on the wrist for having used MailChimp.

But yes, the EU regulatory environment is definitely making it more difficult to cheaply trial some business idea. GDPR isn't part of the problem. By unifying EU data protection law it has become easier to target the EU as a whole. My gripe is more with things like the upcoming copyright directive which is a de-facto link tax.

It is primarily the authors view.

The proposed regulation – like many EU regulations – can also apply to non-EU entities. In this sense, the EU does try to exert extraterritorial jurisdiction.

However, this is constrained to the case where the non-EU entity targets people in the EU, so somehow participates in the EU market. The origins of this “targeting criterion” actually come from consumer protection cases, where it's easy to understand: if you advertise your goods or services to people in a particular country, you'll have to play by that country's rules.

The text in question does define more closely what it means to offer services in the EU. To lawyers (and to anyone who has experience with GDPR compliance) this is not a particularly vague statement. Admittedly, there's no unambiguous bright line definition, but there's a lot of jurisprudence on the matter.

In reality, the question is not whether EU citizens will use these services, but whether the operator of the service is targeting people in the EU, i.e. whether the operator intends or reasonably expects for EU people to use their service. A US service will most likely be fine if their reasoning goes something like this: (1) We primarily intend to serve connections from the US. (2) This expectation is reasonable based on our network topology. (3) But we don't care if someone else connects.

It would not be appropriate to exempt specific organizations since those organizations may change their targeting in the future. It already exempts most non-EU organizations, due to the criterion that they don't target the EU.

We had the same panicking in 2018 when the GDPR came into force and – quelle surprise – there are no fines for random international websites. The EU doesn't insert itself into your affairs if you don't insert yourself into the EU market.

When using a legitimate interest (opt-out) as a legal basis, the interest must be both legitimate AND outweigh the data subject's rights and freedoms. This requires a balancing test between the various factors to be performed first.

Similarly, you can't just legitimize anything with consent (opt-in) – the consent must be valid, and of course can't override more specific laws. You can't consent to something illegal.

So no, failing to use legitimate interest doesn't mean it's illegitimate or that consent could always be used. It could also mean that the balancing test failed, or that laws prescribe a different legal basis. E.g. the “cookie law”prescribes consent for non-necessary cookies and similar technologies.

They were not asking for consent in the meaning used by the GDPR. They are merely "asking" you to agree to updated terms, i.e. their contract with you.

GDPR allows processing of data under various legal bases. They use consent (opt-in) only for things like accessing your camera. For sharing data with other Facebook services, they rely on a "legitimate interest" (opt-out) instead. In theory, you might be able to object to processing under a legitimate interest, but they make it rather cumbersome. Which processing activities they perform under which legal basis is actually well-explained in the privacy policy, if you manage to find the correct section (it has a rather labyrinthine structure).

Since this is about cookies and IP addresses, GDPR is not the most relevant EU law. Instead, we have to look at the old ePrivacy Directive.

For cookies or any other access to information stored on the user's device, that access must either be strictly necessary for performing the service explicitly requested by the user, or consent is required (ePD Art 5.3). This is where those annoying cookie banners come from. LocalStorage isn't any different and would require the same consent as cookies.

For traffic data such as IP addresses, processing is allowed if it's technically necessary for the “transmission”, if the data has been anonymized, if it's required for billing purposes, or if the user has consented (ePD Art 6). There is an argument that security logs might be necessary, other uses like analytics are more dubious. The good news is that Umami seems to properly anonymize the IP address, so this part seems fine.

In cases where ePD mandates using consent, we cannot fall back to another GDPR legal basis such as legitimate interest. Of course this discrepancy between ePD and GDPR is a huge problem, and the promised ePD update has yet to materialize.

Yes, what you say is more precise than what I said. They did a lot of fixing and by now nearly all sites use the same codebase (incl frontend), apparently with feature toggles for some sites (e.g. Stackapps or Meta have unique mechanics but run on the same software). The only surviving forks are probably Area51 and their Enterprise product. Previously, frontend/layout had diverged though, which prevented rapid rollout of new UI features[1]. One of the perks of a graduated site used to be a custom design (not always just CSS changes).

Sites don't need separate database servers but separate database schemas, which is also why migrations between sites are still quite broken. Many processes around launching a new site appear to be very manual.

[1] https://meta.stackexchange.com/q/307862

When SO took off they thought they could replicate the success with other sites. This lead to the original trilogy (SO, SuperUser, ServerFault) and the Area51 site to propose new sites, of which many betas were started. Turns out few beta sites reached the engagement criteria to graduate to a full site, and this creeping scale with hundreds of site led to substantial technical debt. Each site was more or less a fork of the codebase, and required a separate DB instance.

The mistake was to treat each site as a separate community and expect it to eventually grow to an SO-level success. In contrast, Reddit's subreddit approach is much more scaleable and can tolerate niche communities. Reddit is strong not because of any single community, but because of the combination of communities to grab inbound traffic and keep users engaged. SO partially missed out on this.

At this valuation an acquisition is less likely. I would have rather expected an acquisition by Microsoft than a series E, but I guess it's too late now.

That would leave an eventual IPO as the only way out, but they'd need to break even first (or at least have a very compelling vision to break even soon). I don't see that happening with their current Teams product, which is an uncompelling entry in a crowded market.

Also, whether Teams eventually succeeds is largely unrelated to SO's unique value: the community. Teams is not valuable because of the community, it's just a spin-off from their public platform. They really need to find a better path to success, because wasting $85M on the status quo is not going to get them there. Their Careers product was more interesting in this respect, but it seems they failed to achieve sufficient scale there.

Or the sites that don't bother with compliance and just show a message to the effect of 'this site operates under a jurisdiction that may have different privacy laws to your country' and leaves it at that.

That's potentially but not necessarily compliant. To a large degree, it depends on the intent of the website's data controller.

* GDPR Art 3(2) discusses the territorial scope of data controllers that are not in the EU. Their data processing falls under the GDPR if they are offering services to people in the EU.

* GDPR Recital 23 discusses potential factors that indicate an offer. Blocking EU visitors is not necessary: “Whereas the mere accessibility of the controller’s, processor’s or an intermediary’s website in the Union, of an email address or of other contact details, or the use of a language generally used in the third country where the controller is established, is insufficient to ascertain such intention, factors such as the use of a language or a currency generally used in one or more Member States with the possibility of ordering goods and services in that other language, or the mentioning of customers or users who are in the Union, may make it apparent that the controller envisages offering goods or services to data subjects in the Union.

* The EDPB has issued further guidance on the territorial scope. In their guidelines 3/2018 [1] the spend a lot of ink on discussing this “targeting criterion”, and provide some clear-cut examples. Of course, that falls short of actually interesting examples of edge cases :)

[1] https://edpb.europa.eu/our-work-tools/our-documents/riktlinj...