HN user

esrauch

3,755 karma
Posts2
Comments1,671
View on HN
Qwen 3.8 4 days ago

I feel like the pelican test can't be relevant anymore; the whole point was to to something that wouldn't be in the training set at all and now it is?

Sure that part is a specific problem which is that new browsers basically want to declare their hacks profile against people who have already written sites and won't update them.

In practice when you make your new set of hacks the string you can always evaluate whatever cruft in the useragent today, but next browser shows up.

It does make the user agents insane but I don't know if there's any obviously better system for the problem, even with hindsight

principle, you could perhaps statically link protobuf into the CPython executable, but I really don't think Google is doing that.)

Yes, we actually do exactly this (not just protobuf but generally all of the other C code with it for the given application). But there are other approaches that can work.

But also now that header is deprecated anyway

It's deprecated exactly because the preferred/supported/default runtime is the upb based runtime. The odd gencode design is what made that transition possible as an implementation detail. So the deprecation is just the same topic here, that people successfully setting up the dynamic linking is not really realistic/viable unfortunately, and the effort for open source support is spent making the best behavior on the upb runtime which can't have that feature.

I just can't see how it could make sense to define features when a ton of the the behavior wasn't even intentional but just tons of bugs. I recall debugging issues that only reproduced in "IE7 compatibility mode of IE8" which didn't reproduce in either IE7 or IE8. And that was already after the standards were taken at all seriously

I feel like this is with 2026 view where browsers are so mutually compatible.

In the bad old days there were so many differences between html, css and js behaviors that if you wanted your site to be nice you had to change it for the browser. The way css padding worked wasn't even the same. Feature detection was rarely viable for any of this.

No user agent would probably have only entrenched IE6 dominance even more by blocking you from deliberately making a site that works at all on other browsers (including IE7 for that matter)

GoGo was not a completely separate implementation but deeply hooked into the official GoProto implementation. So it wasn't "Protobuf the binary wire format" or "Protobuf the schema language" which changed over time here, changes to the Google's Go library caused it problems. It's like building a library that integrates with Jackson (a JSON library) and Jackson details changed in ways that added toil, versus JSON changing.

Well, there's still two implementations even without getting into the quoted case. But yeah, narrowly about the C++Proto shared memory thing, unfortunately it's so hard to make this work for many different reasons that we don't even advertise support for this functionality at all. This is a bit in the weeds but inside Google we static link everything which makes is far easier to do it, and static linking is how we recommend you use C++Proto in general.

But its kind of a clear example of a surprisingly complicated technical space. The big picture is that it's actually not weird when a simpler thing which has less constraints can be better on some axes users when compared to a complicated thing that delivers on many advanced constraints.

If "readability of the .py gencode instead of the .pyi gencode" was the worst pain point (or pet peeve) here then I actually suspect Buf wouldn't have even bothered, it's just one of these that is easy to see and explain.

Engineer who works on Google Protobuf here, commenting as myself and not as an official statement.

It's great to have a healthy ecosystem in the world around Protobuf. Google can't possibly fill all use cases, there's many tools and Buf makes good tooling. The Protobuf team at Google intentionally tries to enable an ecosystem around Protobuf including examples like this. Google Cloud APIs are intentionally usable with any compatible thing that can understand Protobuf encoding including this one.

Kudos to Buf for making something that I'm sure a ton of people will find useful and which takes conformance so seriously.

Just to chime in with some context about Google's own implementations here though (since that's a lot of the discussion otherwise).

Google definitely takes Protobuf seriously including for the long term: you can't really understand how engrained it is within the Google stack without seeing it for yourself. It's not just RPC layer, it's storage, logging, FFI. Html templating is driven off Protobuf messages. Systems which interact with bank XML based systems uses Protobuf schemas. Internally it's widely used for in-memory library api types even without any direct/obvious connection to serialization just because it makes internal details like logging easier. This extremely large surface does create constraints and use-cases to balance. You can see Buf's reported numbers reflect that it is faster for a usecase they expect is typical, but at scale users do fall into the other buckets shown, affecting the performance of preexisting code is a major concern for our implementations that a greenfield implementation doesn't have.

Wide exposure in critical paths alongside long term support directly causes some quirks: for example some of our APIs followed PEP8 when it was created but PEP8 changed. It looks stupid that we have wrong style APIs but also it would be stupider to break compatibility for style reasons. JavaProto as another example still supports Java8 and the runtime is compatible with 2014 gencode which is a pretty major constraint.

Google Py Proto implementation has one extra interesting choice of the same gencode is reused with 3 different implementations (upb, a complete pure python one, and one that uses C++Proto as the in memory representation which libraries like TensorFlow can use to share memory between Python and C++), which is why design the way that it is with runtime created classes, the pyi files are readable but the .py files not.

This definitely has pros and cons, and the direct approach taken by Buf here really makes a ton of sense. It's just that Google's maintained implementation falls into a different spot in a larger technical tradeoff space.

If you see things that appear to make no sense with the official implementations, feel free to file an issue on GitHub and we can look, sometimes there is no reason and we can fix it, and sometimes there's a reason which we can explain.

Kudos again to Buf here, I'm fully sure this will solve some set of real business needs better than Google's (but not because Google isn't maintaining our offerings too).

98% isn't much 15 days ago

I wouldn't conflate old computers and old browsers. I still use an over 10 year old laptop and it still has a latest browser.

I'm not trying to agitate you here and won't keep looping on the same question after this reply: I just do not understand the threat model here, and you're still being oblique as though its obvious what the threat model by being sarcastic instead of just spelling it out in plain language.

Are you worried about some foreign state actor like Israel targeting you specifically, hacking your devices to listen to you chatting? Or you're worried about US warrantless mass surveillance wiretapping all citizen's smart speakers, and you're worried the US government may spin up such a program?

In the latter case, the scenario you expect will happen is we'll have ~100 million US households are live wiretapped 24/7 without anyone knowing, you'll be part of the remainder living your life blissfully wiretap-free thanks to not having a smart speaker?

The Vespa at 80 16 days ago

Nissan Leaf motor is silent, there's really just objectively not a noise from the motor itself which could offend you.

There's still some possibilities here: maybe you came across someone with a broken car and believed that to be representative.

Most other explanations here just involve some form of confusion on your part to be honest: that it was the backup-alert noise that it makes _because_ the car is otherwise too silent, that it wasn't actually a Leaf but an ICE engine that you saw, or something else.

Sorry, you're going to have to spell out the risk for me here. What other cases have we seen that indicate that mass surveillance via smart speakers is a risk?

We all also have phone in my pocket 24/7 and my laptop on my desk, both with microphones in them. In the event of the government doing warrantless spying on all devices it seems like that is a strictly higher ROI target for them?

The Vespa at 80 17 days ago

There is literally any noise from the engine, whine or otherwise. Whether a noise is bad or not can be subjective but silence is objective.

So at least for that one it kind of draws into question your overall conclusions here. You might be hearing an ICE and thinking it is electric, or some specific badly tuned vehicle or something.

Can you spell out the ramifications for the plebs?

As far as I can tell home smart speakers are being used for warrantless mass surveillance, unlike Flock for example. Do you mean the possible future situation where they are?

I believe the feature is that you have a pending unreleased video and go to an llm for tips. When getting the tips it uses the pending video content and your recent videos info as context. So there's no holding back unlisted info short of not letting the user use it for their upcoming videos at all

And then the attack is to trick this recommendation system into putting a link out

I actually the attack is very likely already soft defeated by an interstitial telling you that you are leaving the site though, it would be weird if they didn't do that in general from this surface

In 2026 things have changed, there's literally whatever tens of thousands of "security" reports that are almost all bogus as a raging crap river.

I think theres very little chance this particular report made it to any engineer who works on product at all, because if they did they would be completely overwhelmed by reports, the filter which has to handle the many thousands of reports based on a playbook almost definitely filtered it out before it made it that far.

The Vespa at 80 18 days ago

What noise pollution do you hear from electric? You mean the noise cars make when they are backing up or what?

Claude Fable 5 1 month ago

It seems it is just like macOS releases, they have a number and they give the numbers arbitrary names to refer to them?

Fingerprinting to detect bots seems mostly relevant for things which are not DOS, so that percentage doesn't seem like the relevant one.

Bots manipulate review scores, posting link spam to other users, crawl your database that isn't open to crawl, etc.

Bot protection with fingerprinting is just an illusion. Any signals like this which is on client side can be spoofed by an above average person.

At the upper bound, fraud can always be committed by paying real people with real accounts to perform the desired action in a way that is 100% truly indistinguishable from organic. There's fundamentally actual prevention technique at the limit.

So the entire game is only "increasing the costs until it's not viable ROI", not "holistically prevent", which is why fingerprinting is a relevant technique here.

It has scrollbars, but there's benefits to having more on an individual page at once. The tradeoff point seems unclear but everyone must recognize some tradeoff there.

Foldable maps allow for getting everything on one view by having the final display area be enormously larger, which isn't an option on laptop screens.

Let's say you are writing into a byte[] and have a LEB128 length-prefix followed by a payload, but that determining the length actually involves nontrivial encoding work. For example, you have a UTF16 string and want to write out a UTF8 string, you want to go over the characters and write them out, but the UTF8 length is not known without doing all of that work.

If you can choose a fixed number of bytes for the length prefix, you can skip that number, do the encoding and find out the length, and then come back and fill in the length-prefix after.

But you actually don't know how many bytes it will take without doing all of the work to know the payload length (since larger payloads take more bytes to represent the length).

If you allow overlong representation you can reserve a few bytes and sometimes it'll just be the effective no-op bytes. If you don't, you won't be able to.

"Moxes/Sol Ring. They are a nice touch if not found in abundance."

Seems odd when followed by every 40 card deck having all color-relevant moxen and sol ring...

There was definitely sync bugs with replays at various points.

There was even desync bugs even in live multiplayer games; there was detection that it desynced which would end the game, which in turn meant exploits that would intentionally cause a desync (which would typically involve cancelled zerg buildings for some reason).

I think "offer unlimited but TOS ban behaviors that would cost too much to support" is actually a very normal way that things work instead of "raise prices until equilibrium is reached", including in credit cards. Credit cards do simply ban people they think are "rewards churning" based on a completely subjective TOS policy for example.

Raising prices is a bad strategy if you have a smaller base that costs enormously larger than the rest. "A million users that cost $1 and one user that costs $10 million, charge everyone $10 equilibrium", you're screwing over almost all of your users. The $20/month sub price is basically just not trying to capture the openclaw users, it doesn't make sense that all of the vanilla Claude users should subsidize them (and in fact it wouldn't even work because they will just go to Gemini or ChatGPT if your cheapest paid plan was very expensive to try to subsidize the other users)