HN user

throwaway287391

672 karma
Posts1
Comments339
View on HN

You only need legislation like this to hold in one major market to make a big difference.

Could you explain this? I understand that for things like emissions regulations on cars, where it’s very expensive to invest R&D and production capacity into different SKUs for different markets. I can’t see why it would help in other markets in this case though (beyond potentially inspiring similar regulations elsewhere).

I really doubt it's that, as opposed to the maintenance cost of an extra flow to a boarding pass. Or perhaps just a perceived complexity/annoyance cost when something breaks in the desktop flow here and there.

I'd think it's only maybe 5-10% of customers at most who both use desktop over mobile to get their boarding pass and use an ad-blocker on desktop. And honestly I don't remember ever seeing an ad (even on Ryanair) when getting my boarding pass on mobile. OTOH I distinctly remember seeing many giant ads on printed boarding passes, most often on printed boarding passes brandished by other customers (usually printed in full color!). I'd think that's hugely more valuable as advertising real estate than the iota of additional data they get to collect on a few adblock users who have been forced to use mobile.

I guess it may be because I've only had a US passport but have primarily taken Ryanair to fly between the UK (where I live) and elsewhere in Europe, which is potentially an edge case they don't care to handle well. I recently gained UK citizenship / a passport so perhaps this will get easier for me going forward.

Either way, I stand by my main complaint about Ryanair letting late arrivers jump the queue even if I only have to "suffer" from it when I'm checking bags.

"99.9% of our passengers don't break the rules, they don't get penalised. The 0.1% of the guys who delay the boarding process, the guys who are there delaying the departure of the aircraft because their bag doesn't fit in the overhead (cabin), they are going to pay and we're going to eliminate them."

Somehow I doubt the compliance rate is anywhere near 99.9% -- that's roughly 1 passenger breaking the rules every 5 flights? Would they really be investing so much (including having the CEO spend time on air to rant) in catching the delinquents if that's the scale of the problem?

As someone who mostly follows the rules, the thing that really bothers me about Ryanair and the like are, I'll get to the airport at least 2 hours early as I'm supposed to, and then there'll be a massive hour+ long check-in queue (which I have to wait in just to show my passport, even if I'm not checking bags, since online check in never seems to work when I need to enter passport info), and all the while they'll have staff shouting "Anyone going to <destination of flight that closes boarding in 20-30 minutes>?" and shepherding those passengers to the front of the queue. It irritates me to no end -- why the hell should I bother arriving early if I'm just going to be punished with a longer wait for it?

But obviously in hindsight, it was the most expensive mistake I've made in my life.

Maybe a tiresome pedantic response, but this is only a "mistake" to the same extent that it was a mistake for [everyone in the world who had the required funds] not to buy 100 BTC in 2015. If you can overcome the cognitive biases (endowment effect, loss aversion, etc.), the fact that you previously owned the 100 BTC has no real bearing on the situation, beyond transaction fees and a couple hours max saved by holding (doing nothing) vs. buying.

That's kind of expected for a research intern -- internships are most commonly done within 1-2 years before graduation. But in any case, the fact that the first author is an intern is just the cherry on top for me -- my comment would be the same modulo the "especially" remark if all the authors were full time research staff.

As someone who used to write academic ML papers, it's funny to me that people are treating this academic style paper written by a few Apple researchers as Apple's official company-wide stance, especially given the first author was an intern.

I suppose it's "fair" since it's published on the Apple website with the authors' Apple affiliations, but historically speaking, at least in ML where publication is relatively fast-paced and low-overhead, academic papers by small teams of individual researchers have in no way reflected the opinions of e.g. the executives of a large company. I would not be particularly surprised to see another team of Apple researchers publishing a paper in the coming weeks with the opposite take, for example.

"There's nothing to prevent any state from funding the universities in their states."

I would've thought one major issue is that a much larger chunk of tax revenue is collected by the IRS than by any state. From googling, CA has the highest state income tax rate but still collects <5% of US federal tax revenue, while having >10% of the population. ~2.5x'ing state taxes to attain similar per-capita revenue would probably lead to a fair number of people leaving the state, or at least get the party who passed that tax hike (presumably Democrats) voted out in the next state election.

OTOH the NSF annual budget is $10B/year, in theory "easily" fundable by CA alone with its $220B/year in tax revenue, in the worst case with a 5% tax increase. The NSF isn't the only federal agency that funds research (seems to provide around 25% of federal research funding) but it is probably enough for one state, even the most productive one. So maybe it really is doable.

IME Netflix is a close 2nd best after Apple, which I don't think I can distinguish from a 4K BluRay. I've found that the quality depends on the platform a little -- for Netflix the native LG app seems to look best on my LG TV, while Apple looks best on the Apple TV app (perhaps unsurprisingly).

Amazon Prime 4K HDR on the other hand looks like garbage on every platform I've used -- the compression is unbearable in any dark scene.

I like JAX but I'm not sure how an ML framework debate like "JAX vs PyTorch" is relevant to DeepSeek/PTX. The JAX API is at a similar level of abstraction to PyTorch [0]. Both are Python libraries and sit a few layers of abstraction above PTX/CUDA and their TPU equivalents.

[0] Although PyTorch arguably encompasses 2 levels, with both a pure functional library like the JAX API, as well as a "neural network" framework on top of it. Whereas JAX doesn't have the latter and leaves that to separate libraries like Flax.

OpenAI O3-Mini 1 year ago

I did read that comment. I don't think that person is saying they were part of the study that OpenAI used to evaluate the models. They would probably know if they had gotten paid to evaluate LLM responses.

But I'm glad you pointed that out, I now suspect that is responsible for a large part of the disagreement between "huh? a statistically significant blind evaluation is a statistically significant blind evaluation" vs "oh, this was obviously a terrible study" repliers is due to different interpretations of that post. Thanks. I genuinely didn't consider the alternative interpretation before.

OpenAI O3-Mini 1 year ago

Yes, I am assuming they evaluated the models in good faith, understand how to design a basic user study, and therefore when they ran a study intended to compare the response quality between two different models, they showed the raters both fully-formed responses at the same time, regardless of the actual latency of each model.

OpenAI O3-Mini 1 year ago

Erm, why not? A 0.56 result with n=1000 ratings is statistically significantly better than 0.5 with a p-value of 0.00001864, well beyond any standard statistical significance threshold I've ever heard of. I don't know how many ratings they collected but 1000 doesn't seem crazy at all. Assuming of course that raters are blind to which model is which and the order of the 2 responses is randomized with every rating -- or, is that what you meant by "poorly designed"? If so, where do they indicate they failed to randomize/blind the raters?

I assumed if I kept reading there would be a line explaining why they can't simply raise prices until the demand becomes manageable with current staffing, such as "We sold all these suits with guaranteed adjustments for £[some heavily discounted number] for life", but I didn't find any such explanation. Shrug

Yes, this is correct -- TikTok's own "shutdown" was never required by law. I'm not in the US so I can't check for myself, but from googling it still seems to be removed from both the Apple and Google stores.

If Apple/Google don't change their minds, TikTok won't be able to get any new US users, and won't be able to distribute updates to current US users. To continue using it in the current state, US users will have to keep the same phone and TikTok will have to continue supporting whatever last version(s) they're on indefinitely. (Modulo the few that might jump through VPN and app store locale setting hoops.)

And I don't see how Apple/Google could change their minds: the ban bill comes with a 5 year statute of limitations, so regardless of how convincing the Trump administration is in their promise not to enforce the law, the next administration inaugurated in January 2029 would still be able to impose the penalties on Apple/Google for 4 years of non-compliance. Those penalties would be cripplingly massive even for the world's largest companies (I'm reading an estimate of $850B [0]).

As far as I can tell, the only events that could end this are (1) TikTok finding and agreeing to sell to a US buyer or (2) Congress overturning the ban.

It's odd that people are talking as though the saga is over now...

[0] https://www.theverge.com/2025/1/19/24347325/tiktok-service-p...

Yes, I'm so completely fed up with recurring subscriptions for things with negligible or no recurring costs for the seller. This one is of course particularly obnoxious given the hardware itself is expensive and the recurring cost is 0. But for example I would've gotten an Oura ring by now if they would just charge twice the price for the ring itself and not require a subscription, even though the subscription fees over the lifetime of the hardware would probably add up to a significantly smaller amount. To me it's just incredibly off-putting -- it reeks of greed and feels like a blatant attempt to fool customers by obscuring the actual cost. I guess it must be working for them, but for me, the cost of anything with a recurring fee gets mentally rounded up to "approximately $infinity".

Given that (as the article mentions) the ban essentially only directs Google/Apple to remove the app from their US stores, what's the rationale on ByteDance's part to immediately revoke existing US users' access? My naive assumption was they'd want to keep it going and support the current dead version of the app for as long as possible to continue squeezing US revenue for at least a few more months until that becomes untenable. Are they instead hoping to rally the user base into mass protests and pressure lawmakers into reversing the ban?

Reflections 2 years ago

FWIW OpenAI themselves give a reasonably specific definition of AGI in their Charter [1]:

highly autonomous systems that outperform humans at most economically valuable work

But I guess the "as we have traditionally understood it" bit from Sam's phrasing may imply that in fact he means something other than OpenAI's own definition?

[1] https://openai.com/charter/

You wouldn't need to use computer vision on a picture of the PDF. arXiv has the tex source for most of the papers. An LLM trained on code could do a pretty good job of translating tex to readable html with a bit of effort.

Gemini AI 3 years ago

A-ha, thanks! Hadn't looked at or heard of the referenced paper, but yeah, sounds like it's almost certainly also 5-shot then.

It would've been more consistent to call it e.g. "5-shot w/ CoT@32" in that case, but I guess there's only so much you can squeeze into a table.

Gemini AI 3 years ago

CoT@32 isn't "32-shot CoT"; it's CoT with 32 samples (or rollouts) from the model, and the answer is taken by consensus vote from those rollouts. It doesn't use any extra data, only extra compute. It's explained in the tech report here:

We find Gemini Ultra achieves highest accuracy when used in combination with a chain-of-thought prompting approach (Wei et al., 2022) that accounts for model uncertainty. The model produces a chain of thought with k samples, for example 8 or 32. If there is a consensus above a preset threshold (selected based on the validation split), it selects this answer, otherwise it reverts to a greedy sample based on maximum likelihood choice without chain of thought.

(They could certainly have been clearer about it -- I don't see anywhere they explicitly explain the CoT@k notation, but I'm pretty sure this is what they're referring to given that they report CoT@8 and CoT@32 in various places, and use 8 and 32 as the example numbers in the quoted paragraph. I'm not entirely clear on whether CoT@32 uses the 5-shot examples or not, though; it might be 0-shot?)

The 87% for GPT-4 is also with CoT@32, so it's more or less "fair" to compare that Gemini's 90% with CoT@32. (Although, getting to choose the metric you report for both models is probably a little "unfair".)

It's also fair to point out that with the more "standard" 5-shot eval Gemini does do significantly worse than GPT-4 at 83.7% (Gemini) vs 86.4% (GPT-4).

Diamonds Suck 3 years ago

After all, [engagement rings] only need to last a few months, or maybe 1-2 years, until the wedding.

I think this is the cultural distinction you're missing -- married American women typically wear both their engagement and wedding ring their entire lives, and the engagement ring is usually the more ornate/expensive one (which might be due to the De Beers marketing throughout the 20th century this article discusses).