HN user

trunnell

1,262 karma

2009-2019 at Netflix.

Posts2
Comments217
View on HN
Precursor 9 days ago

I dislike bots as much as anyone else... when weird inquiries come through my company's lead form, it costs some time and attention to sort them.

But what makes Cloudflare so confident that automation always equates to "fraud and abuse?" If I send my agent to go retrieve some information, do they consider that fraud?

If I block various ad trackers does that trigger their "bot detection" incorrectly? Do I have any recourse? Or is Cloudflare appointing themselves judge, jury and executioner?

And let's not forget this little chestnut: > 4. Privacy by design. Precursor was designed to collect signals that help to distinguish human patterns from automated and abusive patterns.

Ahh, so to "protect" against bots they're standing up a whole new regime of user surveillance and session-level monitoring. And they definitely won't be selling that, they promise. Got it.

This crap should be illegal. In the real world, I can authorize others to act on my behalf. The same should be true with software agents.

Hooray! Glad everyone came to their senses and we can all get on with business.

I bet it'll continue to be messy at the frontier for the foreseeable future as society gradually wakes up to the consequences of strong AI.

Will It Mythos? 29 days ago

Maybe you mean that an expert will use more specific language which in turn triggers the model to give a response that more closely matches the "expert distribution"

Anthropic published a study showing that Claude does more work for the expert user, and experts have a higher rate of "successful sessions" than novices.

https://www.anthropic.com/research/claude-code-expertise

honest angle is that the industry [felt] exposed by the residual risk [and] have a months-long bugfixing backlog exposed by Glasswing

Two problems with this theory.

1. Amazon complaining to the White House wouldn't have been the opening salvo. Amazon and Anthropic would find it much easier to talk to each other than go through the White House. We'd need evidence that Amazon (and probably others) already asked Anthropic to not release a Mythos-class model but Anthropic released it anyway. Are they on record saying this?

2. The jailbreak Amazon found needs to be real. Maybe the White House staffers are not AI experts and they don't really understand what a jailbreak is... but it's much harder to make that claim about Andy Jassy. For the jailbreak to be the real reason for the export control order, the jailbreak would need to be significant and cause material harm to Amazon. Then Jassy might pass it along to the White House assuming he already was refused by Dario.

But there is no evidence the jailbreak was real. There is one story that it amounted to a request, "fix this code." In any case, Anthropic is on record saying the so-called jailbreak didn't enable any vulnerability work that couldn't already be done by other models.

What an amazing achievement by America's adversaries.

The Trump administration lists Anthropic as a security risk and kneecaps its best model, despite the fact that compared to the other frontier US labs Anthropic is more transparent, more safety-oriented, frequently honest to a fault, and is clearly acting with patriotic intent.

Meanwhile, the same administration is hesitating to counter certain Chinese companies' efforts of industrial-scale theft and sabotage due to a fear of angering the CCP!

This administration has it exactly backwards. 4.5 months until election day, 7 months until the next Congress is sworn in.

Oh, I agree distillation isn't stealing "outright" as in it's not theft of 100% of the model. But there's a reason they're doing it. I didn't say anything about Chinese labs innovating -- obviously they are.

What accounts for the difference between your attitude that distillation is no big deal, "common practice," yet Anthropic sees as it as a huge threat?

"Anthropic accused Chinese firms of 'industrial-scale distillation attacks' on its AI models."

"Distillation involves training less capable models on more advanced ones’ output, and can be used illicitly to acquire powerful capabilities cheaply. The AI startup accused China’s DeepSeek, MiniMax, and Moonshot of generating 'over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts,'"

https://www.semafor.com/article/02/24/2026/anthropic-accuses...

After reading their posts and watching interviews with Dario it's abundantly clear that they view Chinese-lab distillation of US frontier models as a threat to US national security. You can argue with them about whether that is true, but not whether distillation is real.

I'll defend Anthropic.

They are clear about the reasons for guardrails: prevent their models from doing harm in dual-use contexts including CBRN or by accelerating research in authoritarian-backed AI labs.

What is the critique against that? It seems pretty reasonable to me. You want AI-accelerated biological or radiological experiments running in your neighbors backyard? You want PRC-backed labs to continue to steal Anthropic's models via distillation?

Mitigating the harms of dual-use tech is notoriously difficult and fraught with trade offs. What I would want to see is cautious rollout and quick response, which is EXACTLY what they're doing.

Instead, this thread is full of bad-faith arguments about Anthropic being dishonest, making a "useless" model, or "the power is going to their heads." You can't read Anthropic's System Cards and come away with any of these impressions. Quite the opposite, in fact. They are honest to a fault, acknowledging problems they discovered even when it hurts them.

If your harmless request was downgraded to Opus, you're billed for Opus. They were 100% clear about that. I'd much rather have a Mythos-class model that falls back to Opus 10% of the time than be capped to Opus 100% of the time. If that doesn't work for you, then make a suggestion for something better!

If you are a white-hat security engineer hitting guardrails, I don't think you have standing to complain. I really don't. Their Glasswing program actually got banks and the industrial sector to take action to fix security vulnerabilities. Do you realize how special that is? A huge portion of the economy runs on vulnerable code and has for decades, despite security experts testifying to Congress, begging business leaders, pleading for intervention-- with no results. But suddenly they're all enrolled in a program that will find *and fix* vulnerabilities! White-hat security people should be rejoicing. Instead some of them are throwing rocks. Unbelievable. Shameful.

Meanwhile, society is screaming at the AI labs to be more conscientious about potential harms of AI. Legislatures are passing laws limiting data center construction. There are protests. And you, the HN community, the vanguard of our profession, have the temerity to demand "NO GUARDRAILS!" "HOW DARE YOU TRY TO PROTECT DEMOCRACY!" "MY SOFTWARE PROJECT IS MORE IMPORTANT THAN KEEPING NUKES AWAY FROM THE BAD GUYS!"

Go ahead HN, downvote me. It'd be an honor.

We've heard:

- It can make kids "overconfident when they see material they think they already know, so they end up not engaging."

- Some programs, particularly RSM, are criticized for valuing speed over depth. Current culture for K-8 math teachers is the opposite, they value depth over speed.

Left unsaid:

- It can make the teacher's job harder when the class has a wide span of abilities.

- Current teaching culture is skeptical of accelerating and/or skipping grades in math.

Notably, we've never heard English teachers be upset about a kid reading a book outside of school that's above grade level, or using advanced vocabulary in an essay. They tend to praise it.

I'm in the SF bay area w/ middle school and high school age kids.

Between San Jose and San Francisco, 15%-30% of kids are in private school (it's 30% in SF where the public schools are extra dysfunctional). That's far above the California statewide average of 8% in private school.

Among our peers, somewhere between 1/4 and 1/3 of kids are doing advanced math outside of school, typically either Russian School of Math or Art of Problem Solving. This group only partially overlaps with the private school group. This is happening despite the fact that both public and private school teachers strongly discourage math outside of school!

So by decelerating math in the public school, incentives were created for privileged parents to take matters in their own hands and put their kids into programs that accelerate math education far beyond what public schools used to do. We now have a system that is creating even wider disparities in outcomes. It stands to reason that it's producing far less equitable outcomes, too, given that extremely bright kids who happen to be in lower-resourced schools have fewer opportunities. Universal screening for giftedness, advanced public school math courses, and the SAT -- all avenues for advancement regardless of background -- were all eliminated.

The Codex App 6 months ago

How about, "tell the agent to write instructions for cloud deployment with a cost estimate"

GPT-5.2-Codex 7 months ago

That’s for future unreleased capabilities and models, not the model released today.

They did the same thing for gpt-5.1-codex-max (code name “arcticfox”), delaying its availability in the API and only allowing it to be used by monthly plan users, and as an API user I found it very annoying.

GPT-5.2-Codex 7 months ago

Why aren’t they making gpt-5.2-codex available in the API at launch?

I hate to tell you this, but it might not be the full story that making search easier is "all it does."

Why doesn't the macOS App Store game search include results from Steam? That would be a very consumer-friendly thing for Apple to do, right?

The answers to both questions are related.

Yeah, it's ok, can't win 'em all. Lots of negativity in this thread. Maybe people have a gut feeling that "Netflix buying WB" fits into the preexisting narrative about media consolidation, and they're reacting negatively to media consolidation being a problem. I think that's more of a problem in the news media than in entertainment media. In entertainment, the bigger story is the tech-centered transitions, esp. to internet distribution. I don't think the consolidation narrative is a perfect fit in this case; this is a pretty different type of consolidation than the others in recent memory.

I think this is about Netflix's model reflecting a fundamental technology shift; any company not participating fully in that shift will be operating less and less efficiently compared to those that are. Look at the inside history of HBO's attempts to build a streaming platform; in the early 2010s their leadership knew they probably should, but were their hearts in it? Did they have executives with competence in this area? No, they outsourced it and mismanaged it. Repeatedly. But like you said, my view includes being a former Netflix employee so maybe I'm biased.

I don't have current information on whether or to what degree studio production capacity is a constraint. Content spending was publicly projected to grow, so studio capacity had to grow, which is why Netflix decided to build giant new studio facilities in New Mexico and New Jersey. Those were referenced in the Q&A Netflix held Friday morning [1]. Wild guess: Netflix's own studios run at full capacity, which is why they're continuing to expand them. I'd love to know if WB studios run at capacity.

I assumed their interest was strictly a content play and the extra studio space might actually be an anchor they were willing to drag along to get the content/IP.

Doubt it. Like I said, I'm not an insider on that question and I'm 6 years out of date. But if I had to guess, it would be that WB studio capacity will be a highly productive asset for Netflix -- most likely, it will be more valuable connected to Netflix's global distribution model that it was when operated under WB's model.

[1] Q&A transcript https://s22.q4cdn.com/959853165/files/doc_events/2025/Dec/05...

Have you considered the possibility that much like App Store rules, Apple's requirements for "catalog indexing" go far, far beyond the Netflix catalog merely showing up in TV app?

Perhaps the judgement about Netflix being anti-consumer might be hard to sustain if you could more fully inspect the details of what Apple requires.

Commenters here seem to be missing the larger David vs. Goliath story...

Netflix was a silicon valley start-up with a tech founder (Reed) who teamed up with an LA movie buff (Ted). They tried to solve a problem: it was too hard to watch movies at home, and Hollywood seemed to hate new tech. The movie industry titans alternated between fighting Netflix and making deals. They fought Netflix's ability to bulk purchase and rent out DVDs. Later, they lobbed insults even while taking Netflix's money for content licensing. Here's Jeff Bewkes, CEO of Time Warner, in 2010:

"It’s a little bit like, is the Albanian army going to take over the world? I don’t think so." [1]

Remember: this was the same movie industry that gave us the MPAA and the DMCA. They were trying to ensure the internet, and new tech in general, had zero impact on them. Streaming movies and TV probably wouldn't exist if Netflix had not forced the issue.

Netflix buying HBO is significant, but also just another chapter in this story of Netflix's internet distribution model out-competing the Hollywood incumbents. Even now in 2025, at least 12 years after it was perfectly clear that streaming direct to the consumer would be the future, the industry is still struggling to turn the corner. Instead, they're selling themselves to Netflix.

I was at Netflix 2009-2019. It was shocking how easily our little "Albanian army" overthrew the empire. Our opponents barely fought back, and when they did, they were often incompetent with tech. To me, this is a story about how competent tech carried the day.

Netflix has been rapidly buying and building studio capacity for a decade now. Adding the WB studio production capacity is a huge win for Netflix. It makes those studios more productive: each day of content production is now worth more when distributed via Netflix's global platform.

Same with WB and HBO catalog and IP: it's worth more when its available to Netflix's approx 300 million members. Netflix can make new TV and films based on that IP, and it will be worth more than if it was only on HBO's platforms.

[1] https://www.nytimes.com/2010/12/13/business/media/13bewkes.h...

Having a bad manager in past roles can be some of the best "manager training."

If one your past managers did something recommended in this article but it caused problems, that's ok! It just means you have seen another failure mode that the author didn't experience.

I remember being in a meeting with a bunch of the best managers at a former company. "Why did you originally want to be a manager?" was one of the first questions passed around the circle. The most common answer was, "I had this one really bad manager and I figured that surely I could do better."

They need to focus on fixing reliability first.

Maybe. What would you rather have?

A) rock solid Sonnet 4 with Sonnet 5, say, next April

B) buggy Sonnet 4 with Sonnet 5, say, next January

Seems like different customers would have a range of preferences.

This must be one of the questions facing the team at Anthropic: what proportion of effort should go towards quality vs. velocity?

https://status.anthropic.com/incidents/72f99lh1cj2c

They recently resolved two bugs affecting model quality, one of which was in production Aug 5-Sep 4. They also wrote:

  Importantly, we never intentionally degrade model quality as a result of demand or other factors, and the issues mentioned above stem from unrelated bugs. 
Sibling comments are claiming the opposite, attributing malice where the company itself says it was a screw up. Perhaps we should take Anthropic at its word, and also recognize that model performance will follow a probability distribution even for similar tasks, even without bugs making thing worse.