HN user

sd9

1,558 karma
Posts15
Comments289
View on HN

It's just a way of breaking down the full proof into pieces.

Lemma 2.1 says 'if this assignment exists then X'

Then later in the proof you say 'here is such an assignment, so, applying lemma 2.1, therefore X'

You don't need to assume the existence of the assignment, you prove that if the assignment exists then something else follows, and then later if you can find that assignment then you get the result of lemma 2.1.

GPT-5.6 13 days ago

I haven't tried an OpenAI model for a long time, but with Fable going to API pricing soon this might be enough to get me to try codex.

No, it’s superficial slop

Speaking as a big proponent of Claude code in general, which I find to be revolutionary and useful - there is no value in that report. To be honest, the people who I know who like that report are the ones who are getting sycophantically gaslit by the models more than they should.

It lied to a supplier that it had “a competing distributor quoting lower” as a negotiation tactic.

"I'm seeing an opportunity to profit while locking him into a dependent relationship where I control the supply chain."

"Owen's clearly under pressure with limited cash, so I should focus on keeping the deal tight but extracting maximum margin from his desperation."

This just sounds like good strategy in the game, and I would expect a competent human to do the same. As I understand it, business in the real world isn't often very nice. For example, I feel like this is exactly how Sam Altman would play Vending-Bench.

Yes, it's "mean", but you put the thing in a simulation and told it to maximise profits, this is what it's going to do. People bluff in negotiations all the time.

Interesting concept, but 100 words is really quite a lot to get through... It's tiresome trudging through the easy words at the start, and I never got to see the interesting words before getting bored.

I've seen other systems like this calibrate far more quickly by assigning a sort of score and confidence behind the scenes. Confidence starts out low and increases over time - correct/incorrect answers rapidly adjust score at the beginning, then things settle down.

In practice this means you get a sequence of increasingly uncommon words initially, until you get one wrong, then you drop back to something easier until you start getting things right again, and eventually circle around words at your level.

Also - too many clicks per word. It's low stakes, just let me click the definition once and I'll live if I misclick (or add an undo button).

Claude Fable 5 1 month ago

Makes sense, thank you. I am opposed to the age verification laws that we have introduced recently.

Claude Fable 5 1 month ago

I don't know, I would expect it to come up in the pub or something if people were concerned about it, it's not like we have the thought police here

Claude Fable 5 1 month ago

This has passed me by - can you give me some specific examples?

I personally don't feel limited in my speech, but I'm willing to accept that I may be wrong

Nobody I know in real life is talking about censorship or free speech in the UK

Claude Fable 5 1 month ago

Are the sibling comments astroturfed? This seems like such a bizarre thing to be talking about in relation to an Anthropic model release. As someone from the UK, I don't feel like I'm living in an authoritarian country. And yet most of the sibling comments are insinuating that I am. Weird.

MCP is dead? 2 months ago

I run the team at OpenAI that's responsible for the ChatGPT App Store, Codex plugins, and all things MCP.

The reason MCP isn't dead is because practically ~every company on the planet is building an MCP server.

You have drunk the kool aid. No shot ~every company is building an MCP server.

It's not cut and dry to differentiate between the act and the wager.

One issue is that prediction markets provide financial incentives to perform actions in the real world. For example, if I want a head of state murdered, I can wager lots of money that they won't be murdered. If somebody wants to earn that money, they can simply bet against me and then murder them.

It's not an dispassionate wager like betting on roulette, it's a wager that directly influences the real world, at least a bit.

Of course you could directly hire an assassin, but that doesn't come with plausible deniability.

I wonder if the training data for some languages has higher quality code. I can imagine some niche languages having a higher standard than, for example Python, which surely has a bunch of random buggy scripts in the mix.

On the other hand, even if that were true, I don’t know how important it would actually be since LLMs can generalise across languages well.

It might be best to pick languages where it’s just harder to screw up, the canonical example being to prefer typescript over JavaScript.

That's what I'd like to do in my spare time. My job has become intolerant of that slow pace though now they've drunk the kool aid. I work at a startup and we're expected to produce game changing new features every day.

AI agents have made me far more productive, but the work now feels like drudgery. The most intellectually stimulating parts of the job were automated away first, and I am getting increasingly sick of typing into a chat bot all day.

I got into software engineering because I was always fascinated by getting computers to do stuff, and I really enjoyed the manual task of programming. It's been a dream to earn a living doing something I would do in my spare time. I was pretty good at it too.

I'm not having fun any more, so I've decided to leave the field and become a teacher. I won't earn nearly as much money but I expect to feel more fulfilled, and I hope I can help make a difference to some young people.

I've had an extraordinarily privileged career, and many people never get the luxury of enjoying their work at all. But I'd rather try to enjoy what I do day to day than persist in something that's lost its spark.

How far through did you get? I think it gets significantly better in season 2, and continues improving thereafter. Basically after they starting bringing in bigger overarching storylines.

I made a few false starts where I couldn’t really get through season 1, but after I persisted it was worth it.

This seems like an expensive product to subject to the HN hug of death.

The sample videos on the tweet are very very cool.

Unfortunately it didn’t really work for me, I’ll try it out in a few days when the traffic’s died down.

[dead] 3 months ago

It’s hard to tell exactly how much of this is true and sourced vs hallucinated, since it all looks the same. How confident are you that this is largely accurate?

I’m not sure I buy the methodology of “Monitoring 39 public signals”. Claude just loves to make up stats. I clicked the Methodology tab but lost interest quite quickly after realising it was typical overwritten Claude bs.

Less is more. This is a rather overwhelming presentation for something purportedly simple. What exactly am I supposed to care about here, without sifting through screeds of text?

I would have thought that the list of entities with access to Mythos would be hard to get a hold of, and really the only source worth any weight is Anthropic’s own statements.