HN user

217

27 karma
Posts6
Comments25
View on HN

It's vibes mostly. Like, fable is better on frontend while 5.6sol is better at long running tasks with a lot of subtasks, harness did made them mostly interchangeable as well as able to have them both participate in solving a hard issue. Longest to learn is the ui (I'm more of a GUI guy but this is just too good) and all hundreds of features where you want to do something and then discover the harness already supports it out of the box but you never knew it, yet

GPT-5.6 13 days ago

Consensus itself does NOT matter, omp is objectively the best harness for power users yet it has 0 hn posts about it, zero.

You're fully free to use and try anything and without caring about what others think is right

98% Isn't Much 16 days ago

while true, the people who will read this and then think twice about implementing and applying things are exactly the people who already doing too much thinking

Claude Fable 5 1 month ago

So essentially there are 2 models, Mythos and Fable, they have the same weights but Fable is very safety-nerfed, and only ultra authorized companies have access to mythos with full capabilities

Reported benchmarks:

swe-bench verified mythos 5: 95.5%; fable 5: 95.0%

swe-bench pro mythos 5: 80.3%; fable 5: 80.0%

terminal-bench 2.1 mythos 5: 88.0%; fable 5: 84.3%

gpqa diamond mythos 5: 94.1%

riemannbench mythos 5: 55.0%; mythos preview: 43.0%; opus 4.8: 34.0%

arxivmath mythos 5: 78.5%

critpt mythos 5: 28.6%; gpt-5.5: 27.1%; opus 4.8: 20.9%

graphwalks bfs 1m mythos 5: 79.4%; mythos preview: 74.3%; opus 4.8: 68.1%

humanity’s last exam mythos 5: 59.0% without tools; 64.5% with tools

browsecomp mythos 5: 88.0% single-agent; 93.3% multi-agent

osworld-verified mythos/fable: 85.0%

gdp.pdf fable 5: 29.8% strict pass; mythos 5: 87.6% with tools on mean criteria pass

officeqa pro fable 5: 57.9% on databricks’ eval

legal agent benchmark mythos 5: 16.91% all-pass; 92.0% mean criterion-pass

healthbench mythos 5: 62.7%

healthbench professional mythos 5: 66.0%

multilingual gmmlu / milu / include 93.2%; 92.9%; 90.5%

biomysterybench 83.9% human-solvable; 46.1% human-difficult

organic chemistry mythos 5: 90.1%

labbench2 patent questions mythos 5: 79.8%

Also why do they just keep doxing people left and right?

Scott Alexander as the most memorable, and then the backlash after they post the backrooms movie creator's house on twitter recently

Just very shortsighted behavior

First ever argument being "People do not realise how much of a toll it takes on you if you actually care about the environment"

GUYS

PLEASE

The impact of ai on the enviroment is one of the dumbest psyops in history, how can you claim to know start with that after claiming you know the technology and what it is doing?

There are hundreds of reasons to hate ai but this is just NOT it

Is This Prime 2 months ago

quite a fun game to bot! got this:

Game over You ran out of time!

You correctly sorted 8389560 numbers.