HN user

wizeyone

10 karma

Founder, Wizey (https://wizey.one) - AI for clinical lab interpretation.

Posts1
Comments6
View on HN

2 things. The headline math 0.95^20 = 0.358 assumes independent errors. "The body argues the opposite - every subsequent action operates on flawed foundations."

Real long chain failure is worse than the math predicts, not equal to it. The headline undersells the problem the article actually describes.

Also DTCM's eval is narrative-QA across 250 stories, reading comprehension over accumulated context, not an agent tool use.

The production failure modes it discusses (wrong tool selection, brittle API contracts, etc) don't obviously map to that benchmark. The 96% number is encouraging but not directly translatable

Author here - we're the team behind Wizey, one of the two AIs in the comparison. A few things up front:

* Methodology was fixed before the runs.

* All outputs are quoted verbatim, including Case 2 (MGUS) where ChatGPT beat us cleanly.

* Panels are reconstructed from published case reports (Blood, Annals of Family Medicine, and others), so anyone can reproduce the experiment on Claude, Gemini, or Grok.

Full verbatim outputs for all five cases: https://wizey.one/blog/2026/04/17/wizey-vs-chatgpt-raw-exper...

Happy to answer anything on methodology or individual cases.

Arms race framing misses it. Insurers have used algorithmic denial scoring for years (ProPublica/Cigna-EviCore, StatNews/UnitedHealth-NaviHealth). Denial works because appealing is expensive for patients and near-free for insurers. Claimable inverts that cost. End state isn't that insurers pay more. It's more like "insurers deny less aggressively up front."

"AI-generated and approved by engineers" is doing a lot of work there. If accepting a 4-character Gemini autocomplete counts, Copilot users hit >90% last year. The useful metric is % of functions where >50% of the body was AI-written before human edits. just my 2 cents