HN user

convexly

151 karma

www.convexly.app

Posts4
Comments72
View on HN

That claim is testable. The 2026 microstructure work on the Kalshi tape (72M trades, Becker) documented a systematic +1.12% excess return for liquidity providing makers and a symmetric -1.12% for takers, plus a longshot bias where 1-cent yes shares pay -41% EV and 1-cent no shares pay +23% EV. That edge is in the market microstructure layer. A patient maker who never trades the frontpage market makes money on the marginal bidder asymmetry alone.

Akey, Gregoire, Harvie, Martineau (2026) on Polymarket find that the top 1% of wallets capture about 84% of all profits, and that the largest whale wallets are NOT the most sophisticated. Large capital systematically bleeds expected value to small order traders. Reichenbach and Walther (2025) document within-Polymarket skill persistence at the 124M-trade scale, so skill differentials are measurable across users distinct from the insider trading question.

I ran this analysis using Convexly, a calibration-tracking tool I'm working on. The Brier scoring math in the post is the same math that runs in the product. Disclosure out of the way.

What I did: pulled trade data for the top 100 Polymarket profit wallets from the public API, resolved 3,651 unique positions, and computed per-wallet Brier scores. Then tested the correlation between calibration and profit using Spearman rank correlation (not Pearson, because the profit distribution is fat-tailed, Hill alpha ~1.6).

The headline finding: Spearman r = +0.608 between Brier score and realized profit. Worse calibration predicts bigger profits. The correlation gets stronger when you drop the top 10 by profit (+0.72), so it isn't outlier-driven. The worst-calibrated whales earn 4.66x the median profit of the better-calibrated whales.

The leaderboard is a concentration ranking. Median single-event concentration: 69.8%. Twenty of the top 100 made their biggest money on the 2024 election.

After separating the four wallets Chainalysis publicly attributed to "Theo" (the French trader who commissioned private YouGov polls), 8 wallets remain in a narrow cluster on the popular-vote markets in a 3-week window around election day. Chainalysis did not link BetTom42 or alexmulti to Theo.

Full 47-column CSV is downloadable from the post. Happy to answer methodology questions.

Both really good points. The research does suggest that the core skill does transfer. The quiz can help with long horizon predictions. The mechanism itself seems to be the actual awareness of overconfidence rather than just domain-specific knowledge. With that being said, the gap between the quiz and real-world application is real, and tracking both over time is part of why I built the decision logging side. For your question about teams, that's a built-in feature already! Submissions are "sealed" so you submit before seeing others. The team feature also has a believability-weighted aggregation based on each submitter's track record, and I also built an IC mode for investment committees. The problem you describe about one calibrated person in a room with two uncalibrated ones is what the sealed model prevents. Everyone makes draws their own conclusion, then they compare!

That's all true. I'm a solo founder and have been using Claude heavily to build this. It definitely shows in many places, and I'll make sure to clean those up. I did not expect to get this many visits from a show HN (almost at 1600 quiz takers from the last few hours alone). The core math is sound, but I agree the presentation needs more care. Appreciate the honest feedback!

Your point on scoring is correct, if you're 100% confident and right on everything you would score a perfect 0. The calibration insight is in how you handle the questions where you don't know the answer. Say you're highly knowledgeable and 95% confident on everything, but get 2 wrong scores compared to someone that says they are 70% confident on those same two questions. That would indicate that you are overconfident compared to the other person!

Great question. Calibration specifically is about whether your confidence in an answer matches your accuracy, not whether you know the answer. Someone who knows a lot but is always 90% confident would score poorly even if they're wrong 20% of the time, as an example.

In terms of research, Tetlock's Expert Political Judgement and Superforecasting were the foundation. He did a 20 year study that showed domain experts were barely better than chance at long-range predictions. The Brier score was the standard metric for that research.