Love the p_step^N framing — maps cleanly to agentic chains where errors compound. Worth naming a second reliability failure that sits one layer below, in single-turn: confident-absent.
Ran "what's the best X for Y?" across 6 LLMs (ChatGPT, Perplexity, Gemini, Claude, DeepSeek, Mistral) for ~200 B2B SaaS tools across 34 categories. In 60%+ of categories, models converge on the same "default three" and everything else is effectively invisible. Not wrong — just erased. Single-turn, so it never shows up in p_step^N.
A verification layer catches "false." But there's no layer catching "the space of correct answers was silently pruned." Curious if your framework could be extended from correctness per step to coverage per response.