Maybe the next summer hit, peak AI
HN user
felix089
https://x.com/felix94123
Thanks yea same, it is both very convincing but at the same time also open to changing its positing when presented with a good argument.
Hey HN! About 2 months ago I launched AI Roundtable here, where 200+ models answer and debate your question (https://news.ycombinator.com/item?id=47507666).
It's now collected 29,605 public sessions and 336,039 individual model responses, and we just published all the stats for everyone to see, updating them daily.
A few highlights:
- In multi-round debates, Claude Opus 4.7 convinced other models to flip their vote almost 3K times, the most of any model. Gemini 3.1 Pro came in second at 2.1K
- Most used model is Gemini 3.1 Pro at 25K sessions, with GPT-5.4 second at 21K.
- Grok 4.1 Fast held its position 88.7% of the time, the highest conviction rate of all models. Probably not surprising.
It's been quite amazing to see all the questions and feedback since launch. Initially the only mode was structured answers (vote yes/no or pick from custom options).
Based on feedback we've added an open questions mode where models answer freely and a roundtable chat where you can join the debate and follow up with individual models to challenge their reasoning.
If you want to give the roundtable a try, it's free to use until community credits run out. All models routed via my startup Opper. Happy to dig into specifics or make more data available if interesting.
Hey HN! About 2 months ago I launched AI Roundtable, where 200+ models answer and debate your question (https://news.ycombinator.com/item?id=47507666).
It's now collected 29,502 public sessions and 334,589 individual model responses, and we just published all the stats for everyone to see, updating them daily.
A few highlights:
- In multi-round debates, Claude Opus 4.7 convinced other models to flip their vote almost 3K times, the most of any model. Gemini 3.1 Pro came in second at 2.1K - Most used model is Gemini 3.1 Pro at 25K sessions, with GPT-5.4 second at 21K. - Grok 4.1 Fast held its position 88.7% of the time, the highest conviction rate of all models. Probably not surprising.
It's been quite amazing to see all the questions and feedback since launch. Initially the only mode was structured answers (vote yes/no or pick from custom options).
Based on feedback we've added an open questions mode where models answer freely and a roundtable chat where you can join the debate and follow up with individual models to challenge their reasoning.
If you want to give the roundtable a try, it's free to use until community credits run out. All models routed via my startup Opper. Happy to dig into specifics or make more data available if interesting.
Hey just fyi the open question feature is now live. Also gave the UI a facelift. Any feedback welcome! Also got a custom domain for easy access: https://askroundtable.ai
It's now live, give it a spin!
Nice! Opus in general is the best debater so far, most models cited Opus for changing their opinion, by a considerable margin.
Cool question! just a quick headsup, they don't have access to tools so what you are seeing are answers based on their training data. They might not know about the latest model version. That said, sonnet is def a great choice.
It's so funny to see the smaller / first gen models make the wrong choices despite overwhelming evidence, almost adorable. I ran the same test with one model from each GPT generation, all but 3.5 Turbo could be convinced. https://opper.ai/ai-roundtable/questions/i-want-to-wash-my-c...
haha good to hear, then the latest update on the roundtable history list seems to work well and the good ones are on top
Thanks, yes this is coming shortly!
Two models changed their minds but from opposite sides so the score stayed the same, that's the first time I've seen this.
Okay since the launch we got about 5k questions asked to the roundtable, really cool stuff! We had much higher usage than expected and had to scale up to keep things running. Thanks for all the feedback, shipped a bunch of updates during the day. Now the history tab has a much better sorting logic, added upvotes, and more filters. You can create final summaries in a couple of voices, which is quite funny I think. There's a couple more things coming shortly, like open questions mode and potentially joining as a participant in the roundtable. Any other feedback just let me know. Thanks!
You can basically already do that, all you need is to create your own API key and put it in navbar/API key. Then all your sessions are unlisted so unless someone has the link nobody will should be able to find it. You can still share them with others if you like. Like unlisted yt videos.
Okay it's done, all fixed!
Yes! Amazing you spotted this, I'm about to push an update, will be live in 1h max.
Glad you like it!
Thanks!
Yea Opus 4.6 is the one that changes opinions the most from what I've seen. Also the maybes or the are you 100% certain framings trigger most models to default to maybe / no. https://opper.ai/ai-roundtable/questions/can-you-be-100-cert... - Or as Shane puts it, Nobody's saying he IS a lizard. They're saying the universe doesn't hand out 100% certificates.
The debate round is actually restricted to only 6 models otherwise I'd get out of hand both quality and financially. And changing position is just one feature of the debate. Seeing arguments from multiple sides is also quite nice, give it a spin!
Yes, much requested feature it will be released shortly!
yea good points, in general the models don't change their mind that much from what I have seen with the current sample size, but worth checking in more detail. The summarizer is just tasked with objective summarization from facts presented, it doesn't have an opinion, so changing model should not really affect anything.
Thanks! :)
Happy to hear! Yes very true I have a version built for open questions already but wasn't too happy with the UI yet. It's not as straight forward as comparing based on answer options. But I'll release a first version of it shortly and let you know
Thank you, and fun use case. Yea this is just v1 I have an open question version, but the UI is not as sleek. But what you can do is download the transcript, put it into claude and generate a chart. Which when I think about it would also be a nice UI idea for the page, custom charts based on the model output data. Will report back on this! And RE costs, most questions are very cheap so I created a credit pool anyone can use. if people keep having fun, I'll keep on filling it up, and it looks good so far
Yea Gemini is the only model that chose based on the correct reason, the other ones got kind of lucky
Thanks, yes bias is one of the most interesting ones for sure
This app cracked the GEO code
Thanks! Yea I think the best ones are when science is actually quite clear but politics get in the way so you see their bias
thanks happy to hear. Yes for debate mode the max number of models is actually only 6. More than that didn't really add anything in my preliminary test. Only for direct comparison in the poll mode you can choose up to 50, then it's kind of nice to see their single responses side by side.