HN user

felix089

242 karma

https://x.com/felix94123

Posts9
Comments99
View on HN

Hey HN! About 2 months ago I launched AI Roundtable here, where 200+ models answer and debate your question (https://news.ycombinator.com/item?id=47507666).

It's now collected 29,605 public sessions and 336,039 individual model responses, and we just published all the stats for everyone to see, updating them daily.

A few highlights:

- In multi-round debates, Claude Opus 4.7 convinced other models to flip their vote almost 3K times, the most of any model. Gemini 3.1 Pro came in second at 2.1K

- Most used model is Gemini 3.1 Pro at 25K sessions, with GPT-5.4 second at 21K.

- Grok 4.1 Fast held its position 88.7% of the time, the highest conviction rate of all models. Probably not surprising.

It's been quite amazing to see all the questions and feedback since launch. Initially the only mode was structured answers (vote yes/no or pick from custom options).

Based on feedback we've added an open questions mode where models answer freely and a roundtable chat where you can join the debate and follow up with individual models to challenge their reasoning.

If you want to give the roundtable a try, it's free to use until community credits run out. All models routed via my startup Opper. Happy to dig into specifics or make more data available if interesting.

Hey HN! About 2 months ago I launched AI Roundtable, where 200+ models answer and debate your question (https://news.ycombinator.com/item?id=47507666).

It's now collected 29,502 public sessions and 334,589 individual model responses, and we just published all the stats for everyone to see, updating them daily.

A few highlights:

- In multi-round debates, Claude Opus 4.7 convinced other models to flip their vote almost 3K times, the most of any model. Gemini 3.1 Pro came in second at 2.1K - Most used model is Gemini 3.1 Pro at 25K sessions, with GPT-5.4 second at 21K. - Grok 4.1 Fast held its position 88.7% of the time, the highest conviction rate of all models. Probably not surprising.

It's been quite amazing to see all the questions and feedback since launch. Initially the only mode was structured answers (vote yes/no or pick from custom options).

Based on feedback we've added an open questions mode where models answer freely and a roundtable chat where you can join the debate and follow up with individual models to challenge their reasoning.

If you want to give the roundtable a try, it's free to use until community credits run out. All models routed via my startup Opper. Happy to dig into specifics or make more data available if interesting.

Okay since the launch we got about 5k questions asked to the roundtable, really cool stuff! We had much higher usage than expected and had to scale up to keep things running. Thanks for all the feedback, shipped a bunch of updates during the day. Now the history tab has a much better sorting logic, added upvotes, and more filters. You can create final summaries in a couple of voices, which is quite funny I think. There's a couple more things coming shortly, like open questions mode and potentially joining as a participant in the roundtable. Any other feedback just let me know. Thanks!

You can basically already do that, all you need is to create your own API key and put it in navbar/API key. Then all your sessions are unlisted so unless someone has the link nobody will should be able to find it. You can still share them with others if you like. Like unlisted yt videos.

yea good points, in general the models don't change their mind that much from what I have seen with the current sample size, but worth checking in more detail. The summarizer is just tasked with objective summarization from facts presented, it doesn't have an opinion, so changing model should not really affect anything.

Thank you, and fun use case. Yea this is just v1 I have an open question version, but the UI is not as sleek. But what you can do is download the transcript, put it into claude and generate a chart. Which when I think about it would also be a nice UI idea for the page, custom charts based on the model output data. Will report back on this! And RE costs, most questions are very cheap so I created a credit pool anyone can use. if people keep having fun, I'll keep on filling it up, and it looks good so far

thanks happy to hear. Yes for debate mode the max number of models is actually only 6. More than that didn't really add anything in my preliminary test. Only for direct comparison in the poll mode you can choose up to 50, then it's kind of nice to see their single responses side by side.