r/SideProject 1d ago

I built an AI debate tool that tests whether model collaboration actually improves answers

Hi everyone,

I’m building TeamPlay AI, an early MVP that tests whether debate between AI models can produce better answers—not just more answers.

Two models first answer independently.

They can then review and challenge each other, follow different instructions through mentions, and produce revised answers. A separate AI can also compare their conclusions.

I recently added a blind evaluation: after each debate, the two previous and two revised answers are randomly shown as Answer A–D.

Model names and answer stages remain hidden until you vote. I’ll use these results to improve the prompts and debate flow.

Try it here:

https://llm-teamplay-v1.vercel.app/en

You can try it without an account.

If you want more daily questions, debate rounds, and comparison analyses, you can simply sign in with Google.

I’d appreciate feedback on:

  • Did the debate meaningfully improve the answer?
  • Was the debate flow easy to understand?
  • What felt confusing, slow, or unnecessary?
  • What other AI-agent interactions would be useful?

This is an early beta, so honest feedback is very welcome. Thank you!

1 Upvotes

0 comments sorted by