AI debate for researchers and analysts: verified multi-perspective analysis
The failure mode of research assistants — human or AI — is agreement. One model summarizing literature will happily confirm your hypothesis, cite plausible-sounding numbers and never mention the study that contradicts you.
A debate flips the incentive. Models from different vendors are pushed to find what the others missed: the confound, the opposite result, the number that doesn't replicate. A verifier then checks each factual claim on the web and marks it VERIFIED, DISPUTED or FALSE, with sources you can follow.
Stress-test a hypothesis, not just summarize around it
State the claim — 'remote work reduces junior developer onboarding speed' — and let Debate mode run. Agents argue for and against, must respond to each other's strongest evidence, and their agreement is tracked numerically per round, so 'the models genuinely converged' is distinguishable from 'one wrote longer paragraphs'.
The Socratic mode works the other way: when your question is still vague, one agent interrogates it with progressively deeper questions until the actual research question emerges.
Fact-checking is a separate pass, not a promise
After the debate, a verifier extracts every self-contained factual assertion from the transcript and checks each one independently with its own web search. The output is a claim-by-claim table — VERIFIED, DISPUTED, FALSE or UNVERIFIABLE — with the sources next to each verdict.
It also flags position shifts: if an agent quietly changed its stance between rounds, that gets called out with before/after quotes.
- Hypothesis testing → Debate with fact-check on
- Vague question → Socratic mode to sharpen it first
- Literature angles → Expert panel (independent positions, merged verdict)
- Survey of options → Brainstorm, then Idea mode to deepen the winner
- Final write-up → Synthesis mode drafts one document with marked edits
Multiple vendors as a bias control
Using models from rival companies is a crude but real bias control: training sets, alignment choices and refusal behaviours differ, so a claim that survives GPT, Claude, Gemini and Grok simultaneously is meaningfully harder to attribute to one vendor's quirks. Council mode (Pro) strengthens this with anonymous peer ranking — models grade answers without knowing whose they are.
A transcript you can cite
Every debate produces a full transcript: who argued what, in which round, what was verified and what was disputed. Export it to PDF or DOCX and attach it to the appendix — 'we asked an AI' becomes a documented, checkable procedure.
FAQ
Do the models have access to academic databases?
Agents search the open web, which covers preprints, abstracts and open-access work. For paywalled corpora, paste key excerpts into the question — agents will debate the material you provide.
Can I trust the VERIFIED label?
Treat it as triage, not gospel: each claim is checked against live web sources that are linked next to the verdict, so you can audit any claim in one click. DISPUTED and UNVERIFIABLE labels are honest about the limits.
Which languages can I research in?
The interface ships in 15 languages and agents reply in the language of your debate — sources found on the web may of course be in any language.
Free tier · bring your own API keys · 15 languages