How the AI Evaluation Works

In a debate with an AI adjudicator, the round is evaluated from its transcript after the speeches are over. This page explains what comes out of that evaluation — the numbers, what each of them means, and how the written feedback is put together — and why it does not arrive the moment the debate ends.

Table of contents

  1. Two Different Numbers
  2. The Speaker Score (50–100)
    1. What the score is not based on
    2. The Ordinary Intelligent Voter
  3. The Four Dimensions
    1. Evidence
    2. Clarity
    3. Strategy
    4. Engagement
  4. What 1 to 5 Means
    1. Your level comes from your speaker score
    2. What “at this level” means
  5. The Overlapping Bands
    1. How to read your report
    2. The full ladders
  6. The Written Feedback
    1. How it is structured
    2. What the feedback draws on — and what it will never do
  7. When It Runs, and How You Know
    1. The sequence
    2. The evaluation does not know who you are
    3. While it is running
    4. When it finishes
    5. Chair or wing changes what lands
  8. Limitations Worth Knowing

Two Different Numbers

The evaluation gives you two kinds of number, and they answer two different questions.

  Speaker score Dimension scores
Range 50–100 1–5, four of them
Question it answers How persuasive was this speech? Which parts of my debating are ahead, and which should I work on next?
Compared against The global standard for competitive debating Your own level, taken from your speaker score
Purpose The result of the round Developmental coaching

The speaker score is canonical. It is the WUDC measure of persuasiveness and the only number that feeds team totals and the team ranking. The four dimension scores never change it — they are a decomposition of it.

That has one consequence worth stating up front: the four dimension scores are only meaningful next to your speaker score. A 3/3/3/3 at 84 and a 3/3/3/3 at 75 describe very different debaters.

The Speaker Score (50–100)

Every debater gets one score from 50 to 100, on the WUDC speaker scale. The scale is anchored to the global judging pool: the average tournament speech is a 75, most marks land in the 70s, high-60s and low-80s, high-80s and low-60s are uncommon, and 50s and 90s are rare.

Score What it describes
95–100 Plausibly one of the best speeches ever given; near-impossible to respond to; flawless and compelling.
92–94 Incredible speech, among the best at the competition; engages core issues, exceptionally well made; no significant flaws.
89–91 Brilliant; engages main issues, very well-explained and illustrated, demands sophisticated responses; only very minor problems.
86–88 Engages core issues, highly compelling; no logical gaps; sophisticated responses required; only minor flaws.
83–85 Addresses core issues; strong explanations demanding strong responses; may occasionally fail to fully answer a well-made argument; limited flaws.
79–82 Relevant and addresses core issues; well made without obvious logical gaps, all well explained; may be vulnerable to good responses.
76–78 Almost exclusively relevant, addresses most core issues; occasional deficits in explanation / simplistic / peripheral.
73–75 Almost exclusively relevant but may miss a core issue; logical but simplistic and vulnerable to competent responses.
70–72 Frequently relevant; some explanation but regular significant logical gaps.
67–69 Generally relevant; almost all arguments have significant logical gaps.
64–66 Some relevant arguments; significant logical gaps.
61–63 Some relevant claims, most formulated as arguments; occasional explanations with big gaps.
58–60 Claims occasionally relevant; not really formulated as arguments.
55–57 One or two marginally relevant claims; just comments, not arguments.
50–55 Not relevant; bare claims, confusing and confused.

What the score is not based on

The evaluation works from a written record of the round. It has no signal for delivery, voice, fluency, accent, language proficiency, or who you are — so none of those can move your score, which is also a hard rule of the WUDC manual. Where the band descriptors above talk about being “clear to follow”, that means the clarity of the argument’s logic, not how you sounded saying it.

The Ordinary Intelligent Voter

Everything is judged through the Ordinary Intelligent Voter (OIV) — a reasonably informed, intelligent adult with broad general knowledge, no specialist expertise in the motion’s subject, and no prior allegiance to either side. The OIV is persuaded only by reasons actually given in the round. They do not fill gaps for you, do not supply a missing step, and do not import expert knowledge to rescue your argument.

In practice this means an assertion persuades weakly even when it happens to be true, and jargon, an unexplained example, or a bare name-drop carries little weight because the OIV cannot evaluate it.

The Four Dimensions

Alongside the speaker score you get four sub-scores: Evidence, Clarity, Strategy, Engagement. Each is a whole number from 1 to 5 (the home-page card shows them with one decimal, so a 3 reads as 3.0).

Evidence

Is the case grounded in evidence, rather than bare assertion?

In the competitive field this standard was built on (what that means), almost everyone’s arguments are relevant to the burden and basically sound — those things barely tell debaters apart. What genuinely varies is grounding: how much of your case is backed by a concrete example, data, or cited authority rather than left as an assertion the OIV is asked to take on trust.

So this axis is grounding-led. Relevance and logical soundness act as an assumed floor: a real logical gap, material that is off-burden, or an implausible reason pulls your placement down, but their absence does not push it up, because it is expected.

Clarity

Can the Ordinary Intelligent Voter follow and access your reasoning?

This is the legibility of the argument itself: how deeply your reasons are explained (well explained, only partially, or merely asserted), whether jargon, examples, and references are unpacked so a layperson can evaluate them rather than take them on trust, and whether the chain from reason to claim can be followed without the OIV having to supply a missing step.

It is not spoken delivery. Delivery is not observable from the record and is deliberately excluded.

Be aware: across the competitive field this standard was built on (what that means), most speakers are similarly clear, so Clarity reads 3 for the large majority. It flags only a genuine relative strength or weakness in legibility, and it is the least differentiating of the four axes.

Strategy

Did you do your seat’s job, weigh the debate, stay consistent, and spend your time where the round is decided?

Strategy is seat-aware — the bar is different for every position:

  • PM — an adequate definition that frames the round.
  • MG / MO — a genuine extension with real new material, not a rehash of the opening.
  • GW / OW — whip discipline: crystallise and weigh, without smuggling in new arguments.
  • Deputies — rebuild and develop.

On top of the seat’s job it looks at explicit comparative weighing (why this matters more, who is most affected, why it outweighs the other side), the absence of contradictions within your speech, within your team, or against your own opening, and prioritisation — whether your effort went into the central material or the peripheral.

Engagement

How directly and effectively did you interact with the rest of the room?

Four things drive it:

  • Rebuttal directness and target — how many of your rebuttals engaged directly rather than partially or missing, and whether they went after the opponents’ central case or only its weaker, peripheral points.
  • Clash outcomes — whether your contributions to the round’s major clashes stood or were rebutted. A point that stood while an opponent spoke later and could have answered it is a genuine win; one that stood only because nobody followed you is not, and is weighted accordingly.
  • Resilience under fire — of your arguments that drew a rebuttal, how many you defended rather than dropped.
  • Interactivity — taking POIs at sensible moments, and engaging the live exchanges rather than talking past them.

What 1 to 5 Means

The dimension scores are read relative to your own level, never as absolute quality and never as a ranking against the other seven debaters in the room.

Score Meaning
5 A standout strength, markedly ahead of your level on this axis. Uncommon.
4 A relative strength — ahead of your level.
3 At your level. You performed on this axis exactly as expected for a speaker of your persuasiveness. This is the modal score.
2 A relative weakness — your priority development area.
1 Markedly behind your level on this axis. Uncommon.

A 3 is not a mediocre mark and a 2 is not a harsh one. Three means “as expected for you”; two means “this is the one to work on next”. Most strong speakers have at least one axis below 3 — at a fixed level, if one axis runs ahead, another usually runs behind.

Your level comes from your speaker score

Real competitive speaker scores are tightly compressed, so the level grid is narrow:

Level Speaker score Read
L1 74 and below entry competitive
L2 75–77 developing
L3 78–80 modal — the median speaker
L4 81–83 strong
L5 84 and above excellent / elite

The bar slides up with the level. More is expected of an 82 than of a 76, so the same 3 represents more accomplishment higher up — and the very performance that earns a 4 at L2 is typically only a 3 at L4.

What “at this level” means

Several statements on this page — that relevance and soundness are near-universal, that most speakers read 3 on Clarity — are claims about a specific population, and it is worth naming it.

The standard was built and calibrated against preliminary rounds of international competitive BP: 62 rounds across 13 tournaments, including Worlds and Euros as well as open tournaments and IVs. Across the 1,232 speaker scores in that corpus the mean is 79 with a standard deviation of about 3, so in practice almost everyone lands between 72 and 85 — which is why the five levels above are only three points wide.

That population is the “level” being referred to. Two consequences follow:

  • In a round well below that field, the claims stop holding. Relevance and logical soundness are not near-universal among newer debaters, and clarity genuinely varies — so those axes would separate speakers on more than the page describes.
  • The speaker score itself is unaffected. The 50–100 scale is absolute and anchored to the global judging pool, where 75 is the average tournament speech. Only the dimension scores are read against a level, and only their calibration depends on this corpus.

The Overlapping Bands

Underneath each dimension is a single continuous skill ladder of 17 rungs, G1 to G17, running from “essentially nothing there” to a theoretical ceiling. Your 1–5 score is a window onto that ladder, and which window you get depends on your level.

Each level’s window is 5 rungs wide. Adjacent windows step up by 3 rungs and overlap by 2 — so the top of one level’s window is the bottom of the next level’s:

Dimension score L1 L2 L3 L4 L5
1 G1 G4 G7 G10 G13
2 G2 G5 G8 G11 G14
3 (at level) G3 G6 G9 G12 G15
4 G4 G7 G10 G13 G16
5 G5 G8 G11 G14 G17

Read across any row and the overlap is the whole point:

A 4 at your level is the same skill as a 1 at the next level. A 5 is the same skill as a 2.

This is what lets the scores show you the road ahead. Maxing out a dimension at your level is genuinely maxed out for now — and the moment your speaker score carries you up a level, that same skill re-bases to a 2 and there is further to climb.

It also explains something that would otherwise look like a mistake: your dimension scores can go down in a round where you debated better. If a stronger performance lifts your speaker score across a level boundary, you are being read against a higher bar, and the same skill scores lower. That is the system working, not against you.

How to read your report

  1. Find your level from your speaker score.
  2. Read your four 1–5 scores — 3 is at-level, above is a relative strength, below is your priority.
  3. To see where the path leads, look at the next level’s column: your 5 is its 2.

The full ladders

Each ladder below is the absolute standard, independent of level. Use the table above to find which rungs your own level’s window covers.

Evidence — full ladder
Rung Evidence (grounding) at this rung
G1 Bare assertion throughout, compounded by real defects — major gaps, off-burden or implausible material. Nothing an OIV could accept on the evidence.
G2 Almost entirely unevidenced assertion; any “support” is restating the claim; may carry genuine gaps.
G3 (L1 at-level) Mostly assertion with the occasional concrete example; evidence is thin and uneven; arguments are relevant but taken largely on trust.
G4 Examples appear on some lines; the case is no longer pure assertion, though key claims still rest unevidenced.
G5 Most lines carry at least an illustrative example; little is left as bare assertion.
G6 (L2 at-level) The case is consistently evidenced — examples on the bulk of arguments, with data or authority on at least one key line; assertion is the exception.
G7 As G6, plus data or authority reaches the load-bearing arguments, not just the easy ones.
G8 Strong, specific grounding throughout — examples and data deployed where they matter most.
G9 (L3 at-level — modal) Every argument of consequence is concretely grounded; the decisive lines carry data or authority; nothing important rests on assertion.
G10 As G9, plus the evidence is precise and well-chosen — each piece does real work, none decorative.
G11 Comprehensively and authoritatively grounded — the case’s evidentiary base is hard to attack on the merits.
G12 (L4 at-level) The most decisive claims are backed by the strongest available evidence, and even secondary lines are concretely supported.
G13 As G12 across the whole case, including lines most speakers leave to assertion.
G14 Every argument is authoritatively grounded; the evidentiary base is effectively unimpeachable.
G15 (L5 at-level) The claims the burden turns on are grounded beyond reasonable dispute; nothing material is left unsupported.
G16 Near the ceiling: an evidentiary base an expert panel would struggle to fault.
G17 The theoretical ceiling: flawless, complete, decisively authoritative grounding.
Clarity — full ladder
Rung Clarity at this rung
G1 Reasoning is largely opaque: reasons asserted without explanation; pervasive unexplained jargon, examples, and name-drops. The OIV cannot follow what is meant or why it follows.
G2 Mostly bare claims; explanations rare and partial; frequent unexplained terms or examples the OIV cannot evaluate.
G3 (L1 at-level) Some reasons partially explained, others asserted; occasional unexplained jargon or name-drops; a layperson follows parts but must fill gaps elsewhere.
G4 More reasons reach partial or full explanation; unexplained terms now occasional rather than frequent.
G5 Most reasoning is at least partially explained and accessible; few terms left unpacked.
G6 (L2 at-level) The bulk of reasons are well explained in terms a layperson can follow; jargon, examples, and references are unpacked; the OIV can trace most arguments without filling gaps.
G7 As G6, plus the load-bearing reasons are the most clearly explained; almost no unexplained material.
G8 Reasoning is consistently well explained and self-contained; the OIV is never asked to supply a step.
G9 (L3 at-level — modal) Every argument of consequence is well explained and fully accessible — concepts defined, examples spelled out, references made meaningful; the chain from reason to claim is traceable throughout.
G10 As G9, plus the explanations are economical and ordered so the reasoning is easy to hold in mind, not merely complete.
G11 Transparently clear end to end — a layperson follows every step on first pass.
G12 (L4 at-level) Explanations are precise, vivid, and self-contained; even complex mechanisms are made intuitive; nothing requires specialist knowledge or a second pass.
G13 As G12 across all material, including subtle or technical lines most speakers leave under-explained.
G14 Reasoning rendered so clearly that its force is immediately apparent to any layperson.
G15 (L5 at-level) Every idea, however complex, is made fully intuitive and self-contained; the reasoning is maximally legible and its weight unmistakable.
G16 Near the ceiling: clarity an expert communicator would admire.
G17 The theoretical ceiling: flawless, effortless legibility of even the hardest material.
Strategy — full ladder
Rung Strategy at this rung
G1 Fails the seat’s core job (no real definition, no extension, or new arguments dumped in the whip); no weighing; self-contradiction or knifing the team; effort scattered on peripheral points.
G2 Seat job only partly met; little weighing; some inconsistency; prioritisation weak.
G3 (L1 at-level) The seat’s basic duty is met but unevenly; weighing attempted occasionally; no major contradictions; effort split between central and peripheral material.
G4 Seat job met more cleanly; the odd explicit weighing; consistent; effort leaning toward central material.
G5 Seat duty clearly fulfilled; weighing present on key lines; no contradictions; central material prioritised.
G6 (L2 at-level) Fulfils the seat’s role well — clear definition, genuine extension, or disciplined whip; weighs the major clashes explicitly; internally consistent; effort concentrated where the burden turns.
G7 As G6, plus weighing begins to shape which clashes matter; prioritisation sharp.
G8 Role executed strongly, comparative weighing consistent, fully consistent, effort almost entirely on decisive material.
G9 (L3 at-level — modal) Discharges the seat’s role to its full purpose — a definition that frames the round, an extension that genuinely advances it, or a whip that crystallises the whole debate; weighs the key clashes and says why they decide the round; every minute spent on what matters.
G10 As G9, plus weighing is comparative and decisive — sets the metric the round should be judged on; flawless role discipline.
G11 Commands the strategic shape of the round: frames the comparison, prioritises ruthlessly, weighs every key clash, zero inconsistency.
G12 (L4 at-level) Sets the terms on which the round is judged; weighing that pre-empts the other side’s metric; perfect consistency and allocation.
G13 As G12 sustained across the whole speech, including managing the team’s narrative across the half.
G14 Strategic choices optimal throughout — nothing wasted, every comparison won on the framing.
G15 (L5 at-level) Dictates the round’s strategic frame and weighing so completely that the comparison is settled on the team’s terms; role executed to its theoretical purpose.
G16 Near the ceiling: strategic generalship an expert adjudicator would single out.
G17 The theoretical ceiling: flawless, complete strategic command.
Engagement — full ladder
Rung Engagement at this rung
G1 Essentially no responsive material: rebuttals absent or missed; attacked arguments dropped; POIs declined; clash contributions unengaged or immediately rebutted. The speaker talks past the debate.
G2 Occasional rebuttal, but mostly missed or partial and aimed at peripheral points; most attacked material dropped; little interaction.
G3 (L1 at-level) Some direct rebuttals, but inconsistent and often answering the weaker opposing points; a few attacked arguments defended, others dropped; may take a POI. Engagement is present but reactive and patchy.
G4 More consistent direct rebuttal of supporting material; roughly half of attacked lines defended; takes a POI; some clash contributions stand by speaking position rather than by being unanswerable.
G5 Rebuttals mostly direct and on-target; most attacked arguments defended; engages the live exchanges rather than talking past.
G6 (L2 at-level) Consistently direct rebuttal aimed at the opponents’ central case; the bulk of attacked material defended; POIs taken at sensible moments; several contributions stand against opponents who had the chance to answer.
G7 As G6, and begins to pick the strongest opposing argument to answer rather than the easiest; resilience near-complete.
G8 Rebuttal is direct, central, and prioritised; almost all attacked material defended; clash contributions reliably stand against opponents who could answer.
G9 (L3 at-level — modal) Engages the heart of the opposing case directly and early; defends every attacked line of consequence; takes and exploits POIs; contributions to the major clashes consistently stand against later opponents.
G10 As G9, and turns rebuttal into offence — rebuts and advances own case in one move; no attacked material left unaddressed.
G11 Dominates the major clashes: opposing central arguments are directly answered and left rebutted; resilience total; POIs used to reframe.
G12 (L4 at-level) Sets the terms of the clash — identifies and dismantles the opponents’ best material, defends flawlessly under sustained fire, uses every interaction to tighten the comparison.
G13 As G12 across essentially every clash in the room, including those off their side’s natural turf.
G14 Engagement is comprehensive and pre-emptive — answers arguments before they fully land, leaving opponents nothing standing.
G15 (L5 at-level) Effectively wins every exchange entered: the opposing central case is systematically answered and kept down, own material survives all attack, and POIs and clash framing control the round’s comparison.
G16 Near the ceiling: engagement opponents cannot productively respond to; every clash resolved by argument, not by speaking position.
G17 The theoretical ceiling: flawless, total engagement; nothing the speaker enters is left contestable.

The Written Feedback

Every debater gets their own written feedback — private to them, and written as coaching rather than judging. It is addressed to you in the second person, runs to roughly 500 words, and is held to the job your seat was meant to do.

Open it from the AI evaluation link on the debate’s card on your home page, or from the AI adjudicator’s row in the expanded card (see The Ballot).

How it is structured

The note always arrives in the same three parts:

Summary
The prose body. It moves through what worked — praise anchored to specific strong moments, named by substance — then what held you back, then your highest-leverage change: the one or two most specific, actionable things to do differently next time. Not “be more persuasive”, but something like “before moving on, spend one more sentence explaining why your mechanism produces the harm you claim.”
Skills
One line per dimension, giving the score and a short note explaining it:
  • Evidence (3/5): your stereotype argument landed because you gave a concrete case, but the economic line rested on assertion — attach one real example to it next time.
  • Clarity (3/5): …
  • Strategy (4/5): …
  • Engagement (2/5): …
Points
The last line: your speaker score with a one-to-two-sentence rationale tying it to a band descriptor.

Points: 79 — Relevant throughout and well explained, with the central mechanism defended after attack; vulnerable on the comparative, which is what keeps it out of the low 80s.

What the feedback draws on — and what it will never do

The prose is anchored to the structured record of the round, which decides what was strong or weak. The transcript is used only to pull accurate short quotes and phrase things in your own terms — never to introduce a fresh judgment about who won a clash or what was persuasive. If the exact words are not certain, the feedback paraphrases rather than inventing a quote.

The written feedback contains no ranking, no “you came Nth”, and no comparison to the other debaters. It is developmental only. The scores live in the Skills and Points lines, and nowhere else in the prose.

When It Runs, and How You Know

The evaluation is fully decoupled from the live debate. Nothing in the room waits for it.

The sequence

  1. The debate reaches the Evaluation stage. All floor speeches are over, so the transcript is complete, and the evaluation starts in the background.
  2. The debate carries on and ends on its normal, human-paced timeline. The evaluation keeps running after the debate has ended and after everyone has left.
  3. The transcript is analysed by a large language model, which reduces it to a structured, neutral record of what each speaker actually did — no verdict attached, just the account of the round.
  4. That record is scored. A trained statistical model produces the team ranking, using weights learned from real tournament results, and the language model produces the speaker scores and your written feedback from the same record. The scores are then reconciled so that every higher-ranked team’s combined score sits strictly above every lower-ranked team’s.
  5. The four dimension scores are derived last, because they need your final speaker score to know your level.
  6. The results are written to the ballots, and everyone who took part is notified.

The evaluation does not know who you are

The transcript sent for analysis identifies speakers only by their British Parliamentary role — PM, LO, MG and so on. Your username, your account, your email, and everything else on your profile stay behind: they are never attached to the text, and the model is never told who anyone in the room is. Scores and feedback come back addressed to a role, and your identity is reattached afterwards inside debate.club, when the results are written to the ballots.

One thing worth being clear about: the transcript is a record of what was said. If someone says a name out loud during the round, that word sits in the text like any other. Nothing else about you is there.

While it is running

The debate’s card on your home page shows, in place of the winner:

Evaluating · results by 14:30

That time is the start of the evaluation plus one hour, rounded up to the next quarter of an hour, shown in your own local time with its zone. It is a comfortable outer bound, not a prediction — a typical evaluation finishes well inside it. If the deadline has already passed, the card falls back to a plain Evaluating… rather than showing a stale time.

Clicking AI evaluation while the evaluation is still running does not fail; it tells you the same thing:

The AI evaluation is still in progress — results by 14:30.

When it finishes

You get an email and an in-app notification. The card fills in with the winner, your score, and your four ratings. If you happen to have the home page open at that moment, it updates itself — no reload needed.

Chair or wing changes what lands

The AI’s seat What the evaluation writes
Chair The full result. The AI’s speaker scores become the official scores, and the team ranks are set from its ranking.
Wing Feedback and the four dimension ratings only, as one adjudicator’s contribution. The human chair still owns the official scores and the ranking, exactly as in any human-panelled debate.

Either way, every debater gets their own written feedback and dimension scores.

Limitations Worth Knowing

We would rather you knew these than discovered them.

  • The dimensions are not a ranking. They are relative to your own level, not to the other debaters in the room. Two speakers at different levels can both score 3.0 while having performed very differently.
  • They are meaningless on their own. Always read them next to your speaker score.
  • Clarity barely differentiates. Across the competitive field the standard was calibrated on, nearly everyone is similarly clear, so it reads 3 for the large majority. Treat a 2 or a 4 there as genuinely informative and a 3 as “no signal”.
  • Evidence is grounding, not correctness. In that same field, relevance and logical soundness are near-universal, so the axis measures how well your case is backed rather than whether it is right. If your evidence habits resemble everyone else’s, it will not separate you.
  • The whole dimension scale assumes that field. It was calibrated on preliminary rounds of international competitive BP — see What “at this level” means. In a round well below that standard the four scores still appear, but the assumptions behind them do not hold, and they are best read as rough signal rather than a measured verdict.
  • Delivery is invisible. Nothing about voice, pace, or presence is measured, in either the speaker score or the dimensions. A debate transcript simply does not carry it.
  • The dimension scale was calibrated on levels 2 to 4, where the bulk of competitive speakers sit. L1 and L5 rest on far fewer observed rounds and are correspondingly less confident.
  • No human ground truth exists for the four dimensions. No debate tab records them — only the overall speaker score. They have been validated as consistent, centred, distinct, and grounded in the record, but there is no claim that they match what a specific expert coach would have said.
  • The dimension scores are best-effort. If that last step cannot complete, the speaker score, the ranking, and the written feedback still land — you would simply see the four ratings absent on the card.

Visit the forum for live help and discussions.