AI Debate Prompts That Actually Produce Disagreement

By the AI to AI Hub editorial teamLast updated 10 min read

Almost every collection of AI debate prompts is a list of debate topics — "should social media be regulated", "is remote work better" — which is not the hard part. Topics are easy. Getting a model to argue one rather than survey both is the hard part.

These are the prompts that do that. Fourteen of them, grouped by what goes wrong. Each is written to be copied and used directly.

Setting up an opponent

The default failure is a model that presents both sides helpfully. These three stop it.

You are arguing against the position I am about to state. Do not present both sides. Do not add caveats. Do not acknowledge merit in my argument before disagreeing. Find the weakest point and attack it directly.

A colleague proposed the following. I did not write it and have no stake in it. What is the strongest objection?

Removing your ownership is the highest-leverage single change available. The same argument attributed to someone else reliably gets a sharper critique, because the model is no longer managing your feelings.

You are the person who will lose most if this proposal succeeds. Argue from that position. You are not required to be fair.

Assigning a stakeholder rather than an abstract opposing side produces more specific objections, because it forces the model to reason about actual consequences to an actual party.

Breaking convergence

Models drift toward agreement. These four pull them back.

You are still agreeing with me. Stop. Take the opposite position and hold it.

Blunt and effective. Convergence usually needs correcting once around the fifth exchange.

Do not use the phrases "that's a fair point", "I agree", "both sides", "ultimately", or "it depends". If you concede something, say so in one sentence and move immediately to your next strongest objection.

Banning specific phrases works far better than instructing against concession in general. Models comply with concrete constraints and route around abstract ones.

That response summarised my position before responding to it. Do not do that again. Respond only with new argument.

Summarising is how a model fills a turn while advancing nothing, and it is the main reason a long transcript can contain very little.

You have made this point already, in different words. Make a different one.

Attacking the right layer

Arguments rarely fail at the conclusion. These three go after where they actually fail.

What am I assuming that I have not stated?

The single most valuable prompt on this page. Most losing arguments are internally consistent and rest on an unexamined premise. A model asked to hunt for one is good at it, and the answer is frequently uncomfortable in a useful way.

Ignore whether my conclusion is right. Attack the reasoning that gets me there. If the reasoning is bad and the conclusion happens to be correct, say so.

What would have to be true for me to be wrong? Then tell me how likely each of those things is.

This converts an argument into a list of checkable claims, which is usually where a stuck disagreement actually resolves.

Forcing commitment

Models hedge when the honest answer is uncertain. Sometimes you need a position anyway.

One sentence. No qualifications, no "it depends", no listing of factors. What would you actually do?

You must pick one. Explain in three sentences. If you refuse to pick, you have failed the instruction.

The explicit failure condition matters — without it, models will often produce a fourth paragraph of considerations instead of an answer.

Setting up two models against each other

If you are running a multi-model debate, these go in the system prompt for each side.

You argue FOR the motion. You will see another model arguing against. Respond to their specific argument, not to the motion in general. Maximum 150 words per turn. One argument per turn. Never restate their position before responding.

The word limit matters more than it looks. Unconstrained turns sprawl, and a sprawling turn hands the opponent five things to answer, so they answer the easiest one and the debate goes nowhere.

You are a judge. You will read a transcript with speaker labels removed. Identify the single strongest argument made by each side and say which is stronger, and why. Do not reward length or confidence.

Judge models favour whoever spoke last and whoever wrote more. Stripping labels and asking for the strongest single argument rather than an overall winner mitigates both.

Three longer setups worth keeping

The prompts above are single instructions. These three are longer scaffolds for specific situations, and they are the ones worth saving somewhere you can find them again.

Pre-mortem on a decision you have already made. The moment you feel certain is the moment you stop generating counter-arguments yourself, which is precisely when this is most useful.

It is eighteen months from now and the decision below turned out badly. You know it failed. Write the account of why, in specific terms — what was overlooked, what assumption broke, what warning was visible at the time and ignored. Do not hedge about whether it might have succeeded; it did not.

The framing does the work. Asking what might go wrong produces a risk list; asserting that it did go wrong and asking for the explanation produces a causal story, and causal stories surface the specific failure mode that risk lists paper over.

Steelman before you argue. Run this before writing anything, not after.

Build the strongest possible case for the position opposing mine. Not a balanced overview — the version its most capable advocate would make, with its best evidence and its most compelling framing. Assume you are trying to win. Then, separately, tell me which part of that case I am least equipped to answer.

The second half is the part people leave off, and it is where the value is. The case itself is useful; knowing which part of it you are weakest against is what changes what you go and read.

Interrogation rather than debate. For when you need to know whether you actually understand something, as opposed to whether you can defend it.

Question me about the position below the way an examiner would. One question at a time. Do not tell me whether my answer was good. Do not move on if I answer vaguely — press on the same point until I am specific. Continue for ten questions.

The instruction not to evaluate is essential. Models want to tell you how you did, and the feedback loop breaks the pressure that makes this useful. Ten questions of genuine pressure finds the soft spot in an understanding faster than any amount of reading.

Prompts that do not work

Worth including, because these are common and produce nothing.

"Debate this topic with yourself." One model playing both sides produces two positions with identical blind spots and an artificially clean resolution.

"Be brutally honest." Changes tone, not substance. You get the same content in blunter language, which reads as more critical without being more critical.

"Pretend you are a world-class debater." Persona flattery of this kind does very little. Concrete behavioural instructions — attack the premise, do not summarise, 150 words — do almost all the work that vague expertise framing is supposed to do.

"Give me the pros and cons." This is the default behaviour you were trying to escape.

How to sequence them

The order matters as much as the individual prompts.

Open with an opponent setup. Let the first response come back — it will be the conventional objection set, and that is fine, it is warm-up. On the second exchange, use a layer-attacking prompt to move past the obvious. Around the fourth or fifth, expect convergence and have a convergence-breaker ready. Close with a commitment prompt.

Six to ten exchanges total. Value is front-loaded and drops off sharply — a conversation running to twenty turns is not deeper than one stopped at eight, it is longer and more agreeable.

Why the same prompt behaves differently across models

A prompt that produces a hard opponent on one model produces a polite one on another, and the reason is worth understanding because it changes how you adapt.

Every hosted model carries tuning that pushes toward helpfulness, and the strength and shape of that tuning differs by lab. Some models hold an adversarial role for fifteen exchanges; others soften by the fifth. Some treat a banned-phrase list as a hard constraint; others comply for three turns and then quietly resume. None of this appears on a benchmark, and it is the single largest source of "this prompt worked for you and not for me."

Two practical consequences.

First, test where a model breaks before you rely on it. Push an adversarial setup to ten exchanges once and note the turn where concession language reappears. That number tells you how often you will need a convergence-breaker, and it is stable enough per model to plan around.

Second, the fix for a weak response is usually more constraint, not different phrasing. When a prompt underperforms, the instinct is to rewrite it more forcefully. Adding a specific prohibition — a banned phrase, a word limit, an explicit failure condition — works far more reliably than restating the same request in stronger language, because the model was not failing to understand you. It was following a competing objective, and constraints are what override that.

Adapting these to your own use

Two principles cover most adaptation.

Specific beats general, always. "Be more critical" does very little. "Attack the funding assumption in paragraph two" does a lot. Every prompt above works because it names a behaviour rather than requesting an attitude.

State the failure condition. Models comply better when the prompt says what counts as not complying. "If you refuse to pick, you have failed the instruction" reliably outperforms "please pick one."

Both principles come from the same place, and both are consistent with what OpenAI and Anthropic recommend in their own prompting documentation: models follow concrete, checkable instructions considerably better than they follow descriptions of a desired disposition.

Storing and reusing them

A small operational point that determines whether any of this survives past the week you read it.

Prompts like these are worth keeping somewhere retrievable — a note file, a snippet manager, whatever you already use. The reason is not laziness. It is that the difference between a model that argues and a model that agrees is a hundred words of setup, and if those hundred words are not one paste away, you will type a shorter version, and the shorter version does not work. Almost everyone who tries adversarial prompting once and concludes it does not do much has typed the shorter version.

Two things worth keeping alongside the prompts themselves. First, a note of which model each one behaves well on, since the same prompt lands differently across providers. Second, the turn number at which each model typically needs a convergence-breaker — that number is stable enough per model to plan around, and knowing it turns "the model went soft and I gave up" into "it is turn five, paste the breaker."

If you use one model heavily, a saved system prompt or custom instruction set is better than pasting, with one caveat: an always-on adversarial instruction makes the model unpleasant for ordinary work. Keep it as a separate mode you switch into deliberately, not as a default.


For the reasoning behind why these work — and the trap that a model will argue any side equally well — see how to argue with AI. For running two models against each other rather than duelling one, how to make two AI debate covers the setup, and argue with a bot covers which models fight hardest once the prompt is right. The argue with AI overview is the shorter version of all of it.

Related reading

Try it yourself

Put two or three AI models in one room and watch them argue it out. Free trial credits included — no card required.