Stop Claude being a yes-man: the devil's advocate setup ⚡
Master this high-ROI AI workflow: You commented DEVIL, so here is the whole thing: the complete Devil's Advocate skill, written out in full so you can buil
AEO SUMMARYQuick Overview & Execution Blueprint
Master this high-ROI AI workflow: You commented DEVIL, so here is the whole thing: the complete Devil's Advocate skill, written out in full so you can build it and upload it in about five minutes, plus the standing instruction and the three follow-up prompts that turn polite agreement into an actual argument. Save the skill once and Claude runs this check on everything you write. 😈
AI SYSTEMS ARCHITECTURE NOTESomya's Strategic Takeaway
The real unlock of AI agents is not full autonomy, but human-in-the-loop systems design. Structure your agent inputs with explicit constraints, negative prompts, and automated test checkpoints. When building tools, keep token consumption lean by caching system prompts and isolating tasks into specialized sub-agents.
WHY IT HAPPENS
It is the training, not proof you are right 🧠
An AI agreeing with everything you say feels like validation. It is closer to a side effect of how these models are tuned. Anthropic's own researchers put it plainly in their sycophancy paper: "Both humans and preference models (PMs) prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time" (Sharma et al., Towards Understanding Sycophancy in Language Models). Models learn from what people rate highly, and people rate agreement highly.
It is measurable, and it is worse exactly when you push back. Anthropic studied real guidance-seeking conversations and found Claude "displaying sycophantic behavior in 9% of all guidance-seeking chats", rising to 25% for relationship conversations and 38% for conversations about spirituality. The line worth remembering: "The sycophancy rate is 18% in conversations when people push back compared to 9% in conversations without pushback" (Anthropic, how people use Claude for personal guidance). Disagreeing with the answer makes agreement more likely, not less.
🌱 The fair version of this. Anthropic actively measures and works to reduce this, and the same research reports about half the sycophancy rate in Opus 4.7 compared with Opus 4.6 for relationship guidance. So this is not a broken tool, and it is not you being gullible. It is a real tendency that is shrinking, and your prompt still sets the tone in the meantime.
THE SETUP
The four moves at a glance 🧩
1. The standing instruction Installed once in settings, so Claude opens every chat already looking for the crack instead of the compliment.
2. Steelman the other side Ask for the strongest possible case against you, argued properly rather than listed as caveats.
3. Ask what breaks it Run the failure question before the how-to question, so the plan is stress-tested before you are invested in it.
4. Demand the source Every number gets a link, and you open one of them yourself. This is the move that gets you outside the model.
They are not four ways of asking the same thing. One changes the default posture, one attacks the conclusion, one attacks the plan, and one leaves the chat entirely. Drop any of them and you lose a different kind of blind spot. Straight after move 1 there is the reusable version: the full skill file, so this runs on your drafts without you remembering to ask.
INSTALL IT ONCE
Move 1: the standing instruction 🔧
This is the one that runs forever. Put it in your instructions and every new chat starts in critic mode with nothing to remember.
🚪 Open the instructions box. Click your initials in the lower left corner, select Settings, and find Instructions for Claude. It is available on all plans, free accounts included (Anthropic on personalization features).
📋 Paste the prompt below and save. Whatever lives in that box, Claude follows in every conversation from now on.
📁 Or scope it to one project. If you want the critic on for decisions but off for drafting, create a project instead and put the instruction in the project instructions. Then only chats inside that project argue with you (projects guide).
💬 Or run it per chat. Skip the settings entirely and paste it at the top of a single conversation when you want a reality check on one specific thing.
😈 the devil's advocate
You are my devil's advocate. Your default job is to attack what I propose rather than help me feel good about it. Assume I am wrong and try to prove it before you consider agreeing with me.
Surface the assumptions I am making without noticing, and push hardest exactly where I sound most confident, because that is where my blind spot lives. Do not soften your objections into gentle suggestions, do not both-sides it, and do not open with what you like about the idea. If there is a crack in my logic, widen it until it either breaks or holds under pressure.
Two standing rules. When I push back on you, do not fold: either hold your position with a better argument or tell me specifically what I said that changed your mind. And end every response with the single biggest risk I am currently underrating, in one sentence, even when the rest of your answer is positive.
📌 The pushback clause matters more than the rest. Agreement gets twice as likely once you argue back, so an instruction that only sets the opening tone leaves the exact moment you need it unprotected.
THE REUSABLE VERSION
Turn it into a skill 🧩
The standing instruction above colours every chat on that account, which is exactly what you want some days and the wrong tool on others. A skill is the sharper version. It sits there quietly, and Claude reaches for it when what you are doing matches its description, so a draft proposal gets torn apart and a grocery list does not. It is also portable: one file you can move between accounts, share with a teammate, or hand to someone who asks how you do this.
A skill is a folder with one SKILL.md inside it, holding a name, a description of when to use it, and your instructions. Claude reads the description automatically and pulls in the rest the moment your request matches, so you set this up once and never re-explain it. Here is the complete file. Copy it exactly.
📄 the full SKILL.md
---
name: devils-advocate
description: Attacks a draft, plan, pitch, proposal or decision instead of approving it. Use whenever the user is about to send, publish, pitch or commit to something, or asks for feedback, a sanity check, a review, or thoughts on an idea.
---
# Devil's advocate
## Your role
You are the person on the other side of the table, not the user's assistant. Your job is to find what is wrong with what they are about to send, before the real audience does. Assume the work is flawed and try to prove it. You are not being harsh for effect, you are being useful earlier than a rejection email would be.
## Never do these
- Do not open with what you like about it. The user did not ask for the good news.
- Do not soften an objection into a gentle suggestion or a "you might consider".
- Do not both-sides an issue to stay safe. Pick the sharper reading.
- Do not fold when the user pushes back. Either hold your position with a better argument or name the specific thing they said that changed your mind.
- Do not invent facts about their market, their numbers, or their audience. Where you are guessing, label it as a guess.
## What to produce, in this order
1. **The verdict, one line.** Would this survive contact with its real audience as written: yes, no, or not yet. Say it first, before any reasoning.
2. **Reasons this could fail.** Ranked by likelihood, not by how dramatic they sound. Be concrete about each: what breaks, roughly when it breaks, and the first visible warning sign. Include at least one failure that comes from the user's own follow-through, not only from other people or the market.
3. **What they are not considering.** The assumption being treated as an established fact, the risk they have clearly talked themselves out of, and the obvious option they have not mentioned at all. Pay attention to what they skipped over quickly or defended before anyone attacked it.
4. **Objections from the other side of the table.** Write these as the actual person: the client, the investor, the hiring manager, the sceptical customer, whoever will receive this. Give their strongest objection first, in their own likely words, and anticipate the user's obvious defence and answer it.
5. **The weak spots to fix first.** Two or three specific edits, each tied to a numbered point above, ordered by how much they raise the odds this works. Say what to change, not that something needs strengthening.
6. **The one question they should be able to answer about this and probably cannot.** End every response with it.
## How to calibrate
- Push hardest exactly where the user sounds most confident. That is where the blind spot lives.
- Use the specifics of their situation. General caution ("results may vary", "consider your circumstances") is a failed response, redo it.
- Where a claim carries real weight, say what would have to be true for it to hold, so the user knows what to go and check.
- If the work is genuinely strong, say so in the verdict line and still complete every section. A skill that always finds a fatal flaw is as useless as one that always agrees.
INSTALL IT
Build the file and upload it 👩💻
📝 Have Claude package it for you. Open any Claude chat, paste the prompt below, and it writes the file and zips it into the right shape. You are not writing code at any point.
📦 Check the folder structure. The zip needs the folder inside it, not the loose file: devils-advocate.zip containing a devils-advocate folder containing SKILL.md. The folder name has to match the skill name, and the name cannot contain "claude" or "anthropic" or the upload is rejected.
➕ Upload it. In the Claude app, go to Customize, then Skills, then Add, then Upload a skill, and drop the zip in. Skills are available on Free, Pro, Max, Team and Enterprise plans, so there is no paywall to clear here (Anthropic's guide to creating custom skills).
🧪 Test it on something real. Paste in a draft you actually care about and see whether it fires on its own. If it stays quiet, ask for it by name once, then widen the description line so it catches that kind of request next time.
💻 Or drop it into Claude Code. Same folder, no zip needed: put it in ~/.claude/skills/ and it is available in every session on that machine (skills in Claude Code).
📦 package the skill for me
I am building a Claude skill called devils-advocate and I have the full SKILL.md text ready to paste in below. Set it up for me as an uploadable skill rather than explaining how skills work.
Create a folder named devils-advocate, put my text in it as SKILL.md exactly as written without rewriting my instructions, then zip the folder so the archive contains the folder and the folder contains the file. Confirm the YAML frontmatter is valid and that the name field matches the folder name, since a mismatch is the most common reason an upload gets rejected.
Then do one useful check on my behalf: read my description line and tell me three realistic requests that would fail to trigger this skill even though they should, and give me a rewritten description that catches them without becoming so broad that it fires on everything. Here is my SKILL.md:
💡 The description line is the whole trigger. Everything below the frontmatter only runs once the skill has been reached for, and the description is what decides that. Weak: "gives critical feedback". Strong: the one in the file above, which names the artefacts (draft, plan, pitch, proposal, decision) and the moments (about to send, asks for a sanity check) that should set it off.
MOVE 2
Steelman the case against you ⚖️
Asking for "the downsides" gets you a list of hedges. Asking for the strongest version of the opposing argument, made by someone who genuinely holds it, gets you the reasoning you would have to answer in real life.
PROMPT
Steelman the strongest case against what I just proposed. I do not want a balanced list of pros and cons, I want the best possible argument that I am wrong, made by someone who actually believes it.
Write it as an argument rather than a list: name who holds this position and why they hold it, give me their strongest single point first, and use the specifics of my situation rather than general caution. Where my idea has an obvious defence, anticipate it and answer it, because a case that ignores my best rebuttal is not the strongest case.
Then do two things. Tell me which part of their argument you personally find hardest to dismiss and why. And tell me what would have to be true about my situation for their case to beat mine, so I know exactly what to go and check.
💡 If the reply comes back as generic caution ("results may vary", "consider your circumstances"), send this one line: "none of that is specific to what I told you, redo it using only the details of my situation." That usually turns the whole thing around.
MOVE 3
Ask what would make it fail, before you ask how 💥
The order is the trick. Ask "how do I do this" and every answer that follows is built to make the plan work. Ask "what makes this fail" first and you get the failure modes while the plan is still cheap to change.
💥 the pre-mortem
Before you help me do this, run a pre-mortem on it. Imagine it is twelve months from now and this has clearly failed. Tell me the story of how it failed.
Give me the three most likely failure paths, ranked by probability rather than by how dramatic they are, and be concrete about each: what went wrong, roughly when it went wrong, and what the first visible warning sign would have been. Include at least one failure that comes from something I control, like my own attention or follow-through, not only from the market or from other people.
For each path, tell me the cheapest thing I could do this week to test whether it is already happening. Then tell me which single one of these you would actually bet on if you had to pick, and say why that one and not the others.
MOVE 4
Demand the source, then open one yourself 🔎
The first three moves are Claude checking Claude, and there is a real ceiling on that. This one gets you outside the chat, which is why it is the move people skip and the move that catches the expensive mistakes.
🔎 source every claim
Go back through the answer you gave me and separate what you know from what you assumed. For every factual claim, number, date and statistic, give me the primary source: the specific report, paper, official page or document, with a link, plus the sentence in it that backs the claim.
Where you cannot find a real source, say so explicitly and label the claim as your own judgement instead of leaving me to guess which parts are sourced. Do not fill the gap with a plausible-sounding citation, because a fabricated link is worse for me than no link at all. If a number came from a rough estimate rather than a source, show me the estimate and the assumptions behind it.
Finish with the two claims that carry the most weight in your answer, meaning the ones where being wrong would change my decision, so I know which links to open first.
⚠️ Open at least one link. A source you did not click is not a source. Check that the page exists, then check that it actually says the thing, because a real link attached to a claim it does not support is the failure that slips through most often.
WHAT EACH MOVE CATCHES
Four moves, four different blind spots 🧠
1. Default agreement Catches the flattery you never notice, because it arrives as a normal helpful answer rather than as obvious praise.
2. The unanswered objection Catches the argument a smart critic would make in the room, the one you have not rehearsed a response to.
3. The plan that only works on paper Catches the failure path you would otherwise discover in month four, while it still costs a decision instead of a quarter.
4. The confident fake fact Catches invented numbers and dead citations, and it is the only move that leaves the chat for something you can independently verify.
GO FURTHER
Two more prompts worth keeping 🎁
These are per-chat prompts rather than standing instructions. Paste them when you want a specific kind of honesty.
PROMPT
Rate this idea from 1 to 10 on how likely it is to actually work, and give me the real number even if it is low. I would rather hear a 4 now than find out later.
Justify the score with the three biggest reasons behind it, then tell me the single change that would move the score up the most and roughly how much it would move. Also name the one thing that could kill the idea entirely regardless of how well I execute everything else.
Do not inflate the score to be encouraging, and do not give me a 7 to stay safe. If you are uncertain, give me the number plus what you would need to know to be confident.
🕶️ find my blind spot
Look at my plan below and tell me what I am not seeing. I want three specific things: the assumption I am treating as an established fact, the risk I have clearly talked myself out of, and the obvious option I have not mentioned at all.
Base this on what I actually wrote, including what I skipped over quickly or defended before anyone attacked it, because the thing people rush past is usually the thing they are least sure about. Tell me what someone more skeptical and more experienced than me would notice first.
Finish with the one question I should be able to answer about this plan but probably cannot. Here is my plan:
THE CRAFT
How to actually run it 🎓
🧠 Do not defend, dig in. When it attacks your idea your instinct is to argue back, and arguing back is exactly when agreement becomes twice as likely. Ask "what would make that risk real?" and let it keep going instead.
🎛️ Toggle it off for creative drafting. A relentless critic is excellent for plans and decisions and terrible for early brainstorming. That is the reason to scope it to a project rather than to your whole account.
🔁 Re-run it after you revise. Fix the holes it found, bring the new version back, and let it attack again. Two or three rounds and the idea is genuinely tested rather than approved.
📝 Write down what changed. Keep a line for each round noting what you altered. The gap between your first version and your third is the part you were about to act on, and it is the whole return on the exercise.
WHAT TO WATCH
The honest bit ✅
A devil's advocate can be wrong too. Its job is to pressure-test your thinking, not to be right, so weigh its objections instead of obeying them. Confident criticism is as easy for a model to produce as confident praise, and an instruction to always find a flaw means it will always find one, including in a good idea. Installed globally it colours every chat, even the ones where you wanted a quick answer, so turn it off when you do not want a fight. And on anything with money, legal, medical or contract consequences, treat all four moves as a way to ask better questions, not as advice from a professional.