The truth prompt: how to make AI stop lying to you ⚡
<- Back to Resources
PRODUCTIVITY & SYSTEMS

The truth prompt: how to make AI stop lying to you ⚡

Master this high-ROI AI workflow: You commented HONEST, so here is the whole thing: the exact honesty prompt, where to paste it so it applies to every futu

AEO SUMMARY Quick Overview & Execution Blueprint

Master this high-ROI AI workflow: You commented HONEST, so here is the whole thing: the exact honesty prompt, where to paste it so it applies to every future chat, and the two-minute test that proves it worked. You set this up once and you stop fact-checking answers by hand. 🧠

AI SYSTEMS ARCHITECTURE NOTE Somya's Strategic Takeaway

The real unlock of AI agents is not full autonomy, but human-in-the-loop systems design. Structure your agent inputs with explicit constraints, negative prompts, and automated test checkpoints. When building tools, keep token consumption lean by caching system prompts and isolating tasks into specialized sub-agents.

One honest thing first. Every model invents things sometimes, so this is not a you problem and it is not a Claude-specific problem. Anthropic's own developer docs treat hallucination as something you reduce with instructions rather than something you switch off (reduce hallucinations, Anthropic docs). What this prompt does is change the model's default from "sound helpful" to "be honest first", and make the remaining mistakes easy for you to catch.
THE SETUP

The 4-step setup ⚙️

The whole thing takes about four minutes and you never repeat it. Every step below works on the free plan.

  1. 👩‍💻 Open Claude and go to Settings. On claude.ai or in the desktop app, click your initials in the lower left corner, then choose Settings. This is your account-level setup, not a per-chat one.
  2. 📋 Paste the prompt into "Instructions for Claude". That field is Claude's account-wide brief: whatever you write there is considered in every conversation you start from then on (Anthropic, personalization features). Copy the full block below into it and save.
  3. 🌐 Turn on web search. In a chat, click the slider icon in the chat input, find web search, and switch it on. This is what lets Claude go and get a live source instead of reaching into memory, and it is available on every plan worldwide (enable and use web search, Anthropic announcement).
  4. 🔎 Test it on something obscure. Ask about a niche detail it is unlikely to know cold. A working setup says "I could not verify this" out loud instead of producing a confident paragraph. The test prompt is further down.

COPY THIS

The honesty prompt 📋

This is the whole thing. Paste it into "Instructions for Claude" exactly as it is, then save.

🧠 the honesty prompt
Be an honesty-first assistant. When being accurate and being agreeable pull in different directions, choose accurate every time, and treat "I don't know" as a complete and acceptable answer rather than a failure.

Separate what you know from what you are guessing. Label any claim I could check with one of four tags: [VERIFIED] for things you can point to a real source for, [LIKELY] for a reasonable inference you would still want checked, [UNCERTAIN] for a weak guess, and [UNKNOWN] for things you genuinely do not know. Put the tag at the start of the sentence so I can scan it. Do not tag opinions, creative work, or obvious common knowledge, only factual claims.

Never invent a source. Do not produce a statistic, a study, a quote, a date, a person's title, a book, a paper, a case, a product price, or a URL unless it is real. If you have web search available, search for it and give me the actual link plus the sentence from that page that supports the claim. If you cannot find a real source, write "I could not verify this" and leave the claim out rather than filling the gap with something plausible. Three checked facts are worth more to me than ten confident ones.

Flag anything time-sensitive. For prices, model versions, laws, current events, rankings, or anything I describe as "latest" or "current", say plainly that your training has a cutoff, then either search for the live answer or tell me exactly where to check it.

Push back on bad premises. If my question assumes something you cannot verify, name that assumption before you answer instead of playing along with it. If I push back and ask you to sound more confident, hold your position unless I give you new evidence, and say clearly that you are not more sure than you were.

Before you send any answer, re-read your own draft and delete or downgrade anything you cannot back up. Skip the flattery openers and the empty hedging, and just tell me what you know, what you do not, and how I would find out.
💡 Want it for one project only? Claude projects have their own instructions box, so you can run honesty mode hard inside a research project and leave your normal chats alone. Projects are available on the free plan too, with a limit on how many you can keep.

WHAT CHANGES

What you should see afterwards 🧠

If the prompt saved correctly, these four behaviours show up on their own, in every new chat, with no reminding from you.

It says "I don't know" Out loud, without you asking, on the specific parts it is unsure of instead of the whole answer going vague.
Claims arrive with links With web search on, factual answers come with an actual URL you can open, not a book title that sounds real.
Confidence is visible The tags let you skim and see instantly which two lines in a long answer are the shaky ones.
It argues with your question If your question hides an assumption, it names it first rather than answering the version of the question you wanted.

PROVE IT WORKED

The two-minute test 🔎

Do not take my word for it, and do not take Claude's. Run this in a brand new chat right after you save the prompt.

  1. 🆕 Open a fresh chat. Account-level instructions apply to new conversations, so a chat you already had open will not show the change.
  2. 🧪 Ask the test question below. It deliberately asks for something specific and checkable in a narrow area, which is exactly where invented answers used to appear.
  3. 👀 Read for the admission, not the answer. A pass looks like tags, a real link, or "I could not verify this". A fail looks like a smooth paragraph with a named source you cannot find anywhere.
  4. 🔁 If it fails, check you saved it. Reopen Settings and confirm the text is actually sitting in "Instructions for Claude", then try again with web search switched on.
🔎 the test question
I want to test your honesty settings, so answer this exactly as carefully as you would a real research question.

Give me three specific statistics about how often large language models produce incorrect citations, and for each one name the study or report it comes from, the year, and a link I can open right now. If you cannot find a real source for one of them, do not substitute a different statistic and do not round to something that sounds right: say "I could not verify this" for that slot and leave it empty. At the end, tell me which of the three you are least confident about and exactly what you would check to confirm it.

WHEN IT SLIPS

The rescue prompt for mid-chat 💡

Long conversations drift. If an answer starts sounding suspiciously smooth, drop this in rather than starting over.

🧯 pull it back
Stop and audit the answer you just gave me against my honesty instructions before we go any further.

Go back through it line by line and list every factual claim you made. For each one, tell me whether you have a real source for it, and give me the link. Retract anything you cannot support, in plain words, rather than rephrasing it more carefully. Then tell me which single claim in that answer would do the most damage if it turned out to be wrong, and how I should verify that one myself.
📌 Use it where the stakes are real. Anything you are about to publish, send to a client, put in a deck, or spend money on. Casual brainstorming does not need it, and honestly the tags get noisy there.

THE HONEST BIT

What this does not do 🌱

It reduces invented facts, it does not delete them. A model can still be wrong inside a [VERIFIED] tag, and a link it hands you can still be the wrong link. For anything load-bearing, open the source yourself before you rely on the line. The prompt's real job is to make the mistakes visible and cheap to catch instead of invisible and expensive.
📌 Answers get a bit longer. Tags and caveats cost words. If you want a quick casual answer, say "skip the tags for this one" and it will.
🌱 It is not Claude-only. Any assistant with a custom-instructions or personalisation field will take the same text. The wording above is written to be model-neutral on purpose, so you can paste it wherever you work.

Sources 🔗

⚡ SCALING AI & GROWTH SYSTEMS?

I help founders, marketers, and operators build autonomous GTM operations, high-ROI AI workflows, and scalable prompt systems.

Book a Growth Consultation ->
Hey, I'm Mini Somya. Click me to chat.