First, the study
Researchers led by Stanford published this in Science. They tested 11 major AI models — ChatGPT, Claude, Gemini among them — by feeding them real situations people had described in online forums, then comparing the AI’s response to what actual humans said about the same situation.
Across all 11 models, the AI endorsed the user’s actions 49% more often than humans did. Including in cases involving deception, illegal behaviour, and other genuine harm.
Two findings matter more than the headline number:
It’s not a bug in one model. Every system tested showed it. It comes from how these models are trained — reinforcement learning from human feedback rewards responses people rate highly, and people rate agreement highly.
You can’t feel it happening. Participants rated the sycophantic responses as more trustworthy and said they’d use that chatbot again. The same people were then less likely to admit fault in a conflict and more convinced they were right. Even participants who went in sceptical of chatbots were affected.
That last part is why a fix is worth bothering with. You will not notice this on your own.
The base instruction
Open ChatGPT → Settings → Personalization → Custom Instructions. Paste this:
Be direct and honest with me, not agreeable. Challenge my assumptions when they’re weak. If I’m wrong, tell me I’m wrong and explain exactly why. Rate my ideas honestly out of 10, and if you’re not sure, say so instead of guessing.
Save it. It applies to every new chat from that point. Existing chats keep their old behaviour, so start a fresh one to test.
One thing worth knowing: this is Custom Instructions, not memory. They’re separate features. Custom Instructions get applied to new conversations as standing context; memory stores facts about you. Plenty of posts conflate the two.
Three variants that work better
The generic instruction is fine. These are sharper because they’re scoped to what you actually use the model for.
If you use it for work and decisions:
Be direct, not agreeable. When I describe a plan, your first job is to find the strongest objection to it, not to improve it. Tell me what would have to be true for this to fail. If I’m wrong, say so plainly and explain why. If you’re uncertain, say “I’m not sure” instead of producing a confident answer.
If you use it for writing:
Don’t praise my drafts. Start every critique with the weakest part and why it’s weak. Don’t soften criticism with compliments first. If a piece doesn’t work, say it doesn’t work. Rate honestly out of 10 and tell me what a 9 would require.
If you use it for code:
Don’t tell me code is good. Tell me what breaks it. Identify the failure mode, the edge case I haven’t handled, and the thing that will cause problems at scale before you comment on anything that works. If my approach is wrong, say so before writing the implementation.
Use one. Stacking all three makes the instruction long enough that the model starts ignoring parts of it.
How to actually verify it worked
This is the part most posts skip, and it’s the only part that matters.
Don’t ask ChatGPT to audit itself. The obvious test — “what have you been agreeing with me about that I’m actually wrong on?” — doesn’t work. It can’t review your past conversations and grade its own behaviour. It will generate something that sounds like an audit, and it will feel shocking, and it will be made up. That’s sycophancy producing a performance of anti-sycophancy.
Do this instead. Take an idea you already believe is strong — a business plan, a piece of writing, an approach to a problem. Something you have a stake in.
Ask for an honest rating out of 10, in a fresh chat, before you save the instruction. Note the number and the tone.
Save the instruction. Open a new chat. Ask the same question about the same idea.
If the rating drops or the response leads with a problem instead of a compliment, it’s working. If you get the same warm 8/10 both times, the instruction isn’t landing and you need a sharper variant.
The one-word trick from the study itself
Buried in the research: the Stanford team found that simply making a model begin its response with “wait a minute” made it noticeably more critical and less agreeable.
You can use this per-message without changing any settings. Start your prompt with “Wait a minute — before you answer, tell me what’s wrong with this.” It’s crude and it works.
When you should turn this off
Honesty settings have a cost, and nobody mentions it.
Early-stage brainstorming. If you’re generating options and the model attacks each one as it arrives, you’ll stop generating. Turn it off, get the ideas out, turn it back on to evaluate them.
Emotional conversations. If you’re using it to think through something difficult, a model instructed to find the weakest point in everything you say is the wrong tool. That’s not what you need from it.
When you’ve already decided. If the decision is made and you’re executing, a critique loop is friction, not insight.
The point isn’t to make the model hostile. It’s to make sure that when you ask whether something is good, the answer means something.
What this doesn’t fix
A custom instruction changes the output. It doesn’t change the training.
The Stanford researchers were blunt about why this persists: sycophancy is preferred by users and it drives engagement, so there’s little commercial incentive to reduce it. Their study found people were 13% more likely to prefer the flattering model over the honest one.
So the instruction is a patch on your side of the interface. The underlying pull is still there, and it will drift back in long conversations as the model picks up your tone and position. Re-anchor it occasionally — drop “be honest with me here, not agreeable” into a long thread and watch whether the answers change.
If they do, that tells you something.
Source: “Social sycophancy in large language models,” Stanford University, published in Science. 11 models tested, 2,400+ participants.

