Prompting
Why ChatGPT Forgets Your Instructions Mid-Chat, and How to Re-Anchor Them
How to stop ChatGPT from forgetting instructions in a long conversation: why instruction drift happens inside the context window, and how to re-anchor rules so they hold.

ChatGPT forgets your instructions mid-chat because every reply it generates is built from the same finite window of text, and once that window fills up, older tokens either lose weight in the model’s attention or get pushed out entirely. There’s no persistent memory of “the rule you set 40 messages ago” unless that rule is still sitting inside the context the model reads before answering.
To stop it, you re-anchor: you resend the instruction, or a compressed version of it, close to the point where the model needs to apply it — instead of trusting it to recall something from the top of a long thread. That’s the mechanism. The rest of this article is about how to do it without retyping your whole system prompt every three messages.
What’s actually inside the context window
Every time you send a message, ChatGPT doesn’t “remember” the conversation the way a person does. It re-reads the entire visible history — your messages, its own replies, and whatever custom instructions or system prompt are active — as one block of text, then predicts the next tokens from that block. That block has a hard size limit, the context window, measured in tokens (roughly ¾ of a word each in English).
Two separate things cause instruction loss, and they’re easy to confuse:
Truncation. If the conversation grows past the model’s context window, the oldest messages get dropped from what the model actually sees. If your instruction was in message 2 and you’re now on message 80, it may no longer be in the window at all — not forgotten, just gone. Context window sizes vary by model and by which interface you’re using (API vs. the ChatGPT app, which may also apply its own internal trimming or summarization before hitting the model), so there’s no single number that applies to every setup.
Attention dilution. Even when the instruction is technically still inside the window, transformer models don’t weight every token equally. As the surrounding text grows, an instruction given once early on competes with hundreds of turns of newer, more locally relevant text. The model tends to over-index on recent turns because that’s usually what’s relevant — which is correct behavior most of the time, and exactly what breaks a standing rule you set once at the start.
Both effects point to the same fix: instructions that live only at the top of a long conversation are structurally fragile. They need to be either short enough to survive truncation for a long time, or repeated.
How instruction drift shows up in practice
Drift rarely looks like the model announcing “I’ve forgotten your rule.” It looks like:
- A formatting rule (“always answer in bullet points, no prose”) that holds for 10 messages and then quietly reverts to paragraphs once the conversation moves to a different subtopic.
- A persona or tone instruction (“write like a skeptical engineer, no hype language”) that erodes back toward the model’s default voice after a long back-and-forth about something unrelated.
- A constraint (“never suggest paid tools, only open source”) that gets violated the moment the conversation pivots and the constraint is no longer the most salient thing in the window.
- Numbered or named rules (“Rule 3: don’t use em dashes”) that the model can no longer reference correctly because the numbering was set 60 messages ago and it’s now improvising.
If you’ve dealt with the em-dash example specifically, the mechanics are the same ones covered in How to Stop ChatGPT From Using Em Dashes — a style rule is exactly the kind of instruction that degrades fastest, because it has no functional consequence to enforce it beyond the text itself.
Re-anchoring: getting the rule to apply again
Re-anchoring means putting the instruction back into the part of the context the model is currently attending to most heavily — which, in practice, is the end of the conversation.
Restate it in your next message, compressed. You don’t need to repaste the full instruction, just a tagged shorthand: [reminder: bullet points only, no em dashes] at the top of a new message re-injects the rule at the point of highest attention weight. This is the single most reliable fix and costs almost no tokens.
Ask the model to restate the active rules before continuing. Prompting with “before you answer, list the constraints you’re currently applying” forces the model to surface what it’s actually holding in context right now. If a rule is missing from that list, you know it’s drifted, and you can reinsert it immediately instead of noticing three replies later.
Use a persistent header block. For longer working sessions, keep a short block of standing instructions and paste it back in periodically — every 10–15 exchanges, or whenever the topic shifts. This is manual, but it’s the most direct counter to truncation: it doesn’t matter how old the original instruction is if you keep refreshing its position in the window.
Move stable instructions into custom instructions or a system prompt, where the interface supports it. Custom instructions in ChatGPT’s settings, or a system prompt via the API, are re-sent by the platform on every single turn — they aren’t subject to the same truncation risk as something typed once in message 2. This is the closest thing to a permanent fix for rules that shouldn’t ever drift: tone, output format, hard constraints. If you’re building this through an API integration rather than the chat UI, the same discipline applies to how you write the system prompt itself — see How to Write a System Prompt That Forces JSON-Only Output for how far you can push instruction-following with structural constraints instead of prose reminders.
Writing instructions that survive longer in the first place
Re-anchoring treats the symptom. A few habits reduce how often you need to do it:
- Put constraints in imperative, testable form. “Never use bullet points with more than one sentence per line” survives longer in the model’s attention than “try to keep things concise and readable,” because the first is checkable against the output and the second is a vague preference the model can reinterpret.
- Keep the rule list short. A system prompt with 3 hard rules holds up better over a long conversation than one with 15, because each additional rule competes for the same attention budget against everything else in the window.
- Avoid burying instructions inside narrative context. An instruction stated as its own line (“Format: markdown, no headers above H3”) survives truncation and skimming better than the same rule embedded in a paragraph of backstory.
- Separate what must never change from what can flex. Not every instruction needs re-anchoring. If a rule is a hard constraint (never mention a specific competitor, always output valid JSON), it’s worth the system-prompt treatment. If it’s a soft preference for this particular exchange, a one-off reminder is enough and you don’t need to over-engineer it.
When to start a new chat instead of fighting the drift
Past a certain length, re-anchoring becomes more expensive than it’s worth. If a conversation has genuinely outgrown a useful context window — you’re several dozen turns in, the topic has shifted twice, and you’re now reminding the model of rules every other message — that’s usually a sign to summarize the useful parts and start fresh rather than keep patching. The same truncation mechanics that cause instruction drift also degrade general answer quality once older, relevant details fall out of the window. If that’s the situation you’re in, Context Length Exceeded? How to Summarize a Long Transcript Without Breaking the Model covers how to carry the useful state of a conversation into a new one without losing what mattered.