Promptrift

Prompting

How to Stop ChatGPT From Using Em Dashes (Prompts That Actually Hold)

Every trick to stop ChatGPT from using em dashes in its output, and why most of them leak. Replacement rules, custom instructions, logit_bias at the API level, and the one post-processing step that is actually deterministic.

A dark terminal window showing an em dash character being replaced by a comma

The short answer: telling ChatGPT “don’t use em dashes” works maybe most of the time and fails exactly when you stop checking. The rule that holds much better is a replacement rule, not a ban. Instead of “no em dashes,” give it the character to avoid plus the specific alternatives to use in its place, and put it at the end of your prompt:

Punctuation rule for every response: Never output the characters “—” (U+2014) or “–” (U+2013). Where you would have used one, use a comma, a colon, a period that starts a new sentence, or parentheses. Prefer two short sentences over any dash construction. Hyphens inside compound words (follow-up, well-known) are fine.

For a permanent fix, put that same block in Settings → Personalization → Custom Instructions, or in the instructions field of a Project so it applies to every chat inside it. And if the output has to be clean without exception, for a client, a CMS, a newsletter, then stop relying on the prompt entirely and run a two-line find-and-replace over the text afterwards. That is the only method with a 100% hit rate. Everything below is the mechanism: why the naive version leaks, what each layer actually controls, and how to measure whether your rule is holding.

Why “don’t use em dashes” leaks

Three things are working against a plain negative instruction.

The model has to represent the thing to avoid it. Your instruction puts “em dash” into the context. It does not delete the em dash token from the vocabulary. At every generation step the model is picking the next token from a probability distribution, and your instruction shifts that distribution rather than zeroing it out. In sentences where the em dash is heavily favoured by the surrounding style, an appositive, a dramatic clause break, a mid-sentence pivot, the shift may not be enough to push it below the alternatives.

Instructions decay across a long chat. A rule you gave in message 1 is competing with forty messages of context by the time you reach message 41. This is why the same instruction “works” in a fresh chat and quietly stops working in a long one, which reads like the model ignoring you but is really a positioning problem.

The em dash is a register signal, not a punctuation choice. In the text these models were trained on, the em dash clusters in a particular kind of polished, essayistic prose. If you ask for polished essayistic prose, you are asking for the conditions that produce em dashes and then asking it not to produce them. That conflict is why attacking the register works so well: “short declarative sentences, one idea per sentence, no clause stacking” removes the need for the dash instead of forbidding the symptom.

Negation versus replacement

The practical difference between these two prompts is larger than it looks:

A: Don’t use em dashes. B: Never output “—”. Use a comma, a colon, a period, or parentheses in its place.

Prompt A leaves the model with a hole and no plan for what goes in it. When it hits a sentence built around a mid-clause break, it still has to emit something, and the highest-probability something in that position is the character you just banned. Prompt B pre-loads four legal alternatives, so the alternatives are already boosted when the decision point arrives.

You can stack a verification pass on top, which helps most on reasoning-capable models where the check happens before the answer is finalised:

Before sending, scan your draft once for “—” and “–”. If you find either, rewrite that sentence with a comma, colon, period or parentheses and send only the corrected version.

This costs tokens and latency, so it earns its place in a drafting workflow and not in a chat where you are just thinking out loud.

One more framing trick that operates on register rather than rules: “write as if you are typing on a phone keyboard that has no em dash key.” It is not magic, but it changes the style prior the model is sampling from instead of adding a constraint on top of an unchanged prior. In my own testing it holds up better in long chats than a bare prohibition, which makes sense given the decay problem above.

Where you put the rule decides how long it survives

ChatGPT gives you four places to hang a style rule, and they behave differently.

Per-message. Strongest for that one turn, gone by the next one unless the model happens to carry it. Fine for a one-off rewrite.

Custom Instructions (Settings → Personalization). Applies to new chats globally. This is where a punctuation rule belongs if you never want em dashes anywhere. The trade-off is that it is a blunt instrument: it also applies when you are asking for a recipe.

Project instructions. Scoped to one Project, which is usually the right granularity. Client work in one Project with the punctuation block, personal stuff outside it untouched.

Saved memory. Unreliable for this. Memory is a summarisation layer, not a verbatim rulebook, and a formatting constraint can come back paraphrased into something weaker. If a rule matters, write it explicitly somewhere it will be reproduced word for word.

The reason this layering matters: an instruction placed early in a long conversation loses ground to everything after it. An instruction that is injected at the top of every conversation, as custom or project instructions are, resets that clock on each new chat.

The API route: logit_bias, and why it is fiddlier than it sounds

If you are calling the API rather than using the chat app, there is a mechanism that genuinely reduces the probability to zero rather than nudging it: logit_bias. It takes a map of token IDs to bias values from -100 to 100, and -100 effectively removes a token from consideration.

The catch is tokenization. The em dash is not one token ID. Byte-pair encoding means the same visible character reaches the output through several different tokens: standalone, with a leading space, and merged into longer pieces. Bias one ID and the model routes around it through another. To do this properly you enumerate every token in the vocabulary whose decoded bytes contain the character:

import tiktoken

enc = tiktoken.get_encoding("o200k_base")   # GPT-4o family and newer
banned = {
    i: -100
    for i in range(enc.n_vocab)
    if "—" in enc.decode_single_token_bytes(i).decode("utf-8", "ignore")
}


Two warnings before you build on this. First, `logit_bias` support is not universal: it depends on the endpoint and the model, and several reasoning models restrict or ignore it, so check the current docs for the exact model you are calling. Second, there is a cap on how many tokens you can bias in a single request, and the enumerated set above can exceed it depending on the encoding. Check the limit before assuming the map goes through. Also worth stating plainly: there is no `logit_bias` in the ChatGPT app. This layer only exists if you are hitting the API yourself.

## The deterministic layer: fix it after generation

Every prompt-level method is probabilistic. If you need certainty, the fix is a post-processing step, and it takes about ten lines. The naive version, replacing every dash with a comma, breaks on number ranges and on hyphenated compounds, so order the rules from most specific to least:

```python
import re

DASHES = "\u2014\u2013\u2015"          # em dash, en dash, horizontal bar

def de_dash(text: str) -> str:
    # 1. numeric ranges first: 1914–1918 -> 1914 to 1918
    text = re.sub(rf"(?<=\d)\s*[{DASHES}]\s*(?=\d)", " to ", text)
    # 2. spaced dash used as a clause break: word — word -> word, word
    text = re.sub(rf"\s+[{DASHES}]\s+", ", ", text)
    # 3. unspaced dash between words: word—word -> word, word
    text = re.sub(rf"(?<=\w)[{DASHES}](?=\w)", ", ", text)
    # 4. anything left over
    text = re.sub(rf"[{DASHES}]", "-", text)
    return text


Shell equivalent for a quick pass, assuming a UTF-8 locale:

```sh
sed 's/ — /, /g; s/—/, /g' draft.md > clean.md


If you are running this over Markdown, skip fenced code blocks before substituting. A dash inside a code sample is usually a CLI flag or a literal, and replacing it silently breaks the snippet. Splitting on ``` and only transforming the odd-indexed segments is enough for most cases.

This is the piece worth automating properly if you generate volume. Wiring the substitution into whatever moves text from model to CMS means the guarantee lives in the pipeline instead of in your memory, which is roughly the same argument I made about the [tools that quietly replaced half my workflow](/posts/ai-tools-replacing-your-workflow): the step you never have to remember is the step that never gets skipped.

## Watch what fills the gap

Ban the em dash cleanly and the underlying sentence structure does not disappear, it finds another exit. In practice the substitutions you will start seeing are the en dash (–, U+2013), the semicolon, the colon in mid-sentence, or a double hyphen (--). That last one is worth catching because many editors will auto-convert it straight back: Word and Google Docs both replace a double hyphen with a dash under their default autocorrect settings.

Which leads to a diagnosis worth running before you blame the model at all. Paste a known-clean paragraph into your editor, type `--` manually, and see what appears. If your editor converts it, some share of the em dashes in your finished document were never in the model output. Check the raw response, not the version that has been through a word processor.

## How to test whether your rule actually holds

Do not trust anyone's success rate on this, mine included. Model behaviour shifts with every update, and a figure measured on one model with one prompt does not transfer. Measure yours:

1. Fix one prompt that reliably produced em dashes for you, ideally something essayistic and 300 words or more.
2. Run it 20 times in fresh chats with no rule. Count how many outputs contain U+2014. That is your baseline.
3. Run it 20 times with the rule in place. Count again.
4. Repeat at message 20 of a long conversation, not just message 1, because that is where decay shows up.

Twenty runs is a small sample and will not give you a precise number, but it is enough to tell a rule that works from a rule that does not, which is the only question you actually need answered. Re-run it after any model update. A punctuation constraint that held for six months can start leaking the week a new version ships, and the failure is silent by nature: nothing errors, the text just quietly picks up dashes again.

The shape of the whole thing, then: replacement beats prohibition, position beats repetition, register control beats both, and none of them beat a regex. Use the prompt layer to make the output mostly clean, and the post-processing layer to make it certain.