Promptrift

AI tools

Context Length Exceeded? How to Summarize a Long Transcript Without Breaking the Model

The context length exceeded error when summarizing a long transcript is an arithmetic problem, not a model problem. Here's how input, output and reserved tokens are counted, and the four mechanisms that bring a transcript back under the limit.

Terminal window showing a token count overflowing a model's context limit

A “context length exceeded” error means one thing: the total token count of your request — system prompt + transcript + conversation history + the space reserved for the model’s answer — is larger than the model’s context window. It is arithmetic, not a bug, not a bad API key, and not something a retry will fix. The transcript is the obvious culprit, but in practice the two things that push a request over the edge most often are (a) timestamp and speaker-label overhead in the raw subtitle file, which can cost more tokens than the actual speech, and (b) a max_tokens value that reserves a chunk of the window before your input is even counted.

So the fix path, in the order that resolves the most cases: strip the transcript down to plain speech before it ever hits the API, count the tokens locally with the model’s own tokenizer, lower the reserved output budget to something the summary actually needs, and — only if it still doesn’t fit — split the transcript into chunks and summarize in two passes. Moving to a model with a bigger window works too, but it changes the failure mode rather than removing it, and it comes with its own costs. Below is what each of those mechanisms is actually doing, and what changes depending on your provider, your file format and your workflow.

What the error is really telling you

The typical OpenAI-style message follows this shape:

This model’s maximum context length is 128000 tokens. However, your messages resulted in 131,420 tokens. Please reduce the length of the messages.

Three details matter here.

The window is shared, not split. There is no separate “input budget” and “output budget”. Input and output draw from the same pool. If a model has a 128k window and your prompt is 120k tokens, you have at most 8k tokens of answer left — assuming the provider even allows an output that long, since most models cap single-response output well below the full window regardless of how much room is free.

Reserved output counts before it’s generated. If you set max_tokens: 4096, many providers validate input + max_tokens ≤ context_limit up front and reject the call. A request that would have squeaked through with a 500-token summary gets rejected because you asked for headroom you were never going to use. On newer reasoning-capable models the parameter is often max_completion_tokens, and internal reasoning tokens are billed and counted as output — so a request can consume far more of the budget than the visible answer suggests.

This is not the same as a rate limit. A 429 with “rate limit reached for tokens per minute” is an org-level throughput ceiling; the same request will often succeed 60 seconds later. A context length error will fail identically forever. If retry logic in your workflow is quietly re-firing a context error, you’re just burning attempts. (In n8n specifically, a node that fails on every retry with an unchanged payload is worth checking against the same class of silent-failure causes as a webhook that stops firing in production: something upstream is failing the same way every time, not intermittently.)

Where a transcript’s tokens actually hide

Speech is not token-dense. A rough English rule of thumb is ~0.75 words per token, or ~4 characters per token, for the common OpenAI tokenizers. Human speech runs somewhere around 130–150 words per minute in conversation. Put those together and an hour of talking is on the order of 8,000–9,000 words, or roughly 11,000–12,000 tokens of pure speech. That fits inside any modern window with room to spare.

So why do transcripts blow up? Because almost nobody sends pure speech. They send the WebVTT or SRT file that came out of Whisper, or a diarized JSON export.

A worked example, with the assumptions stated. Take a three-hour interview transcribed to WebVTT. Assume ~140 wpm, so ~25,000 words of speech, which at 0.75 words/token is around 33,000 tokens of content. Now assume the transcriber emitted cues averaging 8 words each — that’s 3,125 cues. Each cue carries a timestamp line like 00:12:34.500 --> 00:12:38.200, and because digits and colons tokenize badly, that line alone lands in the low double digits of tokens. Add a cue index, blank lines, and a SPEAKER_01: prefix, and you’re at roughly 20 tokens of scaffolding per cue.

3,125 cues × 20 tokens = 62,500 tokens of overhead wrapped around 33,000 tokens of speech. The file is ~95,000 tokens, and nearly two thirds of it is punctuation the model does not need to write a summary.

Those numbers are illustrative — your cue length, tokenizer and diarization format will move them, sometimes a lot. The point is the ratio, not the figure. Run the count on your own file before you assume anything.

The reductions available from format alone, in rough order of impact:

  • Drop timestamps entirely if the summary doesn’t need to cite times. If it does, keep one timestamp per topic block or per minute rather than per cue.
  • Merge consecutive cues from the same speaker into paragraphs. This kills the per-cue scaffolding and, as a side effect, gives the model better-connected prose.
  • Collapse repeated speaker labels. SPEAKER_01: on every line becomes one label per turn.
  • Strip filler and stutters (um, uh, repeated false starts) if the transcriber preserved them. Modest gain, and it can distort verbatim accuracy, so it depends on what the summary is for.

Count before you send

Guessing token counts from file size is how you end up debugging in production. Every major provider exposes a way to count exactly:

  • OpenAI: tiktoken, using the encoding that matches the model (o200k_base for the GPT-4o generation, cl100k_base for older GPT-4 and GPT-3.5-turbo). The wrong encoding gives a wrong number.
  • Anthropic: a dedicated token-counting endpoint that accepts the same message structure you’d send to the model, so the count includes system prompt and message overhead.
  • Local/open models: the tokenizer shipped with the model, usually via transformers.

Message formatting itself costs tokens — role markers, separators, tool definitions if you’re passing any. On a 100k-token request that’s noise; on a request sitting 200 tokens under the ceiling it’s the difference between a 200 and a 400.

A practical guard in a pipeline: compute input_tokens, subtract from the model’s window, subtract your desired output length, and branch on whether the remainder is positive. That turns an unpredictable API failure into a deterministic routing decision you control.

The chunking mechanisms, and what each one trades away

When the transcript genuinely doesn’t fit, you summarize in passes. There are three standard patterns and they fail in different ways.

Map-reduce

Split the transcript into N chunks, summarize each independently (the “map”), then summarize the collection of summaries (the “reduce”).

Behaviour: chunks are processed in parallel, so wall-clock time stays flat as the transcript grows. Cost scales roughly linearly with length. Failure mode: each chunk is summarized blind. A decision made in chunk 7 that reverses something agreed in chunk 2 gets reported as two separate facts, because no single call ever saw both. Callbacks, running jokes, and “as we said earlier” references dissolve. Where it fits: long content with weak narrative dependency — panel discussions, conference recordings, batches of unrelated support calls.

Refine (rolling summary)

Summarize chunk 1. Feed that summary plus chunk 2 to the model, asking it to update the summary. Repeat through the transcript.

Behaviour: the state carried forward means later chunks are interpreted in light of earlier ones. Contradictions and resolutions survive. Failure mode: strictly sequential, so latency scales with chunk count. And the summary degrades as it travels — details from chunk 1 get compressed again at every step, so the beginning of a long transcript fades while the end stays sharp. This is the mechanism behind “the model forgot the first half”. Where it fits: negotiations, therapy or coaching sessions, anything where the ending only makes sense given the beginning.

Hierarchical / tree

Map-reduce with more than two levels: summarize chunks, summarize groups of summaries, then summarize those. Useful when even the concatenated chunk summaries exceed the window — which happens once you’re processing dozens of hours rather than one long file.

Choosing chunk boundaries

Splitting on a fixed character count cuts sentences and speaker turns in half, and a chunk that starts mid-answer produces a summary that misattributes it. Better boundaries, in order of preference: speaker turn changes, paragraph breaks, then sentence breaks. Add a small overlap — repeating the last few hundred tokens of each chunk at the start of the next — so a point spanning a boundary appears intact in at least one chunk. Overlap costs tokens proportional to overlap × chunk_count, so a 500-token overlap across 20 chunks is 10,000 extra tokens you’re paying for.

One structural detail that quietly improves every chunked summary: put an explicit instruction in each map prompt saying the chunk is an excerpt from a longer recording and the model should not editorialize about what’s missing. Without it, models frequently open with “This excerpt appears to be part of a larger conversation…”, and you spend the reduce pass cleaning that up.

When a bigger context window is and isn’t the answer

Windows have grown enormously — Anthropic’s Claude models advertise 200k tokens, Google’s Gemini Pro line has offered 1M+, and OpenAI’s flagships sit at 128k and above. Vendors change these numbers frequently, so treat any figure here as a pointer to check their pricing page rather than a fact to build on.

Fitting everything in one call removes the coordination problem entirely, and for many transcripts that’s genuinely the cleanest path. But three things change when you do:

Cost is linear in input tokens. A single 400k-token call costs the same as chunking that content, minus the reduce pass — the saving is real but smaller than it feels. Prompt caching (available on several providers) cuts the cost of re-sending the same transcript across multiple queries substantially; it does not change the context ceiling, which is enforced on the uncached token count.

Latency is not free. Time-to-first-token grows with prompt length. A 300k-token summarization request can take long enough to trip default HTTP timeouts in orchestration tools, which then surfaces as a generic connection error rather than anything informative.

Retrieval quality degrades unevenly across a long context. The “lost in the middle” effect — models attending more reliably to information at the start and end of a long prompt than to the middle — is a documented behaviour in long-context research and it hasn’t disappeared with larger windows. A summary of a 500k-token input can be confidently wrong about the middle third. Chunking sidesteps this by guaranteeing every section gets focused attention, which is one reason map-reduce sometimes beats a single giant call even when the giant call would fit.

A diagnostic order that resolves most cases

  1. Confirm it’s a context error, not a 429. Different problem, different fix.
  2. Count the tokens locally with the model’s own tokenizer, including the system prompt.
  3. Check what you reserved. max_tokens / max_completion_tokens is subtracted from your available room. A summary rarely needs 4k.
  4. Strip formatting. Timestamps, cue indices, repeated speaker labels. On subtitle files this alone frequently halves the count.
  5. Compare the cleaned count to the window. If it fits with output headroom, you’re done.
  6. If it doesn’t, chunk — map-reduce for independent content, refine for content with narrative dependency, splitting on speaker turns with a modest overlap.
  7. If chunk summaries themselves overflow the reduce call, add a level and go hierarchical.

Building this in an automation platform doesn’t change the mechanics, only where the arithmetic lives. In n8n, that’s typically a Code node computing the token count, a Split In Batches loop over chunks, and an IF node routing anything oversized down the chunking branch rather than firing it at the API and hoping. If you’re assembling that kind of pipeline for the first time, the tooling landscape in 5 AI tools that quietly replaced half my workflow covers what tends to sit around it.

The underlying rule doesn’t move: the window is finite and shared between what you send and what you ask for. Every technique above is a different way of making one of those two numbers smaller.