Promptrift

AI tools

How to Turn a 90-Minute YouTube Video Into a Blog Post With AI

The exact pipeline for turning a 90-minute YouTube video into a blog post with AI: pulling the transcript, chunking it so no model chokes, and prompting for structure.

A YouTube video transcript being split into chunks and reassembled into a blog post draft

Turning a 90-minute YouTube video into a blog post with AI comes down to three mechanical steps: pull a text transcript out of the video, feed that transcript to a language model in pieces it can actually digest, and prompt the model to reorganize the raw, spoken transcript into an article shape — headings, an intro that answers the topic up front, pull quotes, takeaways — instead of a wall of “so basically what I wanted to talk about today is.”

A 90-minute video at a normal conversational pace produces roughly 10,000 to 14,000 words of transcript. That’s already longer than most blog posts on its own, and it’s too much to paste into a single prompt without either splitting it up or using a model with a large enough context window to hold the whole thing at once. Everything below is about handling that length problem correctly, because that’s the part that actually breaks when people try to automate this.

Estimated transcript length for a 90-minute video, by speaking pace
110 wpm (slow) 9,900 words 130 wpm (average) 11,700 words 160 wpm (fast) 14,400 words
Worked example: 90 minutes × words per minute

Step 1: Get the transcript out of the video

There are two ways to get text out of a YouTube video, and which one you need depends on whether decent captions already exist.

If the channel uploaded manual captions or the auto-captions are clean, you can pull them directly — either by downloading the .srt/.vtt file from YouTube’s own caption menu, or with one of the small transcript-fetching libraries that wrap YouTube’s caption API. This is the fast path: no audio processing, and you keep timestamps for free, which matters later if you want to link back to specific moments in the video.

If there are no captions, or the auto-generated ones are the usual mess — no punctuation, misheard product names, no speaker separation on a two-person podcast — the more reliable route is running the audio through a dedicated speech-to-text model (Whisper-family models are the common choice) and generating your own transcript. It costs more time and, if you’re using a paid API, a small amount of money per minute of audio, but the output is far more usable as raw material than scraped auto-captions.

Either way, the output you want at the end of this step is plain text with timestamps attached to reasonably sized segments — not just one giant unbroken paragraph.

Step 2: Deal with the length problem before it deals with you

This is the step people skip, and it’s why so many “AI blog from YouTube” attempts come back garbled or cut off halfway through. A 12,000-word transcript is a lot of tokens. Depending on the model and how much of its context window is already spent on your system prompt and instructions, dumping the whole transcript in one shot can get truncated, ignored past a certain point, or summarized so aggressively that the back half of the video disappears from the draft entirely.

The fix is chunking: splitting the transcript into pieces small enough for the model to actually reason over each one, summarizing or restructuring each chunk, and then combining those pieces into a single outline before writing the final draft. There are two practical ways to decide where the chunks break:

  • Fixed-size chunks — split every N words (2,000–3,000 is a common range), regardless of what’s being said at that boundary. Simple to automate, but you can cut a sentence or a train of thought in half.
  • Chapter- or topic-based chunks — if the video has YouTube chapter markers, use those as natural boundaries. If it doesn’t, you can ask a model to scan the transcript once and flag topic-shift points, then chunk on those. This produces cleaner summaries per chunk because each one covers one coherent idea instead of an arbitrary word count.

The number of chunks you end up with scales directly with how small you cut them, which also determines how many separate model calls the pipeline needs:

Chunks needed for a 12,000-word transcript, by chunk size
500 words 24 chunks 1500 words 8 chunks 2500 words 5 chunks 4000 words 3 chunks
Worked example: transcript length ÷ chunk size, rounded up

Smaller chunks are safer against the model losing track of a section, but they mean more API calls and a harder job stitching the summaries back into one coherent outline without repetition. Larger chunks are cheaper and faster but risk hitting the same truncation problem you were trying to avoid, just at a slightly bigger scale. The mechanics of picking chunk size without losing the thread of a long document are covered in more depth in Context Length Exceeded? How to Summarize a Long Transcript Without Breaking the Model — worth reading before you automate this part, because the failure mode there (the model quietly drops the middle of the document) is the same one that shows up here.

Step 3: Turn the transcript into a blog structure, not a summary

Chunked summaries alone don’t make a blog post — they make a shorter transcript. The step that actually produces an article is a second prompt pass that takes the combined outline from step 2 and asks for a specific structure: an opening that states the answer or main point in the first two paragraphs, ## section headings that match how someone would search for the topic, and a closing that doesn’t just repeat the intro.

Splitting this into two separate prompts — one that extracts a clean outline from the transcript, a second that expands that outline into full prose — tends to produce more coherent drafts than asking one prompt to do both at once, because each prompt has a narrower job and less context to juggle. That’s the same logic behind running two models (or two prompts) in sequence instead of forcing one call to do everything, which is broken down in How to Chain Two AI Models in One Workflow.

A workable second-pass prompt looks something like: “Here is an outline extracted from a 90-minute video transcript. Write a blog post section for each outline point, in the speaker’s own claims and examples only — do not add statistics, names, or facts that aren’t in the outline. Use ## for each section heading.” That last instruction matters more than it sounds like it should.

What the AI won’t fix on its own

Two failure patterns show up consistently in transcript-to-blog drafts, and neither gets caught by a better prompt alone:

  • Invented specifics. Models asked to “expand” a bullet point will sometimes add a statistic, a product name, or a comparison that was never said in the video, just to make the paragraph sound more complete. Every number or claim in the final draft needs to be checked against the actual transcript before publishing.
  • Structure that doesn’t match search intent. A model summarizing a rambling 90-minute conversation will often produce headings that follow the order the speaker talked, not the order a reader searching for the topic would want the information. That reordering is usually a manual edit, not something worth prompting harder for.

Automating the whole pipeline

Once the steps above work reliably by hand, they map cleanly onto a workflow trigger: new video ID comes in → fetch or generate transcript → chunk → summarize each chunk → build outline → expand to draft → drop the draft somewhere for review before it goes anywhere near a CMS. That “review before it ships” step is the one worth keeping manual, for the fact-checking reason above — the same pattern of piping an AI-generated summary to a review channel instead of publishing it directly shows up in Build an AI Email Triage Workflow in N8n With GPT and in How to Auto-Send AI Zoom Meeting Summaries Straight to Slack, both of which are the same “transcribe, summarize, hand to a human” shape applied to different source material.

The mechanics don’t change much between a Zoom call, an email inbox, and a 90-minute YouTube video — what changes is the transcript length, and that’s exactly the part that needs chunking done right before anything else in the pipeline is worth building.