Automation
Build an AI Email Triage Workflow in N8n With GPT (Node by Node)
How to build an email triage automation in n8n with GPT: the exact node chain, the classification prompt, and how to route and label mail without a human reading it first.

An email triage automation in n8n with GPT works by pulling unread messages with a trigger node, sending the subject and body to an OpenAI node with a classification prompt, parsing the JSON it returns, and routing the message through a Switch node to the right label, folder, or downstream action. The whole thing is five to seven nodes. No custom code required, though a Code node cleans things up if you want strict formatting.
The part that actually determines whether this works long-term isn’t the trigger or the labels — it’s the classification prompt and how you force the model to return something your Switch node can parse without falling over. Get that wrong and you’ll spend more time debugging malformed JSON than you saved by not reading your inbox manually.
What the workflow actually does
The shape is always the same regardless of which email provider you use:
- Trigger — new or unread email arrives
- Fetch — pull subject, sender, body, maybe attachments metadata
- Classify — send it to GPT with a prompt that returns a category
- Parse — turn GPT’s text response into a usable JSON object
- Route — Switch node sends it down one of several branches
- Act — apply a label, move to a folder, forward, or create a task
Everything downstream of step 3 depends entirely on step 3 being reliable. That’s where most of these workflows break in production, so it’s worth spending real time there instead of rushing to the fun part (auto-replies, task creation, Slack pings).
Node 1: the trigger
If you’re on Gmail, use the Gmail Trigger node set to poll on a schedule (every 1-5 minutes is typical) filtered to is:unread. If you’re on Outlook or a generic IMAP account, n8n has native Microsoft Outlook Trigger and Email Trigger (IMAP) nodes that work the same way conceptually — poll, return new messages, pass them downstream.
One thing worth knowing before you build the rest of the workflow: polling triggers behave differently in production than they do when you’re clicking “Execute Workflow” manually in the editor. If you later add a webhook-based step anywhere in the chain (for example, triggering triage from a form submission instead of polling), the workflow has to be active and saved for the production URL to listen — the same gotcha covered in N8n Webhook Not Triggering in Production Mode: 7 Causes and Fixes.
Node 2: fetching the data you’ll classify
If your trigger only returns a message ID, add a Gmail — Get (or equivalent) node right after it to pull the full message: subject, from, plain-text body, and date. Strip HTML if you’re pulling the HTML body — feeding raw HTML tags to GPT wastes tokens and adds noise the model doesn’t need to classify intent.
A Set node here to normalize field names (emailSubject, emailBody, emailSender) makes the rest of the workflow easier to read and easier to hand off to someone else later.
Node 3: the GPT classification node
This is the core of the automation. Add an OpenAI node (or HTTP Request if you’re calling a different provider) configured as a chat completion. The system prompt is what does the work:
You are an email triage classifier. Given an email subject and body, return ONLY a JSON object with this exact shape:
{ “category”: “urgent” | “action_needed” | “fyi” | “newsletter” | “spam”, “confidence”: 0.0-1.0, “reason”: “one short sentence” }
Rules:
- “urgent” = requires a response within the same day (deadlines, angry customers, anything with words like “ASAP”, “today”, “ending soon”).
- “action_needed” = requires a response but not urgently.
- “fyi” = informational, no response needed, but not a newsletter.
- “newsletter” = bulk/marketing content, digests, automated updates.
- “spam” = unsolicited sales pitches, phishing patterns, irrelevant offers.
Return valid JSON only. No markdown, no explanation, no code fences.
Set the model’s temperature to 0 — you want consistent, repeatable categorization, not creative variation. If your provider supports a JSON response mode (OpenAI’s response_format: { type: "json_object" } does), turn it on. It doesn’t guarantee your schema, but it does guarantee the output parses as JSON, which eliminates one entire category of failure.
Getting the model to actually stick to that schema every single time, including on edge cases like empty bodies or emails in other languages, is its own problem — the same one covered in How to Write a System Prompt That Forces JSON-Only Output (With Test Cases). Worth reading before you wire this into production, because the failure mode isn’t “the model refuses” — it’s “the model adds one sentence of preamble before the JSON,” which breaks a naive parser silently.
Node 4: parsing the response
Add a Code node (JavaScript) right after the OpenAI node:
const raw = $input.item.json.message.content;
let parsed;
try {
parsed = JSON.parse(raw);
} catch (e) {
parsed = { category: "fyi", confidence: 0, reason: "parse_failed" };
}
return { json: parsed };
That fallback matters. Without it, one malformed response from GPT stops the entire workflow execution instead of just mis-routing a single email. Defaulting to `"fyi"` on parse failure means a bad classification lands in a low-stakes bucket instead of silently vanishing or crashing the run.
## Node 5: routing with a Switch node
Add a **Switch** node reading `{{ $json.category }}`, with one output per category: `urgent`, `action_needed`, `fyi`, `newsletter`, `spam`. Each branch does something different:
- **urgent** → Gmail "Add Label" (e.g. `Triage/Urgent`) + a Slack or Telegram notification node
- **action_needed** → label `Triage/Action` + optionally create a task in Notion, Todoist, or wherever the team tracks work
- **fyi** → label only, no notification
- **newsletter** → label `Triage/Digest`, optionally skip inbox (Gmail's `removeLabelIds: INBOX`)
- **spam** → move to a review folder rather than deleting outright — auto-deleting based on a single model call is how you eventually lose a real email that got misclassified
That last point is worth being deliberate about. Confidence score from step 3 is useful here: if `confidence < 0.6`, route to a manual review label regardless of category, instead of trusting a low-confidence guess.
## Handling edge cases
A few things that surface once this runs against a real inbox instead of a handful of test emails:
- **Long threads.** Gmail bodies on long threads include the entire quoted history. Truncate to the first 2,000-3,000 characters before sending to GPT, or you'll burn tokens on quoted text the model doesn't need.
- **Non-English emails.** GPT handles multilingual classification fine without extra instructions, but if your team wants the `reason` field in a specific language, say so explicitly in the system prompt — it defaults to whatever language the email is in.
- **Attachments.** The classification node doesn't need attachment content, just the fact that one exists (`hasAttachment: true/false`) if that should influence urgency.
- **Automated sender addresses.** `noreply@`, `notifications@`, and similar patterns are a strong signal for `newsletter`/`fyi` that's cheaper to catch with a simple **IF** node before the GPT call, saving you an API call on obviously automated mail.
## Testing before it touches a real inbox
Don't point this at a live inbox on day one. Pull 15-20 real emails covering each category you care about — including a couple of intentionally ambiguous ones — and run them through the classification prompt as a batch before wiring up the routing and labeling nodes. The approach in [How to Test One Prompt Against 20 Inputs at Once (Cheap Prompt Regression)](/posts/test-one-prompt-20-inputs) applies directly here: it catches prompt drift (a category the model consistently confuses) before it's silently mislabeling live mail.
## Common failure points
Most breakages in this kind of workflow trace back to one of three things: the OpenAI node timing out on unusually long emails, the parse step failing on unexpected model output, or the trigger not firing because the workflow was saved but not activated. None of these are exotic — they're the standard n8n gotchas that show up in any node chain with an external API in the middle, which is exactly why the parsing fallback and the confidence threshold matter more than any clever routing logic you add on top.