Promptrift

Automation

How to Process a Large CSV in N8n Without Timing Out (Split In Batches, Explained)

A large CSV in n8n times out for three different reasons, and Split In Batches only fixes one of them. Here's how the Loop Over Items node actually works, how to size batches, and how to keep a 180k-row file from killing the execution.

A CSV file being split into numbered chunks feeding a looping workflow diagram

The short version: you split the file into chunks with the Loop Over Items node (the node formerly and still internally called splitInBatches), wire the processing nodes back into that node’s input so it iterates, and take the final result off the done output. The node holds all incoming items in memory, releases n of them per iteration through the loop output, and only fires done when the queue is empty. That turns one enormous run of a node into many small ones, which is what stops per-request HTTP timeouts on the APIs you’re calling.

But — and this is the part that gets skipped — batching does not stop the execution timeout, and it does not stop out-of-memory crashes. There are three separate clocks running when you process a big CSV in n8n, and Split In Batches only touches one of them. If your run dies at exactly 100 seconds, or at exactly 60, or the worker restarts with no error at all, the batch size isn’t your problem. Below is what each mechanism does, so you can tell which one is firing.

The three things that “time out” in n8n

1. The per-node HTTP request. Your HTTP Request node sends 12,000 rows to an API in one call, the API takes 40 seconds and then gives up. Or n8n’s own request timeout fires. This is the one batching genuinely fixes — 12,000 rows at 200 per call is 60 requests that each finish in under a second.

2. The workflow execution timeout. This is a wall clock on the whole run. On self-hosted n8n it’s governed by EXECUTIONS_TIMEOUT (default -1, meaning no timeout) and capped by EXECUTIONS_TIMEOUT_MAX. A workflow can also carry its own timeout in Workflow Settings, which overrides the default. n8n Cloud plans apply their own execution limits — those change with pricing tiers, so check the current plan page rather than trusting a number from a forum post. The key thing: batching makes this worse, not better. Sixty sequential API calls take longer end to end than one bulk call. If your run is being killed at a fixed number of minutes, smaller batches move you further from safety.

3. The HTTP layer in front of n8n. If the workflow is triggered by a webhook set to respond “When Last Node Finishes”, the caller’s connection has to stay open for the entire run. nginx’s proxy_read_timeout defaults to 60 seconds. Cloudflare returns a 524 after 100 seconds on the standard plans. Neither of those cares that n8n is happily still working — the response is already dead. If you’re getting a 504 or 524 while the n8n UI shows the execution as successful, this is your layer, and no batch size will help. Related failure modes are covered in N8n Webhook Not Triggering in Production Mode.

Symptom-to-cause, roughly: a node-level red error with a status code is #1. The execution stopping cleanly at a suspiciously round time with a “timeout” message is #2. A gateway error at the client while n8n keeps going is #3. And an execution stuck on “running” forever, or a worker that silently restarts, is usually memory, which is its own section.

What the loop node actually does

Loop Over Items (node type n8n-nodes-base.splitInBatches, v3) has one input and two outputs:

  • Output 0 — done: fires once, after the last batch. In v3 it emits everything that came back into the node’s input across all iterations, accumulated.
  • Output 1 — loop: fires once per iteration with batchSize items.

The wiring that trips people up: the last node of your processing chain connects back into the Loop Over Items input. That back-edge is the loop. Without it, the node runs exactly once, hands you the first batch, and stops — which looks like “it only processed 100 rows”.

Two options matter:

  • Reset: if this is on while you have a back-edge wired, the node restarts its internal counter every time items arrive, and you get an infinite loop. It exists for nested-loop scenarios. Leave it off unless you know why you turned it on.
  • Batch size: default 1, which is almost never what you want for a CSV.

Inside the loop, $runIndex gives you the current iteration (0-based), which is what you use for logging progress or for offset math.

Sizing the batch: the tradeoff nobody spells out

Batch size 1 feels safe and is usually the worst choice. Each iteration walks every node in the loop, and n8n tracks execution data per node per run. A 50,000-row CSV at batch size 1 is 50,000 iterations × however many nodes are in the loop. The overhead is not the API — it’s n8n’s own bookkeeping.

A worked example, with stated assumptions: say the loop has 4 nodes, per-iteration overhead is roughly 30 ms, and the API call is 400 ms regardless of whether you send 1 row or 200. At batch size 1, 50,000 rows = 50,000 × 430 ms ≈ 6 hours. At batch size 200, that’s 250 iterations × (30 ms + 400 ms) ≈ under 2 minutes. Those overhead numbers are illustrative — yours depend on your database, your nodes and whether you’re in queue mode — but the shape of the curve is real, and it’s why “make the batches smaller” often makes timeouts more likely, not less.

The ceiling on batch size is whatever your slowest constraint is: the API’s payload limit, its rate limit, and the memory cost of holding one batch’s worth of enriched data. Practical starting range for API-bound work is 50–500. If you’re calling an LLM per row, the batch size is effectively dictated by the context window and by the fact that you’re paying per token — different problem, and chunking strategy for that is in Context Length Exceeded? How to Summarize a Long Transcript.

The memory problem batching doesn’t solve

Here’s the mechanism that surprises people: Loop Over Items receives the whole dataset before it releases the first batch. It’s a buffer, not a stream. So if the node before it is an Extract From File node that just parsed a 400 MB CSV into JSON items, all of that is in Node’s heap before batch one exists.

And it’s worse than the file size suggests. A CSV row stores column names once, in the header. The parsed JSON item repeats every key on every row. A 30-column CSV with short values can expand 4–8× going from disk to n8n items. Then n8n keeps per-node execution data, so a chain of six nodes can hold several copies of that expanded payload.

What that looks like when it breaks: JavaScript heap out of memory in the container logs, or the process just dying and the execution sitting at “running” forever with nothing useful in the UI. It is not a timeout, even though it presents like one.

Things that move this needle:

  • NODE_OPTIONS=--max-old-space-size=4096 raises Node’s heap ceiling (in MB). This buys headroom; it doesn’t change the expansion factor.
  • N8N_PAYLOAD_SIZE_MAX defaults to 16 (MB) and limits incoming webhook/request payloads. If you’re POSTing the CSV to a webhook, this is what rejects it — separate from heap.
  • Don’t carry columns you don’t need. A Code node or Edit Fields node right after parsing, dropping 20 of 30 columns, cuts everything downstream proportionally. This is the highest-leverage change in most large-CSV workflows and costs nothing.
  • Split the file before n8n sees it. split -l 5000 big.csv chunk_ on the source machine, or reading from S3 with byte ranges, or — best of all — paginating the source system instead of exporting a CSV at all. The CSV parser nodes have row/range options that vary by node version; check what yours exposes before designing around it.

Making a long run survive: sub-workflows and static data

If the total work genuinely exceeds your execution timeout, the fix is architectural, not a bigger batch.

Sub-workflow per chunk. The Execute Sub-workflow node has a Wait for Sub-workflow Completion toggle. With it off, the parent fires each chunk and moves on immediately — the parent finishes in seconds, and each chunk becomes its own execution with its own timeout budget and its own memory that gets freed when it ends. The tradeoff is real: you lose the return values and you have to handle failures per chunk yourself, because the parent won’t know.

Static data as a cursor. $getWorkflowStaticData('global') persists between executions of the same workflow. Storing { lastRow: 12000 } there lets a scheduled workflow pick up where the last run stopped, so a 500k-row job becomes 50 runs of 10k. Note that static data only persists for production executions, not manual test runs from the editor — which is why it looks broken when you’re testing it.

The Wait node’s 65-second rule. A Wait node set below roughly 65 seconds holds the execution in memory. Above that threshold, n8n saves the execution to the database and resumes it later, freeing the process. That distinction matters if you’re using waits for rate limiting inside a big loop: many short waits keep everything resident, which is fine for pacing but does nothing for memory pressure.

Respond immediately. If a webhook starts the job, set the response mode to respond immediately (or use a Respond to Webhook node early) and return a job ID. The caller gets a 200 in milliseconds and the reverse-proxy clock stops mattering entirely.

Config that changes how long a big loop takes

Two settings quietly dominate loop performance on self-hosted instances:

  • Save execution progress (EXECUTIONS_DATA_SAVE_ON_PROGRESS, default false). When on, n8n writes execution data to the database after every node. In a 250-iteration loop with 5 nodes, that’s 1,250 database writes. Excellent for debugging a workflow that crashes mid-run; expensive as a permanent setting.
  • EXECUTIONS_DATA_SAVE_ON_SUCCESS. Keeping full data for every successful run of a high-volume workflow grows the executions table fast, and a bloated table slows down every subsequent write. EXECUTIONS_DATA_PRUNE and the max-age settings exist for this.

On queue mode with multiple workers, chunk-as-sub-workflow also buys you real parallelism, bounded by N8N_CONCURRENCY_PRODUCTION_LIMIT. Sequential batching inside one execution never uses more than one worker, no matter how many you’re running.

The debugging order

When a large-CSV workflow dies, the useful sequence is: check whether the execution itself errored or the caller errored (that separates layer 3 from layers 1 and 2). If the execution errored, check whether it stopped at a consistent time — a consistent time means a timeout, an inconsistent one means memory or an upstream API. Then check the container logs for a heap error before touching batch size, because heap errors and timeouts look identical in the n8n UI and have opposite fixes: for a timeout you make batches bigger to reduce overhead, for memory you make the dataset smaller before it ever reaches the loop.

Version details on these nodes and env vars move quickly — the splitInBatches node has already changed output semantics between major versions. Check the docs for the version you’re actually running before wiring anything permanent.