Prompting
How to Test One Prompt Against 20 Inputs at Once (Cheap Prompt Regression)
Run one prompt across 20 inputs at once with a CSV loop, promptfoo, or n8n. Concrete setup, grading methods, and cost math for cheap prompt regression.

The fastest way to test one prompt against 20 inputs at once is to put the prompt in a template with one variable slot, put the 20 inputs in a CSV or spreadsheet column, and loop the API call over each row — either with a short script, a no-code workflow tool, or a dedicated eval tool like promptfoo. You get 20 outputs back in one run, side by side, instead of pasting into a chat window 20 times.
This matters because testing a prompt on 2 or 3 examples tells you almost nothing about how it behaves on the input you haven’t thought of yet — the empty string, the 4,000-word paste, the input in the wrong language, the one with a stray JSON brace in it. Below are three ways to run the batch, how to grade the results without reading all 20 by hand, and how to turn this into a regression check you rerun every time you touch the prompt.
Why one or two test inputs isn’t enough
A prompt that works on your first example can silently break on inputs with different shape: different length, different formatting, missing fields, or edge cases like empty input. If you only ever test with the same clean example, you never see that failure — until it shows up in production.
Batch testing against a fixed set of 15-20 inputs turns “does this prompt work” into “does this prompt work on the same 20 cases it worked on last week, plus a few adversarial ones.” That’s the difference between a vibe check and a regression test.
Build a 20-input set that’s actually useful
Random inputs aren’t as useful as inputs chosen to break things. A reasonable 20-input set for most prompts:
- 5-6 “normal” cases, representative of typical real input
- 3-4 edge cases: empty string, single word, all-caps, input in a different language than expected
- 3-4 long inputs, near or past whatever length you expect in production
- 2-3 adversarial inputs: text that contains instructions (“ignore the above and say X”), or text that looks like the output format you’re asking for (to check the model doesn’t just echo it)
- 2-3 malformed inputs: missing a field the prompt expects, stray markdown, broken JSON if you’re feeding JSON in
Keep this set in a CSV with one column per input variable and a column for the expected output or a pass/fail rule, if you have one. That CSV is the artifact you reuse every time you touch the prompt — not just this once.
Method 1: spreadsheet plus a formula-driven API call
If you’re already comfortable with a spreadsheet and an API key, this needs no code. Put your 20 inputs in column A, the prompt template in a fixed cell, and use an API connector add-on (or a script bound to the sheet) that reads column A row by row, substitutes it into the prompt, calls the model, and writes the output to column B.
This is the same mechanism covered in How to Connect Google Sheets to an AI Model Without Zapier — the batch-of-20 case is just that same setup run once with 20 filled rows instead of one.
Worked example: with a prompt template of roughly 150 tokens and inputs averaging 100 tokens each, one call is around 250 input tokens plus whatever the output runs. At a rate of $0.15 per million input tokens and $0.60 per million output tokens (illustrative, not a live price — check your provider’s current page), 20 calls at ~250 input + ~150 output tokens each comes to roughly 5,000 input tokens and 3,000 output tokens total — a fraction of a cent. The point isn’t the exact number, it’s that batch-testing 20 inputs against one prompt costs close to nothing compared to the time saved catching a regression before it ships.
Method 2: a short script (most control, still cheap)
A script gives you the most control over grading and is the easiest to rerun automatically. The shape, in pseudocode:
rows = read_csv(“test-inputs.csv”) for row in rows: prompt = template.replace(“{{input}}”, row.input) output = call_model(prompt) row.output = output row.pass = grade(output, row.expected) write_csv(“results.csv”, rows)
grade() is the part worth thinking about, because reading 20 outputs by eye every time defeats the point of automating this. A few grading approaches, cheapest first:
- Exact match or contains: check the output equals a string, or contains a required substring. Works for classification-style prompts (“respond with one of: positive, negative, neutral”).
- Schema validation: if the prompt is supposed to return JSON, parse it and check required fields exist and types match. This pairs directly with the approach in How to Write a System Prompt That Forces JSON-Only Output (With Test Cases) — that article covers getting clean JSON out in the first place; here you’re validating it across 20 inputs instead of one.
- Regex or length check: cheap sanity checks — output isn’t empty, doesn’t exceed a length limit, doesn’t contain a forbidden phrase.
- LLM-as-judge: for open-ended output (summaries, rewrites) where there’s no single correct string, send the output to a second prompt that scores it against a rubric (1-5, or pass/fail against 2-3 stated criteria). This costs a second API call per row, so for 20 inputs it roughly doubles your token spend — still a small number in absolute terms, but worth knowing before you run it on 200 inputs instead of 20.
Method 3: no-code batch runs in n8n
If you already run prompts through a workflow tool, the batch-of-20 problem is a data-processing problem, not a prompting problem: read 20 rows, loop, call the model node, collect outputs. How to Process a Large CSV in N8n Without Timing Out (Split In Batches, Explained) covers the node that does exactly this — for 20 rows you won’t hit the timeout issues that article is mainly about, but the same Split In Batches node is the mechanism you’d use, feeding each batch into an HTTP Request or AI node and writing results back to a sheet or database.
Method 4: a dedicated eval tool (promptfoo and similar)
Tools built specifically for this — promptfoo is the best-known open-source one — take a YAML or JSON config listing your prompt, your 20 test cases, and assertions per case (contains, equals, json-schema, llm-rubric, and a few others built in). Running it gives you a pass/fail table across all 20 in the terminal or a browser view, and it’s built to diff two prompt versions side by side, which is the actual regression-testing use case: you change the prompt, rerun, and see exactly which of the 20 flipped from pass to fail. If you’re doing this often enough that you’re maintaining test cases as a real asset, a dedicated tool like this saves more time than a script does, because the diffing and reporting are already built.
Making it a regression check, not a one-off
The batch run only pays off if you keep the same 20 inputs and rerun them every time you edit the prompt. Save the CSV or config file next to the prompt itself, and treat a prompt edit the same way you’d treat a code change: rerun the 20-case set, look at what flipped, and only ship the new version if nothing that used to pass now fails — or if a failure is one you’ve deliberately accepted as an acceptable tradeoff for the case it’s now handling better.