A field-test runner is the script that fires your LLM calls at scale — dozens of prompts, hundreds of configs, thousands of runs — and collects the results. Hardening one means making it survive timeouts, hangs, and runaway cost without losing a single sweep. I learned this the expensive way.
Last month I ran a benchmark across 4,768 LLM calls. Zero lost sweeps. Zero blown budget. But the first version of that runner? It died at run 312 and took 40 minutes of results with it. This is the story of what broke, and the four fixes that held.
What does a "sweep" actually mean, and why do they get lost?
A sweep is one full pass through your test matrix. Say you're comparing 8 prompts across 4 models across 3 temperature settings. That's 96 calls per sweep. Run 50 sweeps and you're at 4,800 calls.
Related: I've explored this before in ai writing assistant free.
Losing a sweep means the runner crashed, hung, or hit a rate limit mid-pass and you have no clean way to resume. You either re-run everything (wasting money) or you patch together partial results (corrupting your data).
I've watched both happen. Neither is fun.
Related: This connects to what I wrote about prompt engineering for content creators.
Here's the thing most tutorials skip. The hard part isn't the API call. It's everything around the API call — the retry logic, the state tracking, the cost ceiling. That's where runners actually break.
3 failure modes that killed my first 4,768-run attempt
I logged every crash. Three patterns accounted for 91% of them.
Related: For more on this, see ai for writing a book.
- Silent hangs. A request would open a connection and just... sit there. No timeout, no error, no response. The runner blocked forever on a single call while 4,767 others waited in queue.
- Rate-limit cascades. One 429 triggered a retry, which triggered three more 429s, which triggered a thundering herd of retries. Cost spiked 6x in under two minutes.
- State loss on crash. The runner held all results in memory and wrote to disk only at the end. A crash at run 4,700 meant 4,699 calls were gone.
That third one hurt the most. I had the data. I just hadn't saved it.
Fix #1: Hard timeouts on every call, no exceptions
Every LLM request needs two timeouts: a connection timeout (how long to wait for the socket to open) and a read timeout (how long to wait for the response body).
I set connection timeout to 10 seconds and read timeout to 90 seconds. If a call exceeds either, it's killed and logged as a timeout failure — not retried indefinitely.
Rule of thumb: your read timeout should be roughly 3x your p95 latency. If your median call takes 8 seconds, a 90-second ceiling catches genuine hangs without killing slow-but-valid responses.
The OpenAI API docs recommend configuring client-side timeouts explicitly, because the default behavior varies by SDK and language. Don't trust defaults here. Set them yourself.
Fix #2: Exponential backoff with jitter (not naive retries)
Naive retry logic is a budget killer. If you retry immediately on a 429, you're just adding fuel to the rate-limit fire.
What worked for me: exponential backoff starting at 2 seconds, doubling each attempt, capped at 60 seconds, with random jitter of ±20%. Max 4 retries. After that, the call is marked failed and the sweep continues.
AWS's architecture guidance on timeouts, retries, and backoff with jitter explains why jitter matters: without it, all your retries fire in sync and recreate the exact spike that caused the failure.
Once I added jitter, my 429 cascade problem disappeared. Same API. Same rate limits. Just smarter retry timing.
Fix #3: Checkpoint every single call to disk
This is the fix I should have made first. After every successful call, append the result to a JSONL file. One line per call. Flush immediately.
JSONL (JSON Lines) means each line is a complete, independent record. If the runner crashes at run 4,700, you have 4,699 valid lines on disk. Restart the runner, it reads the checkpoint, skips completed calls, and resumes at 4,700.
No lost sweeps. Ever. The cost is a few milliseconds of disk I/O per call. Worth every millisecond.
Fix #4: A hard cost ceiling that halts the run
Track cumulative token cost as you go. Set a ceiling before you start. If the runner crosses it, stop everything and alert.
I set mine at $40 for the 4,768-run benchmark. The runner hit $38.20 and finished clean. But I've seen a runaway retry loop burn $200 in an afternoon because nobody capped it.
According to Anthropic's pricing documentation, costs scale linearly with tokens — which means a retry bug scales linearly too. A ceiling is your circuit breaker.
What the hardened runner actually looks like
Four components, in order:
- A task queue that holds every (prompt, model, temp) combo as a job.
- A worker pool that pulls jobs, calls the API with hard timeouts, and retries with jittered backoff.
- A checkpoint writer that appends every result to JSONL and flushes.
- A cost monitor that halts the pool when the ceiling is crossed.
That's it. No fancy orchestration framework. No distributed queue. Just four things that each do one job well.
I ran the full 4,768 calls in 6 hours and 12 minutes. Two timeouts, both recovered by retry. Zero lost sweeps. Final cost: $38.20.
The part nobody warns you about
Hardening a runner is unglamorous work. It's timeouts and file flushes and retry math. Nobody tweets about it.
But here's the honest truth: the quality of your results depends entirely on whether your runner survives long enough to produce them. A clever prompt matrix means nothing if the sweep dies at run 312.
If you're running LLM experiments at any real scale — even 500 calls — build the four fixes in from the start. Retrofitting them after a crash is a bad afternoon.
The same principle applies to the content side of AI work. When I'm generating a batch of product descriptions or blog drafts, the tooling that survives a 200-item run is the tooling worth using. That's the gap AI-Mind was built for — you describe what you want, pick a content type, and it handles the prompt engineering and the batch execution without you writing retry logic or managing a queue. The first 30 generations are free, which is enough to see whether the output holds up before you commit to a bigger run.
The lesson from 4,768 runs is simple. Reliability isn't a feature you add later. It's the foundation you build on. Checkpoint early. Cap your cost. Set your timeouts. Then run.
Key Takeaways
- Set hard connection and read timeouts on every LLM call — roughly 3x your p95 latency for the read timeout.
- Use exponential backoff with jitter for retries; naive retries cause rate-limit cascades that spike cost.
- Checkpoint every result to JSONL immediately so a crash never loses a completed sweep.
- Track cumulative token cost and halt the run at a preset ceiling to prevent runaway spend.
- A hardened runner is four simple components, not a complex orchestration framework.
Sources
- Amazon Web Services, Timeouts, Retries, and Backoff with Jitter, 2019. Engineering guidance on retry timing and jitter to avoid synchronized retry spikes.
- OpenAI, Rate Limits and Timeout Configuration, 2025. Official guidance on client-side timeouts and handling 429 responses.
- Anthropic, API Pricing and Token Costs, 2025. Reference for how token usage maps to cost at scale.
- JSON Lines, JSONL Specification, 2024. Format reference for append-only, crash-safe result logging.
Frequently Asked Questions
What's the difference between a timeout and a hang in an LLM runner?
A timeout is a call that exceeds your configured time limit and gets killed cleanly. A hang is a call that opens a connection but never responds and never errors — it blocks indefinitely unless you have a hard timeout set. Hangs are more dangerous because they silently stall your entire run. Setting explicit connection and read timeouts is the fix for both.
How do I stop retries from blowing up my LLM API cost?
Two levers: exponential backoff with jitter so retries don't fire in sync, and a hard cost ceiling that halts the run when cumulative token spend crosses a preset limit. Cap retries at 3-4 attempts and mark the call failed after that rather than retrying forever. Together these prevent the runaway loops that burn hundreds of dollars in an afternoon.
Why use JSONL for checkpointing LLM run results?
JSONL stores one complete JSON record per line, so each result is independent and self-contained. If your runner crashes mid-sweep, every line already written is valid and recoverable. You restart, read the checkpoint, skip completed calls, and resume. It's append-only, crash-safe, and trivial to parse — which is exactly what you want for long batch runs.