Partial Failures in Multi-File Jobs: When Ninety-Nine of a Hundred Made It
"Did the invoice run work last night?" "It says OK." "Then where is invoice four?" A nightly job moves a hundred files to a partner. Tonight, ninety-nine arrive and one does not, and the deceptively simple question, did the job succeed, has no right answer. If it reports success, someone downstream is missing a file and nobody knows. If it reports failure, someone may re-run the whole batch and deliver ninety-nine duplicates to chase one straggler. Both answers are wrong, because the question is wrong: a multi-file job does not have an outcome. It has a hundred of them.
This article is about engineering for that fact. We will look at the two honest philosophies for batch outcomes, all-or-nothing and per-file accounting. We will look at the manifest that serves as the job's version of the truth. We will cover the resume-versus-restart decision after something goes wrong, and reporting that tells the real story instead of a comforting boolean. It is part of our Retry Logic and Error Handling series, and it extends the per-file thinking introduced by the dead-letter pattern to the whole batch. A boolean is a comforting thing to be handed at six in the morning. For a batch of a hundred, it is also the wrong data type.
The Lie of the Single Status
Most batch scripts inherit their shape from single-file scripts: do the work, exit zero or nonzero. Applied to a hundred files, that single bit becomes a lie in one of two directions. The bit is not lying on purpose; it only has two positions.
The optimistic lie is the script that loops over files, ignores individual errors, and exits zero because the loop finished. Its log says "job completed." Its monitoring shows green. The missing file surfaces days later as a business question — where is invoice such-and-such? The investigation starts from nothing, because the system that lost the file recorded no memory of losing it. This is the classic silent partial failure, and it is far more common than total failure. It is also the more polite of the two, which is the problem.
Northgate Retail found this out the slow way. Their nightly price push looped over one file per store, logged each error, and exited zero because the loop had finished. So monitoring had shown green for months. One store's file had been rejected on its name for four nights running before a store manager asked why the shelf labels still showed last week's prices. The fix was a nonzero exit whenever any file failed, plus a count of failures in the morning report. Nobody had done anything wrong; the script had been asked whether the loop finished, and it had answered truthfully.
The pessimistic lie is the script that exits nonzero at the first error, abandoning the ninety-nine files after the failed one. Worse still is the re-run of the entire batch that a bare "FAILED" invites, a reasonable move given the information on the screen. Now the partner receives most of the files twice, and whether that is harmless or a disaster depends entirely on the receiving system. (Re-delivery is a whole topic of its own; our duplicate detection and idempotency series covers what happens when the same file arrives again.)
A partial failure — some items succeeded, some did not — is the normal failure mode of multi-file work. Handling it well requires exactly one mental upgrade: the outcome of a batch is a list, not a bit. Everything else in this article follows from that.
All-or-Nothing vs Per-File Accounting
Before building anything, decide what a partial delivery means for each flow, because there are two legitimate and opposite answers.
All-or-nothing flows treat the batch as one unit: the files only make sense together, so a partial delivery is worthless or actively dangerous. A financial export where a header file summarizes detail files is the classic case. Process the details without the header (or vice versa) and downstream systems compute nonsense. These flows deliver into a staging area and only "publish" once everything has landed. Typically, they do that by uploading a completion marker file last. Or they upload into a temporary directory and rename it into place as the final step. The receiver's side of the contract is to touch nothing until the signal appears. The mechanics of markers, temp names, and atomic renames live in our partial file safety series. The same tools that protect a single half-written file protect a half-delivered batch.
Per-file flows treat each file as independent: an invoice is an invoice, and ninety-nine delivered invoices are ninety-nine useful things regardless of the hundredth. Here, partial success is real success that must be preserved. The job's duty is to deliver what it can, account precisely for what it could not, and never let one bad file take healthy ones hostage.
Choosing is usually easy once the question is asked out loud:
| Question about the flow | Points to all-or-nothing | Points to per-file |
|---|---|---|
| Do the files reference each other? | Yes — header/detail, dataset parts | No — each stands alone |
| Is a late complete batch better than a prompt partial one? | Yes — consistency beats speed | No — deliver what you have |
| Can the receiver detect an incomplete set? | No — so you must guarantee completeness | Yes — counts or manifests are checked |
| Typical examples | Database exports, software releases | Invoices, images, documents, logs |
All-or-nothing does not remove the need for per-file tracking; the sender still needs to know which files remain to be staged when a run is interrupted. It only changes what the receiver is allowed to see. Internally, every well-built batch job does per-file accounting; all-or-nothing adds a publication gate on top.
The Manifest: The Job's Version of the Truth
Per-file accounting needs a place to live, and that place is the manifest. It is a small file (or table) the job maintains, one row per work item, recording what should happen and what actually did. The job builds it at the start of the run by snapshotting the intake — this is the work list — and updates it as each file progresses. An excerpt from a working manifest, after a night with some weather:
# flow=invoices-out run=Mar 14 02:00 host=app01 # file bytes sha256(first8) attempts status detail inv_YYYYMMDD_0001.pdf 48211 9f31ab02 1 verified - inv_YYYYMMDD_0002.pdf 51877 0c77d1e9 1 verified - inv_YYYYMMDD_0003.pdf 2380 bb40129c 3 verified retried: 421 then ok inv_YYYYMMDD_0004.pdf 0 - 1 quarantined zero-byte file inv_YYYYMMDD_0005.pdf 49902 71d3c880 1 verified - ... # totals: 100 listed / 98 verified / 1 retried-then-verified counted above / 1 quarantined
Three details in that excerpt carry most of the value. The status vocabulary is small and unambiguous — pending, sent, verified, failed, quarantined. And "verified" is distinct from "sent" because an upload that completed is not the same as an upload confirmed intact at the far end. The difference is the whole subject of verifying transfers end to end. The size and hash columns pin down exactly which bytes were delivered, which settles later disputes and enables safe re-runs. And the attempts and detail columns preserve the story — file three needed three tries tonight. That is invisible in any single status bit but is precisely the early-warning data that predicts tomorrow's trouble.
If the word manifest sounds familiar from the integrity world: yes, this is the same idea seen from the other side. A sender-provided checksum manifest, as described in checksum files and manifests, lets the receiver confirm a batch is complete and intact. The job manifest here is the sender's internal ledger. Mature flows have both, and they should agree — when they do not, one side is lying and you want to know which.
Keep the manifest itself crash-proof, because it is only useful if it survives the failures it describes. Three habits do it. Give each run its own manifest file, named with the run stamp (run_YYYYMMDD_HHMMSS.manifest), so runs never overwrite each other's history. Update it as an append-only log or by the write-then-rename trick, so a crash mid-update cannot leave a half-written ledger. Store it beside the job's logs, under the same retention rules. That way, the answer to "what happened on the fourteenth" is still there when the question arrives weeks later. A manifest that lives only in the script's memory evaporates with the process; the entire point is that it does not. I once inherited a job whose whole ledger was an array in memory. It knew everything, right up to the moment it exited.
The manifest earns its keep three times over. During the run it is the work list. After a failure it is the recovery checkpoint. The next run reads it and knows exactly what remains, the mechanism behind jobs that recover themselves. After the run it is the raw material for honest reporting. Few files in the estate work that hard.
Resume or Restart?
Something went wrong at file sixty-one and the job stopped. The next run has two options, and picking the wrong one is expensive in opposite directions.
Resume means trusting the manifest: skip everything marked verified, retry everything marked failed or pending. It is fast, it minimizes duplicate deliveries, and it is the right default — provided the manifest can be trusted. That trust has one sharp edge: a file marked sent but not verified is in limbo. The upload may have completed just before the crash, or not. Resume logic must treat limbo items as unfinished and redo them. That is why the redo must be safe to repeat, either because an overwrite of the same bytes is harmless or because the receiver deduplicates.
Restart means discarding run state and doing the whole batch again. It is the right call when the state itself is suspect. The manifest might be missing or damaged, or the intake might have changed mid-run. Or the failure might suggest the environment was so unstable that even "verified" rows deserve skepticism. Restart is only safe when re-delivery is safe; if the receiving system processes whatever appears, a hundred re-sent files are a hundred incidents. When in doubt, resume. If you must restart, restart into the same verification discipline, letting size and hash comparisons turn most re-uploads into cheap no-ops.
One special case deserves its own sentence: a single large file that died partway through is not a batch problem but a within-file partial. Protocol-level resume and the safety rules around half-written files are covered in our partial file safety series. The manifest treats that file as one item still owed.
Where Retries Fit in a Batch
Per-file accounting changes retry placement. The wrong design retries the job because one file failed; the right design retries the file. In practice: attempt each file with a small inline retry — two or three tries with short backoff, per the backoff article. If it still fails, mark it failed in the manifest and move on to the next file. The batch keeps its momentum; one struggling file costs seconds, not the night.
What happens to the failed file afterward follows the patterns you already know. If the manifest shows it failing across multiple runs, it is a poison candidate and takes the dead-letter exit. Suppose the whole tail of the batch failed at once — files one through sixty fine, everything after file sixty dead. That is not sixty poison files but one environment failure mid-run. In that case, the correct response is the connection-level retry and, eventually, the give-up-and-alert path. The manifest makes the two cases look different at a glance. Poison is a lone row with a high attempt count; an outage is a wall of failures sharing one timestamp.
The exception to never-stop-at-first-error is the flow where order matters — file two must not be delivered unless file one was. Those flows exist (sequenced database changes, chunked archives) and for them, stopping at the first failure is correct per-file accounting. In those flows, mark the failure, mark everything after it blocked, and let the report say so.
Remember: in a multi-file job, the unit of failure is the file, but the unit of alerting is the run. Handle failures per file; tell humans per run. One summarizing message beats a hundred individual ones every time.
The Batch That Was Short to Begin With
Per-file accounting has one blind spot worth naming: it can only account for files it saw. If the upstream system produced ninety files on a night it should have produced a hundred, the manifest lists ninety rows. All of them end the run verified, and the report says OK with a clear conscience. The job did its work perfectly, on an incomplete batch. Ten files are missing and nothing in the transfer layer knows. The report is accurate, tidy, and ten files short.
Closing that gap takes an expectation from outside the run. Sometimes the expectation is explicit: the sender provides a manifest or count of what the batch should contain. That is the receiver-side use of checksum files and manifests. In that case, the job compares listed-versus-received before declaring victory. Sometimes it is statistical: this flow normally carries between ninety-five and a hundred and ten files. So a night with ninety deserves at least a warning line in the report. And sometimes it is a deadline: the absence of any batch at all by a certain hour is its own alarm. No per-run report can raise that alarm because no run happened. That last case — alerting on files that never arrived — is a monitoring concern rather than an error-handling one. It is covered in our transfer job monitoring series. The point for this article is modest but important: a green manifest proves the job delivered what it received, never that it received what existed.
Reporting That Tells the Real Story
The end-of-run report is the manifest folded into one line a human can absorb, plus the details a responder needs. The line that works looks like this:
invoices-out Mar 14 02:47 PARTIAL: 98 verified, 1 retried-then-verified, 1 quarantined (inv_YYYYMMDD_0004.pdf: zero-byte)
Counts first, then names — but only the names of the exceptions. The healthy majority appears as a number; every failure appears by name with its one-line cause. And the status word is PARTIAL, a first-class outcome alongside OK and FAILED, because that is the truth a binary vocabulary cannot say. What each of these lines should contain, and where they should go, is the craft covered in error messages worth logging and in what to log.
Watch what that one line does to the morning routine. The responder reads it at their desk and already knows: the flow ran, nearly everything landed, one named file is parked with a stated cause. Triage is a decision, not an investigation. Open the dead-letter sidecar for the zero-byte file, ask the upstream team why their export wrote an empty PDF, and requeue after the fix. Compare that with the alternative morning: a red "job failed" with no particulars. There is an hour of log spelunking to discover that ninety-nine files were actually fine. There is a nervous debate about whether re-running will double-deliver. Same failure, same night; the difference is entirely in the accounting. I have had both of those mornings. Only one of them ended before lunch.
For the scheduler, map the three outcomes onto exit codes deliberately: zero for a fully verified run, one distinct nonzero value for partial, another for total failure. Monitoring can then treat partial as "page nobody, but do not let it age." The file that failed tonight must be delivered or quarantined by tomorrow. A partial that repeats nightly is a defect, not weather. How those signals feed dashboards and alerts is the territory of our transfer job monitoring series.
Two places where tooling helps rather than replaces this thinking. A scheduled-transfer product such as Sysax FTP Automation brings task-level retry and error handling plus email notification on failure out of the box. That is the run-level half of the story. The per-file ledger remains your design to keep. And on the receiving side, a server like Sysax Multi Server logs every session and file operation to file and database. That gives you an independent per-file record from the other end of the wire. When your manifest says sent and limbo strikes, the server's log is the tiebreaker that says whether the bytes actually arrived.
The Shape to Carry Away
Build multi-file jobs so that the question "did it work?" has a precise answer. That takes a manifest row per file, a small inline retry per item, and a dead-letter exit for repeat offenders. It takes a PARTIAL outcome that reaches humans as counts-plus-named-exceptions, and a resume path that trusts verified rows and redoes limbo ones. From here, designing transfer jobs that recover themselves turns the manifest into a full checkpoint-and-recover mechanism. The article on poison files and the dead-letter folder details the eviction path this article kept pointing at. Ninety-nine out of a hundred is a fine night's work — as long as the job can name the hundredth.
Frequently Asked Questions
What is a partial failure in a transfer job?
Should a batch job stop at the first failed file?
What is a transfer manifest?
After a partial failure, should I resume or restart?
What exit code should a partially failed job return?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
