Safe Reprocessing: Running Yesterday's Files Again
Most of this series is about stopping files from being processed twice by accident. This article is about the mirror image: processing files twice on purpose. Sooner or later a legitimate reason arrives — a fixed bug, a restored database, an auditor's question. Someone needs last Tuesday's files, or last month's, to flow through the pipeline again. If your duplicate protection is any good, it will do exactly what you built it to do: refuse. The ledger sees files it has already recorded and skips every one of them.
At that moment, teams without a plan start improvising — hand-copying files, commenting out checks, deleting ledger rows. The improvisations cause more damage than the original problem. Teams with a plan flip a switch that was designed for this: an explicit reprocess mode. It is scoped to an exact set of files, logged with who asked and why, and coordinated with the people downstream. This article — part of our duplicate detection and idempotency series — is the design for that switch. It covers the flags, the scoping discipline, the downstream etiquette, and the audit trail that explains the rerun months later.
Reprocessing Is Routine, Not an Emergency
Start by normalizing the need, because pipelines are often designed as if reprocessing were unthinkable. The requests that will actually arrive, all of them reasonable:
- The fixed bug. The transform mangled a column for three weeks before anyone noticed. The code is fixed; now the affected files must run through the corrected version.
- The restored consumer. A downstream database was recovered from backup and is missing four days of loads. The files still exist; they need to be applied again.
- The corrupted output. Something damaged the pipeline's results after processing — the inputs are fine, the outputs are not, and re-deriving them is the cleanest fix.
- The new consumer. A new system wants history: "load us everything from the last two quarters" is a reprocessing request wearing a project plan.
- The audit question. Someone needs to demonstrate that a given input produces a given output — which means running an old file again, carefully, without disturbing anything.
Notice what these have in common: none is caused by the pipeline misbehaving, and none can be refused. A pipeline that runs for years will field several of these. Designing the door in advance is the difference between a calm afternoon and a weekend of hand-surgery. It also completes the idempotency story from the plain-words article. Rerun-safety question eight — "is there a documented, deliberate way to reprocess a file on purpose?" — is this article.
The Hand-Edit Anti-Patterns (and How Each One Bites)
Without a designed mode, reprocessing gets done in one of four improvised ways. Each works, in the sense that the files get processed. Each also plants a mine:
- Copying files back into the intake folder. The live pipeline eats them — along with whatever else it does to fresh arrivals. Notifications fire again, forwarding steps forward again, and the replayed files interleave with tonight's real traffic. So nobody can later tell which run produced what. If the ledger is working, the copies are skipped and the exercise silently does nothing, which is its own confusion.
- Commenting out the duplicate check "for a minute." Now the entire pipeline is unprotected, not just the two files in question. The change is invisible, uncommitted, and famously forgotten. The retry storm that arrives while the check is disabled writes the incident report for you.
- Deleting rows from the ledger. This "works" — the files look new again — but it falsifies history. The ledger now claims those files were processed once, on the wrong date. When someone later reconciles ledger against database against server logs, the numbers disagree and trust in the whole record collapses.
- Cloning the script and editing paths. The clone starts life one edit away from correct and drifts further every time. Eventually someone runs the clone thinking it is the original, or the original with the clone's half-applied edits.
Never disable duplicate protection globally to reprocess specific files, and never delete ledger history to make files look new. Both convert a scoped, explainable operation into an unbounded, unexplainable one. The ledger is a diary: you may add today's entry — "processed again, here's why" — but you do not tear out last week's page.
The Design: Reprocessing as a First-Class Mode
The alternative is one honest principle. The pipeline itself, running its normal code path, should know how to reprocess. It should do so under an explicit flag, within an explicit scope, leaving an explicit record. Concretely, a good reprocess mode has six properties:
- Explicit activation. A command-line flag or job parameter —
--reprocess— that a human sets deliberately. It is never the default, and it cannot be triggered by file arrival. - Mandatory reason. The mode refuses to start without a stated reason and requester — the two facts the audit trail needs most and hand-edits never record.
- Bounded scope. An explicit file list or a tight date range. No bare wildcards; the mode should make it hard to reprocess more than you meant. That is because "backfills gone wide" is a classic duplicate source — see where duplicates come from.
- The same code path. Reprocessing runs the very pipeline it is replaying — same validation, same transform, same load. The mode changes only the ledger interaction and the side-effect behavior. A separate "replay script" is a clone with a flag's worth of justification.
- Truthful ledger updates. Instead of skipping recorded files, the mode processes them and appends a new event: result
reprocessed, pointing at the reason. History gains a line; it never loses one. - Deliberate side effects. Notifications, forwarding, and triggers are suppressed by default in reprocess mode. They are re-enabled only by a further explicit flag when the replay genuinely should fire them.
Here is the shape as an operator sees it — a command line you can adapt to any scripting stack, and the behavior contract behind each flag:
ingest --reprocess \
--files reprocess_list.txt \
--reason "TICKET-0042: transform bug fixed, re-deriving outputs" \
--requested-by jjohnson \
--dry-run
# Flag contract
--reprocess activate the mode; without it, ledger skips apply as normal
--files FILE explicit list, one name per line, sourced from archive/
--from / --to alternative bounded range (both required; no open ends)
--reason TEXT mandatory; written to every audit line and ledger event
--requested-by ID mandatory; who asked (not who typed — record both)
--dry-run resolve and print the scope, touch nothing (default first step)
--notify re-enable downstream notifications (default: suppressed)
--force-refetch pull files from the source server instead of archive/
Two of those defaults deserve emphasis. --dry-run should be the mode's reflex: run it first, read the resolved list, count the files, then run without it. And sourcing from archive/ rather than re-downloading matters because the archive is your preserved, already-verified copy of what was actually processed. Refetching from the partner's server risks getting a regenerated file that differs from history. Which raises a dependency worth saying plainly: you can only reprocess what you kept. A reprocess mode is only as deep as your archive retention. So the retention window for processed files should be set with replays in mind. The tradeoffs live in retention basics for admins.
Scoping: Deciding Exactly What Runs Again
Most reprocessing mistakes are scoping mistakes — too many files, the wrong window, one boundary day doubled. The discipline that prevents them:
- Build the list from records, not memory. Query the ledger for files processed in the affected window, or list the archive folder against the date range. Suppose the trigger was "the transform was broken between two deployments." The ledger's run identifiers tell you precisely which files went through the broken version.
- Cross-check against arrivals. Before trusting the list, compare it with what actually arrived in the window, from the transfer server's own records. A server that logs every transfer to a queryable store makes this a two-minute check. Sysax Multi Server, for instance, logs each transfer to file and database. So "every file received between those dates, with sizes and timestamps" is a query. Its answer either matches your list or tells you the list is wrong before the rerun, not after.
- Mind the boundaries. Off-by-one days at the edges of a range are the classic replay bug. State boundaries in full ("from the 3rd through the 9th, inclusive") in the ticket. Make the dry run print the first and last file it resolved.
- Start with one. For any replay bigger than a handful, reprocess a single representative file first, verify the result end to end, then run the rest. The single-file rehearsal catches wrong-version code, wrong-target config, and permission surprises at the cheapest possible size.
Generating the list is usually one query or one command against the ledger. With the database-table ledger from the detection article, "everything the broken code touched" looks like this. With a flat-file ledger, a filtered read of the same fields does the same job:
-- files processed by the runs that used the broken transform SELECT file_name, sha256, processed_at FROM processed_files WHERE run_id BETWEEN 'run_0042' AND 'run_0055' AND result = 'loaded' ORDER BY processed_at; # flat-file equivalent: pull the same window out of the ledger awk -F'|' '$5 >= " run_0042 " && $5 <= " run_0055 "' processed.ledger
Save the output as reprocess_list.txt, attach it to the ticket, and hand the same file to --files. The list you reviewed is then, verbatim, the list that runs — no second transcription, no wildcard reinterpreting the scope at execution time.
Warning Downstream: Replays Are Visible to Other People
A replay that is safe inside your pipeline can still ambush the people after it. The outputs change, or re-arrive, or double, depending on how the consumer works. A consumer who was not warned treats all three as incidents. The etiquette is short:
- Tell them before, not during. What is being replayed, why, the expected window, and the expected volume ("about 40 files, roughly 50,000 rows, landing between 14:00 and 15:00").
- Make replayed output identifiable. Carry the run identifier into the output — a marker column, a header line, or simply the reprocess run id in the delivery filename. That way, the consumer can distinguish replayed data from fresh data without calling you.
- Agree on replace-versus-append. If the consumer loads whatever arrives, a replay doubles their data unless they also key on file identity. The conversation "when we resend a file, do you overwrite or append?" is ten minutes now versus a reconciliation project later.
- Close the loop. When the replay finishes, confirm completion and final counts. Your jobs may already announce themselves. Scheduled tasks in Sysax FTP Automation can send email notifications when a task completes or fails. If your jobs do this, point those notifications at the stakeholders for the duration of the replay. That way, "is it done yet?" answers itself.
Timing matters as much as notice. Run replays in a window when the live pipeline is quiet — after the nightly run, or with tonight's schedule deliberately paused. That way, replayed traffic never interleaves with fresh arrivals. Interleaving is not dangerous to a well-gated pipeline, but it is confusing to every human reading the logs afterward. Confusion is what post-replay questions are made of. A replay that owns its window produces a log anyone can read: normal traffic, a clearly-bracketed reprocess block, normal traffic again.
The Audit Trail: Explaining the Rerun Months Later
The final property of safe reprocessing is that it explains itself later, because "later" always comes. A reconciliation shifts, a number changes between two report runs, and someone asks: why does the settlement table say these rows were loaded twice? The answer must be in the records, findable by someone who was not there. Two records carry it: the log and the ledger.
The log gets a structured block — a header when the mode starts, one line per file, a summary at the end:
Mar 21 14:02:11 REPROCESS start run_0057 Mar 21 14:02:11 reason="TICKET-0042: transform bug fixed" requested-by=jjohnson operator=mchen Mar 21 14:02:11 scope=reprocess_list.txt files=38 source=archive/ notify=suppressed Mar 21 14:02:14 REPROCESS settle_YYYYMMDD.csv sha256=9f86d081... prior=run_0042 result=reprocessed rows=1240 Mar 21 14:02:17 REPROCESS orders_YYYYMMDD.csv sha256=60303ae2... prior=run_0042 result=reprocessed rows=311 ... Mar 21 14:09:52 REPROCESS done run_0057 files=38 ok=38 failed=0 duration=7m41s
And the ledger gains one appended event per file — original entry untouched, new entry pointing back at it and at the reason. Read together, they answer every future question: what ran again, when, against which prior run, on whose request, for which ticket, with what outcome. This is the same evidence discipline that good pipelines apply to normal runs. Our guide to what to log covers the general craft. Here the discipline is applied to the runs most likely to be questioned. A replay with this trail is an operation; a replay without it is a rumor.
One last organizational note: pair the trail with a short runbook in your operations wiki. That way, the mode gets used the same way by everyone who touches it:
- Confirm the need and open a ticket; the ticket id becomes part of
--reason. - Build the scope list from the ledger, cross-check it against the server's transfer records, and attach it to the ticket.
- Dry-run; review the resolved list, the first and last file, and the count.
- Rehearse with one file; verify its output end to end.
- Warn downstream: window, volume, how replayed data is marked.
- Run in a quiet window with the live schedule paused.
- Verify final counts against the dry run; confirm with downstream.
- Close the ticket, quoting the reprocess run id and summary line.
Reprocessing done this way is so undramatic that the runbook's main job is convincing people it really is that boring — which is precisely the point.
The Version to Tell a Colleague
Reprocessing is running already-processed files again on purpose — and it is a routine need, not an emergency. The failure mode is improvisation: files hand-copied into intake, checks commented out, ledger rows deleted. The fix is a first-class reprocess mode: explicitly activated, requiring a reason and requester, bounded to an exact file list sourced from the archive. It runs the normal code path, appends truthful ledger events instead of rewriting history, and suppresses side effects by default. It logs a block that explains the whole operation to whoever asks next quarter. Scope from records, dry-run first, start with one file, and warn the people downstream.
The machinery this mode leans on — the ledger and its atomic claims — is built in detecting duplicates: names, sizes, hashes, and ledgers. The reasoning behind treating reruns as normal life is in idempotency in plain words. To see a reprocess mode installed as part of a full retrofit — incident, design, and after-state — finish the series with the worked example.
Frequently Asked Questions
Why not just delete the ledger entries and let the files look new?
Should reprocessed files come from the archive or from the sender?
Do notifications and downstream triggers fire during a replay?
How is reprocessing different from a backfill?
What is the minimum viable reprocess mode for a small script?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
