Catch-Up: Recovering the Batch After a Bad Night
06:20. "The morning reports are empty, finance can't reconcile, can you just rerun the night?" Somewhere in the small hours the batch went wrong. What happens in the next four hours separates shops that have a recovery discipline from shops that have a recovery adventure. The instinct under pressure is always the same, rerun everything, fast, and it is precisely the instinct that turns one bad night into three. A batch ecosystem half-ran. Rerunning work that already completed is how you double-pay claims and post the same cash twice. The night does not remember what it finished. You have to find out.
This article is the morning-after playbook. It covers how to assess what actually ran before touching anything, and how to choose a replay order the dependency chain will accept. It covers the guards that make reprocessing safe, and how to communicate while the catch-up is still in motion. It is part of our Nightly Batch Ecosystems series and leans on two earlier pieces. These are the night sheet from The Anatomy of a Batch Night and the register from Mapping Batch Dependencies. A recovery is only as good as the map it runs on.
Four Ways a Night Goes Bad
"The batch failed" describes at least four different situations. The first job of the morning is knowing which one you have, because each implies a different shape of catch-up:
- A feed never arrived. A partner file missed its cutoff or a transfer silently failed. The chain either held (good — it noticed) or ran without the data (bad — everything downstream is incomplete but "successful"). Recovery is dominated by the late file's fate, the territory of our cutoffs article.
- A job died mid-chain. The classic: adjudication crashed at 01:40, everything after it starved. Clean if downstream held; messier if clock-scheduled downstream jobs ran anyway against stale inputs.
- Infrastructure ate the night. A host rebooted for updates, a network link flapped, storage filled. Many jobs are affected at arbitrary points, including some that half-ran. Scheduled tasks may have simply not fired at all, the misfire family covered in surviving reboots and misfires.
- Everything ran, the data was wrong. The poisoned night: a malformed input sailed through, every job "succeeded," and the wrongness surfaced in a human's eyebrows at 08:00. The hardest case, because the state census below reports all green and the recovery is really a controlled reprocessing.
The recurring villain is not the step that failed but the step that ran when it should not have. Failed steps are honest; wrongly-ran steps are landmines, and finding them is what assessment is for. A crashed job at least has the decency to say so.
First, Do Nothing: The Assessment
Before any rerun, take fifteen minutes to establish the truth. The unit of assessment is the node — each job and each feed on the night's chain. The question for each node is not "did it error?" but four sharper ones: Did it start? Did it finish? Does its output exist and look sane — right size, right record count, control totals that add up? And has anything downstream already consumed that output? Fifteen minutes of reading has yet to make a morning worse.
Three records answer those questions: the scheduler's task history (what fired and what it returned), each job's own log, and the transfer server's activity log. The activity log is the night's black box, the neutral record of what actually landed and left, independent of what any job believed. This is where server-side logging pays for itself. A server like Sysax Multi Server keeps its activity log in both file and database form. So "did partner B's file arrive, when, and did anything download it?" is a thirty-second query at 06:30 rather than a scroll through the night's noise.
Mark every node with one of four states: GOOD (finished, output verified), DID-NOT-RUN, FAILED-PARTWAY (started, died; output suspect), or RAN-ON-BAD-INPUT (finished "successfully" against missing, stale, or partial data). Then find the high-water mark: the furthest point along the chain where everything upstream is verifiably GOOD. Everything before it stands; everything after it is in scope. Recovery is, in one sentence, replay from the high-water mark, in dependency order, with guards.
Remember: a green status is a claim, not a fact. "Completed successfully" means the job believes it finished — not that its input was complete or its output correct. Verify outputs (existence, counts, totals) before declaring a node GOOD, or your high-water mark is fiction.
The Morning-After Runbook
Here is the playbook as a numbered runbook — the version to adapt, print, and keep beside the night sheet. After all, 06:30 is not the hour to compose one:
- Stabilize. Pause the schedules and triggers that have not fired yet, so the mess stops spreading while you think. The night's remaining steps, this morning's dependent daytime jobs, and any auto-retries all wait until you have a picture.
- Preserve evidence. Copy the relevant logs aside; move or delete nothing. Half-written outputs and stray inbound files are exhibits first, cleanup later.
- Census every node on the chain: started / finished / output sane / consumed downstream. Write the four-state verdicts on the map itself.
- Find the high-water mark — the last point where everything upstream is verified GOOD.
- Choose the recovery shape: replay the remainder tonight-style (there is still time before business hours bite). Or choose partial catch-up (recover the critical path now, let slack branches wait). Or roll forward (accept the miss, run a double batch tomorrow — see below).
- Clear the cause or route around it. Fix the crashed job, obtain the missing file, or consciously proceed without it per the cutoff rules. Do not replay into an unfixed fault.
- Replay in dependency order from the high-water mark — upstream before downstream, verification gate after each step (counts, totals, expected files) before releasing the next.
- Apply the double-processing guards (next section) at every step that runs twice or ran partway.
- Release outputs deliberately. Reports, extracts, and partner files produced by the catch-up go out when verified — with a note when their timing or content differs from normal.
- Protect tonight. Check the collision between the catch-up and the coming night's schedule: will the rerun still be running at 19:30? Which cutoffs tonight are now tight?
- Debrief within a day: ask what broke, what the census missed, which step was scary to rerun. Turn each answer into a register update, a new check, or a runbook edit.
Step 5's third option deserves defending, because it feels like defeat and is often the disciplined choice. Some nights are cheaper to skip than to chase. Suppose catch-up cannot finish before the business day consumes the systems. In that case, running tomorrow's batch with two days' data — where the jobs support it — beats a recovery that collides with production hours and tonight's run. The cutoff calendar's "if missed" column should already say which feeds tolerate this. Choosing to skip is a decision. Running out of morning is not.
Replay Order: The Chain Decides
Replay order is not a judgment call; it is a property of the dependency map. Work strictly upstream-to-downstream from the high-water mark — the topological order of the graph. Any step replayed before its inputs were refreshed just manufactures a second generation of bad output. Three refinements make the order efficient rather than merely correct:
- Scope by descent, not by fear. Only nodes downstream of the break need replaying. The cash-application branch that completed cleanly off the settlement file does not rerun because adjudication died; consult the map, not anxiety.
- Run independent branches in parallel. Once the broken node is redone, its separate downstream branches — warehouse load and letter build, say — can catch up simultaneously. The map shows which recoveries share no edges.
- Triage by the critical path. Minutes spent on the critical path move the finish line; minutes on slack branches do not. Recover in the order that serves the deadlines still alive, and formally abandon the ones already lost. The 05:30 print run is gone by 06:20. Pretending otherwise wastes the exact minutes that could still save the 10:00 regulatory extract.
The two-crew pattern is worth stealing. One person (or crew) runs the recovery; another protects the normal schedule — fielding questions, watching the daytime feeds, preparing tonight. The recovering crew must not also be the communicating crew, or both jobs get done badly. I have watched a recovery run from inside a status meeting; it proceeded at meeting speed.
Double-Processing Guards
The signature catch-up injury is not failing to recover — it is recovering twice. Consider a replayed consolidation that re-includes an already-processed partner file, a ledger posting run a second time, a resent partner extract. Each turns a visible outage into invisible wrong data, which is strictly worse. The guards:
- Know, per step, whether it is idempotent. An idempotent step produces the same result run once or five times — the property explained in idempotency in plain words. Your runbook should carry a per-step verdict: safe to rerun / safe with preconditions / never blind-rerun. Steps that regenerate an output wholesale (extracts, report builds) are usually safe; steps that apply data (postings, payment releases) usually are not.
- Check-before-apply on the dangerous steps. Before rerunning ledger posting, ask the ledger: is batch
YYYYMMDDalready posted? A processed-batch ledger — even a table with one row per business date — converts "I think it didn't run" into a query. The broader toolkit (processed-file ledgers, quarantine folders, rename-on-consume) is laid out in safe reprocessing patterns. - Respect consumed-file markers. If the normal flow renames or moves files as it consumes them, the catch-up must honor those markers, not bulldoze them. A replayed consolidation should pick up exactly the unconsumed files — which is the point of designing consumption to leave tracks.
- Complete partial sets; do not blindly resend them. When three of five files in a multi-file delivery made it before the failure, the recovery sends the missing two — the reasoning covered in partial failure in multi-file transfers. Resending all five "to be safe" hands the receiver a duplicate-detection problem they may not have.
- Replay under the original business date. A catch-up on the 15th of the night of the 14th processes
YYYYMMDD-the-14th files, writes the 14th's outputs, and posts to the 14th — everywhere. The moment today's date leaks into yesterday's replay, downstream reconciliation breaks in ways that take days to untangle.
The census may have exposed steps that cannot be rerun safely — the job with no check, the script that appends blindly. If so, note them in the debrief as engineering debt. Making individual jobs rerunnable is a design discipline of its own, covered in designing for recovery. The ecosystem-level payoff is that every safe-to-rerun step makes some future morning shorter.
Bluewater Bank added its check-before-apply the morning after it needed one. After a storage failure, a rerun of the overnight fee posting was launched from the top "to be safe." It reapplied a business day that had already posted before the disk filled. The totals reconciled to exactly double, which is at least easy to spot, and a reversal batch went in before anything reached a statement. The posting job now asks the ledger whether the business date is already there before it writes a single row, and refuses if it is. The runbook's verdict for that step changed from "safe to rerun" to "safe with preconditions," which is the answer it should have carried from the beginning.
A Worked Recovery
Put the pieces together on one bad night at our worked insurer, Harborview Mutual. At 01:40 the adjudication run dies on a malformed record. The downstream ledger posting, clock-scheduled at 02:25, correctly refuses to start — its input check finds no adjudication output. But the letter build at 04:35 has no such check and runs happily against yesterday's adjudication file. The 06:20 census therefore reads: consolidation GOOD, adjudication FAILED-PARTWAY, GL posting DID-NOT-RUN, warehouse load DID-NOT-RUN, letter build RAN-ON-BAD-INPUT, cash application GOOD (separate branch). High-water mark: the consolidated claims file.
The recovery order falls straight out of the map. Quarantine the letter build's output first — it is wrong and must not reach the mail vendor. Fix the malformed record, rerun adjudication under the original business date, and verify its claim counts against the consolidation totals before releasing anything. Then GL posting — but first the check-before-apply: query the ledger for the business date, confirm nothing posted overnight. Then warehouse load and letter build in parallel, since they share no edges. The 05:30 print cutoff is already lost, so nobody chases it. The letters go out tomorrow, and the print vendor gets a courtesy note by 07:00. Reports land at 10:40. Total business impact: one report pack four hours late, one day's letters delayed. That is instead of the double-posted ledger a panicked 06:30 full rerun would have produced.
The debrief yields two tickets. One is an input-freshness check for the letter build (the step that ran when it should not have). The other is a malformed-record validation at consolidation so the poison surfaces at 23:35, not 01:40.
Communicating While You Catch Up
During a catch-up, silence is the accelerant: unanswered consumers escalate, escalations pull the recovering crew into meetings, and the recovery slows, generating more silence. The antidote is cheap and mechanical — early, bounded, business-language updates on a promised cadence:
BATCH STATUS — 06:55 update
Impact: Morning reports and warehouse data are STALE (yesterday's).
Claims payments for last night are NOT yet processed.
Not
affected: Cash application, lockbox posting, all daytime systems.
Cause: Overnight job failure at ~01:40; recovery underway.
Plan: Replaying processing chain in order; reports expected ~10:30.
Next
update: 08:00, or sooner if the picture changes. — Ops (H. Alvarez)
These habits make this work. Send the first message before you have answers (state what you know and when you will know more). Name what is not affected, which prevents more escalations than the impact line. Promise the time of the next update rather than the time of completion. That way, a slipping recovery does not also become a broken promise. Route partner-facing notes — your file will be late, or please do not resend — through whoever owns that relationship. Use the same channel used for cutoff changes. The first message may be mostly blank. Its job is to exist.
Making the Next One Smaller
Every catch-up buys information at retail prices; the debrief is where you keep the goods. Three questions turn an incident into infrastructure. What would have surfaced this at 23:40 instead of 06:20? (Usually an absence or freshness check — add it.) Which step was frightening to rerun? (Make it idempotent, or give it a check-before-apply.) Where did the map disagree with reality? (Fix the register while the evidence is fresh.) I have sat through the same debrief twice for the same step; the second time we wrote the ticket.
The transfer layer deserves a specific mention here, because it is where the cheapest prevention lives. A large share of bad nights begin as transfer failures that nobody heard. A partner upload died at 40%, or a pull hit a transient timeout and gave up. An automation client with genuine retry and error handling helps here. Sysax FTP Automation retries failed transfers, treats exhausted retries as errors, and emails a human when that happens. Such a client converts many would-be catch-ups into a 23:15 page and a five-minute fix. That is the best catch-up of all: the one that never becomes one.
The Discipline in One Paragraph
Assess before acting: census every node, verify outputs, find the high-water mark. Replay from there in dependency order, gates between steps, critical path first, abandoned deadlines abandoned out loud. Guard every rerun with idempotency knowledge, check-before-apply, and the original business date. Communicate early, on a cadence, in business terms. Then debrief until the next bad night is smaller. That is the whole playbook — and every piece of it leans on the map. That is why dependency mapping is the best pre-payment on a recovery you will ever make. The night still does not remember what it finished; now there is somewhere to look it up. When bad nights start coming from a window that has grown too tight, the series continues with When the Batch Window Shrinks.
Frequently Asked Questions
Why not just rerun the whole night from the top?
What is the high-water mark in a batch recovery?
How do I know if a job is safe to run twice?
Should the catch-up use today's date or yesterday's?
When should we give up on catching up and just skip the night?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
