Home › Topics › Nightly Batch › Modernizing

Modernizing a Batch Ecosystem Without a Big Bang

The slide is always the same. "Retire the nightly batch": new platform, streaming feeds, real-time everything, the old chain sent off with honors. Every team that operates an aging batch night eventually drafts it, and I have drafted it myself, with a diagram and a timeline. The slide is usually right about the destination and catastrophically wrong about the route. That is because a batch ecosystem has a property most modernization plans ignore: it runs tonight. And tomorrow night. There is no freeze window, no maintenance quarter, no moment when payroll, claims, and the ledger will politely wait while the replacement finds its feet. The night has no opinion about the slide. It just runs.

This article is the other route: an incremental path that improves the night in the order that keeps it safe. Visibility comes first, then rerun safety, then event-driven feeds where they actually pay. The article finishes with a worked migration of a single feed from clock-guessed to arrival-driven. That is the pattern you repeat until the ecosystem has quietly become a different thing. The article closes our Nightly Batch Ecosystems series. It assumes the assets the earlier articles built: the night sheet, the dependency register, and the runway log from When the Batch Window Shrinks.

Why the Big Bang Fails at Night

Batch ecosystems punish rewrite projects for structural reasons, not for lack of nerve:

  • The system is always in production. A night that runs every night offers no rehearsal space. The replacement must be built beside a moving machine and swapped in without stopping it. That is exactly what incremental methods are for and big bangs are not.
  • Half the ecosystem belongs to other people. Partners, banks, bureaus, and vendors own the far ends of your feeds. They will not re-platform on your schedule; whatever you build must keep speaking yesterday's interfaces at every boundary, indefinitely.
  • The requirements are archaeological. Decades of quiet fixes live in the current night — the tolerance for partner B's odd encoding, the ordering quirk finance depends on without knowing it. A rewrite rediscovers each buried requirement by breaking it, one bad morning at a time.
  • Parallel worlds cost double. The long transition every big bang actually becomes means running and reconciling two ecosystems with one team — the phase where most replacement projects quietly die.

One more correction to the slide: the destination is not "no batch." Consolidated end-of-day processing is not a legacy embarrassment. It is the correct shape for work that genuinely needs the day's complete data. It is astonishingly efficient at moving millions of records. The honest goal is a smaller, observable, rerunnable night. That means batch where batch is right, events where events pay, and nothing running at 02:00 that could safely run at 14:00. Batch is not a legacy. Batch at 02:00 for no reason is.

The Order of Operations

Incremental is not the same as unordered. The three phases below are sequenced so that each one reduces the risk of the next. That is why doing them backwards is the most common way incremental modernization fails. It means starting with the exciting event-driven part, on a night with no visibility and no rerun safety:

  1. Visibility. Instrument the night until every arrival, departure, run, and duration is recorded and absences raise alarms. Zero risk — it changes nothing the night does — and it is the instrument panel for every later change.
  2. Rerun safety. Make each step safe to run twice. This converts future mistakes — including modernization mistakes — from incidents into retries. It is the insurance policy the rest of the work is written under.
  3. Event-driven feeds. Move work out of the night by processing arrivals when they arrive. This is the phase that shrinks the window and de-fragilizes the cutoff scramble. It is only safe because the first two phases happened.

Phase One: Visibility

You cannot safely change what you cannot see, so the first phase adds eyes and touches nothing else. Concretely, four instruments, all buildable in weeks, none requiring a single job to change:

  • A central arrival-and-departure record. The transfer layer is the natural place to collect it: every partner upload, every job's pickup, every outbound send, timestamped in one queryable place. A transfer server that logs activity to a database as well as a file turns the night's file movements into a table you can chart. Sysax Multi Server does both. That table is the raw material for the night sheet, the runway log, and every "what changed?" question the next two phases will ask.
  • Expected-file checks on every register edge, so absence becomes an event instead of a silence — the discipline from freshness checks and expected files, applied edge by edge off the dependency register.
  • Failure notifications that reach a human at night, from the transfer tasks and the jobs alike, so problems surface at 23:15 while the night can still be saved.
  • The runway log records start, finish, and per-step durations trended monthly, as built in the window-shrinkage article. Modernization needs a scoreboard. And "the night finishes 40 minutes earlier than last quarter" is the sentence that keeps the program funded.

Visibility also pays a political dividend. The first month of real data usually retires several beliefs about the night — which feed is really late, which job really grew. It usually replaces the loudest anecdote with a chart. Every subsequent decision gets easier. The loudest anecdote rarely survives its first chart.

Phase Two: Rerun Safety

The second phase makes the night forgiving. Work through the chain one step at a time — most-feared step first. The debrief list from Catch-Up is a ready-made priority queue. Give each step one of the standard rerun-safety properties:

  • Regenerators become idempotent: extracts, builds, and report jobs rewrite their whole output for a business date. Running twice therefore produces the same files, not two generations of them.
  • Appliers get check-before-apply: posting and payment steps consult a processed-batch record before acting, so a replay is refused rather than doubled.
  • Consumers leave tracks: processed inputs are renamed or moved on consumption. So "what has been handled" is visible in the filesystem itself. These are the patterns detailed in safe reprocessing patterns.
  • Everything takes a business date parameter, so any step can be rerun as of the night it belongs to, instead of assuming "today."

At Harborview, the worked insurer from earlier in this series, phase two opened with the two steps its last bad night exposed. Ledger posting gained a processed-batch table and a check-before-apply (the step everyone feared rerunning). The letter build gained an input-freshness gate so it can never again run happily against yesterday's adjudication output. Neither change took a week; both were rehearsed by deliberately rerunning a completed night's step against the guard and watching it refuse. That rehearsal habit turns rerun safety from a documented hope into a tested property. Prove the guard by trying to defeat it on a quiet morning. We rehearse ours quarterly, and the quarter we skipped is the one I remember.

The per-job engineering behind these properties is covered in designing for recovery; what matters at the ecosystem level is the compounding effect. Each step made rerunnable shortens every future catch-up — and, less visibly, licenses every future change. Once mistakes cost a retry instead of a reconciliation project, the team can modernize at ten times the pace with a tenth of the ceremony. Rerun safety is not cleanup before modernization. It is modernization — the part that changes what the system can survive.

Remember: the order is the strategy. Visibility de-risks rerun-safety work; rerun safety de-risks event-driven work. Teams that skip to phase three are performing surgery with the lights off and no way to close the incision.

Phase Three: Event-Driven Feeds Where They Pay

The third phase changes the night's shape. Instead of everything waiting for a clock, feeds are collected, validated, and staged when they arrive. A partner file that lands at 21:15 is checked and staged at 21:16. Its problems surface while the partner's operators are still awake. The cutoff moment stops being a scramble of fetching and becomes a gate that closes on already-staged input. The building blocks are the watch-folder family of patterns — introduced in event-driven transfers and developed in the hot folder pattern. Both ends of the transfer layer can supply the trigger. On the receiving server, Sysax Multi Server's event triggers (in its Pro and Enterprise editions) can run an action the moment a partner's upload completes. On the automation side Sysax FTP Automation's folder monitoring launches a transfer task when a watched folder gains a file. Neither is a workload orchestrator, and this phase does not need one. It needs reliable arrival-triggered moves at the seams, which is precisely the transfer layer's job.

Two honesty clauses keep this phase out of trouble. First, event-driven does not abolish cutoffs. Consolidation still needs a moment when the set is declared complete; arrival-driven staging moves the work earlier, not the deadline. Steps that need all N inputs keep their gate — they just find everything already staged when it opens. Second, an arrival is not a file. Triggering on a half-uploaded file replays the classic partial-file disaster. So every arrival-driven step honors an arrival contract — completion markers, size-settle checks, debounced triggers — per arrival contracts and debouncing. Half a file arrives with the same fanfare as a whole one.

Where does it pay? Score each feed before touching it:

FEED SELECTION SCORECARD — convert the high scorers first

+2  arrives well before its consumer runs (dead air to reclaim)
+2  arrival time varies night to night (clock schedules guess badly)
+1  validation failures currently surface after the cutoff
+1  upstream of the critical path (minutes here move the morning)
-2  consumer needs many inputs at once (the gate stays anyway)
-2  producer cannot signal completeness (no marker, no stable size)
-1  feed or partner changing soon for other reasons (let it settle)

A Worked Migration: One Feed Goes Event-Driven

Here is the pattern applied once, at our worked insurer Harborview Mutual, to the highest-scoring feed: claims partner A. This partner uploads between 21:00 and 22:30. The consolidation step pulls and validates everything in a rush after the 23:00 cutoff. The migration, step by step:

  1. Write the contract down. One page: filename pattern (claims_ptnA_YYYYMMDD.csv), the completion marker the partner already sends (.done file), expected window, validation rules, and what happens on failure. Making the interface explicit before changing anything is the discipline our files-as-integration-glue series calls the file interface contract. Note that the partner's side changes not at all.
  2. Build the arrival path beside the old one. An event trigger on the receiving server fires when the .done marker lands. It launches the existing validation script, which stages the checked file into a shadow staging folder. The 23:00 clock path keeps running untouched, still feeding the real consolidation.
  3. Shadow-run and compare. For two weeks, every night produces two staged copies — clock path and arrival path. A small nightly diff (checksums, record counts) proves the new path stages byte-identical input. The visibility layer shows its timing: staged by 21:40 on average, problems surfacing two hours earlier than before.
  4. Cut the consumer over, keep the net. Consolidation now reads the arrival-staged file. The old clock pull demotes to a 22:45 fallback sweep that only acts if the arrival path staged nothing. Rollback is one config change, and rerun safety (phase two) means even a botched night is a replay, not a crisis.
  5. Retire the scaffolding on evidence. After a clean month, the fallback sweep becomes an absence alarm instead of a second path. The contract page then joins the register, and the migration is done. Total elapsed: about six weeks; total nights at risk: zero.

The visibility layer showed the measurable result of that one migration. The post-cutoff scramble shed eleven minutes (partner A's share of the fetching and validating). Two malformed files in the following quarter were caught and fixed with the partner before 22:00 instead of poisoning the 23:30 consolidation. The 23:00 cutoff gate itself became a two-second staging check. No consumer noticed anything except that nothing goes wrong anymore — which is the correct experience of good infrastructure work. Good infrastructure gets no thank-you notes, only fewer phone calls.

Then repeat — partner B next quarter, the settlement file after that. Each conversion banks its minutes of dead air and moves another failure mode from midnight to evening. This shadow-then-cutover shape is the same parallel-run discipline that governs any transfer migration, treated fully in our migrating transfer workloads series.

Keeping Payroll Running While the Ground Shifts

The program's operating rules matter as much as its phases, because the night must run every single night while you renovate it:

  • One change per cycle, measured. Land a change, watch a full week of runway and exception data before the next. The visibility layer is the safety interlock; use it.
  • Never change a feed and its consumer in the same week. One end of every edge stays still, so a regression has an obvious cause and an obvious rollback.
  • Freeze around the cliffs. No cutovers in month-end week, pay-cycle nights, or the nights before a regulatory deadline. The calendar from the cutoffs article marks the frozen zones.
  • Every change carries its rollback. The fallback sweep pattern above generalizes: the old path is demoted before it is deleted, and deletion waits for a month of evidence.
  • The register is the change ledger. Every migration updates the dependency register in the same ticket — the map stays true, and the next migration inherits accurate ground.
  • Partners hear about changes that touch them, and only those. The beauty of modernizing at the seams is that most conversions leave the external interface untouched. Partner A never learned its feed went event-driven. When a contract genuinely must change (a new marker file, a naming tweak), it travels through the same notice process as a cutoff change. That includes a shadow period on their side too.

Meridian Parts wrote the "one end of every edge stays still" rule after breaking it once. In a single week they moved a supplier price feed to arrival-triggered staging. Since the consumer was open anyway, they also changed the pricing load to read the new staging folder that week. The Thursday load ran on an empty folder, and the morning went on deciding which change had emptied it, because either could have. Both were rolled back, which fixed the night and proved nothing. They redid the two changes a week apart the following month, each with its own week of evidence. The second attempt was uneventful in the way the first was supposed to be.

What "Modern" Looks Like From the Morning

Run this program for a few quarters and the ecosystem transforms without ever having been replaced. Feeds validate themselves on arrival all evening, so the night begins with clean, staged input. The window holds only what genuinely needs the complete day — consolidation, posting, the loads — and finishes with runway to spare. Every step reruns safely, so the rare bad night is an hour's replay instead of a war room. The map is current, the runway chart is boring, and the on-call phone has learned to sleep. It is still a batch night — honestly, proudly so — but it is a batch night that can be seen, survived, and changed. The slide, not the chain, gets retired with honors.

That is the arc of this whole series. Learn the night's anatomy, map its dependencies, respect its cutoffs, drill its recoveries, and measure its window. Then, from that foundation, modernize it one safe step at a time, without a big bang, while it keeps doing its job in the dark.

Frequently Asked Questions

Is the goal of modernization to eliminate batch processing?
No. Work that needs the complete day's data — consolidations, postings, end-of-day loads — is correctly batch, and efficiently so. The goal is a smaller, observable, rerunnable night: batch where batch is right, arrival-driven processing where waiting for a clock adds nothing but risk.
Why does visibility come before everything else?
It is zero-risk (it changes nothing the night does) and it de-risks every later step. Without a record of arrivals, runs, and durations, you cannot verify a shadow run, score a feed for conversion, or prove an improvement. Visibility also usually corrects several wrong beliefs about where the night's problems actually are.
Do event-driven feeds get rid of cutoff times?
No — steps that need a complete set of inputs still need a moment when the set is declared closed. Event-driven staging moves the fetching, checking, and fixing earlier in the evening; the cutoff remains as a gate that now opens onto work already done.
How do I pick the first feed to convert?
Score candidates. Favor feeds that arrive long before they are consumed, arrive at unpredictable times, or fail validation late. Avoid feeds whose consumer needs many inputs at once or whose producer cannot signal a complete file. Convert one high scorer, shadow-run it, and let the evidence recruit the next one.
How long does modernizing a batch ecosystem take?
Visibility lands in weeks. Rerun safety proceeds step by step over months. Feed conversions run one or two a quarter under the one-change-at-a-time rule. Expect the character of the night to change within a year — with zero nights bet on any single cutover along the way.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.