Validating the Migration and Keeping Rollback Real
"Looks like it all works." It is nine on the Monday after cutover, the coffee is still warm, and somebody has just said the sentence where migrations go to die quietly. "Works" on day one means the loud failures didn't happen. It says nothing about the file that will silently miss its window on the first month-end. It says nothing about the flow that drifted in a way nobody has looked for yet. Between cutover and the day you power off the old platform stretches the migration's most neglected phase, and its most decisive one. It is the phase where a cutover weekend either stays a weekend or quietly becomes a fortnight.
This closing article of our Workload Migration series covers the two jobs of that phase. First comes validation: assembling proof of equivalence — same files, same windows, same results — flow by flow. The work continues until "it works" becomes a claim with evidence attached. Second, rollback: keeping the escape hatch genuinely usable, with triggers decided in advance and a path that is rehearsed rather than merely written down. The migrations that end well are the ones where rollback stayed real until the moment the evidence made it unnecessary. Relief is not evidence; it just arrives earlier.
Equivalence: Same Files, Same Windows, Same Results
Define what you are trying to prove, or you will collect noise. Equivalence for a migrated flow has three axes:
- Completeness — every file the old platform would have produced or received exists on the new one. Nothing missing, nothing extra.
- Timeliness — the files move in the same windows. The nightly manifest that always landed between 01:30 and 02:30 still does. Nothing now misses a cutoff that downstream systems and partners built their day around.
- Fidelity — the contents and the observable behavior match: hashes agree, names follow the same patterns, permissions and encodings survived, the upload-then-rename convention still holds.
Two disciplines follow from the definition. Equivalence is judged per flow, never per platform. "The platform works" is just the average of fifty flows, and averages hide the one that matters. And equivalence is judged against recorded old behavior. That is why the shadow-run comparisons from migrating the automation and the arrival-time history in your logs are worth more than anyone's memory of how things used to run. The end-to-end habits in verifying transfers end to end supply the fidelity mechanics.
Meridian Parts ran forty flows through the shadow period, and thirty-nine of them were clean from the first day. The fortieth, a nightly price extract to a distributor, matched on count and names and failed on size, by one byte, on one file each night. The byte was a trailing newline the new job wrote and the old one never had. The distributor's import, it turned out, would have rejected the file for it. Nobody would have noticed until the first Monday after cutover, when the distributor's price list failed to load and the phone rang. The fix was one setting; the streak counter reset to zero; five clean days later the flow was validated for real. Thirty-nine clean flows is a fine average, and the average was never the point.
The Evidence Pack
For each flow, assemble a small, boring bundle that would convince a skeptic — a new team member, an auditor, or yourself at the decommission decision. It reads like this:
Evidence pack -- FLOW-041 (Bayside manifests, inbound)
1 reconciliation log shadow-period comparisons: five consecutive
clean days (counts, names, sizes, hashes)
2 first-live-run record server log extract, Mar 14: login 02:09,
manifests_YYYYMMDD.csv landed 02:14,
downstream import consumed it at 04:15
3 timing series arrival times for ten cycles before and
ten after cutover -- window unchanged
4 partner confirmation Bayside reply, Mar 15: "receiving normally,
reconciliation totals match"
5 sign-off flow owner's initials and date, in the
register row
The flow's owner assembles it; the register links to it; and its existence is what the word "validated" means in this series. The pack pays for itself three times. It is the decommission gate's raw material. It is the answer when an auditor asks how you knew the move was safe. And — most practically — it lets you decline a panicked rollback demand when some unrelated outage gets blamed on the migration, as one will be.
The Platform-Level Comparison
Per-flow packs are the foundation, but fifty green rows can still hide an estate-level lean. Add one roll-up view that compares the platforms as wholes, day by day. Compare total files in and total files out, per day, old platform's historical normal against the new platform's actuals. Compare total bytes moved, failure counts, and the spread of arrival times for the flows with deadlines. None of this needs tooling beyond the activity logs and a small script. It is the same counts-names-hashes ladder from the job-level reconciliation, aggregated upward. Fifty green rows can still add up to a leaning wall.
The roll-up catches a class of problem the per-flow view misses. A platform that is slightly slower per transfer can pass every individual comparison while the nightly batch as a whole finishes forty minutes later. That is fine on a quiet night, but a missed cutoff at month-end volumes. Watch the trend of the daily totals and completion times across the soak, not just each day's pass/fail. Drift that grows is a finding even while every number is still inside tolerance. Set the tolerances in writing before the soak starts (how late is late, how many failures are normal). A threshold invented on the day it is crossed will be invented generously. I have invented one myself, and it was very generous.
The Soak Period
The soak is the stretch after cutover when the new platform carries production while the old one stands ready. The old one is deliberately kept alive, not lingering by neglect. There are two things the soak is not. It is not "waiting to see if anyone complains," because the whole lesson of this series is that transfer failures are silent. It is not optional because cutover night went smoothly. Cutover night proves only that connections work. The soak is what proves the rhythm works — the month-end surge, the Monday catch-up, the quarterly job's first live run. Its length is set by the calendar, not the clock: every flow must pass through its real rhythm at least once. That means a minimum of one full business cycle, including a month-end if anything in the estate runs monthly. For most estates that lands between one and two months. Rare quarterly flows either extend the soak for their season or get an explicitly planned first-live-run watch of their own after the formal soak ends. The calendar decides, and the calendar does not take meetings.
A soak without instruments is just waiting. Three instruments do the work, and the diagram below shows how they drive the loop:
- Freshness checks per flow — the expected-arrival monitoring from freshness checks for expected files, which alarms on the absence that every other monitor misses. If the migration leaves one permanent gift behind, make it this.
- Failure and retry rates against the old baseline — not "are there errors" (there always are) but "are there more errors than the old platform's normal." Your own outbound jobs should announce their failures. A scheduler such as Sysax FTP Automation can send a notification the moment a push fails. That keeps the soak review fed with same-day facts.
- The human channels plus the old endpoint's logs provide partner tickets and helpdesk mentions on one side. On the other, they show every login still arriving at the old platform. Each one is a straggler to chase through the partner coordination ladder before decommissioning can even be discussed.
Hold the review daily, ten minutes, same short cast: instruments read, exceptions assigned, clean days counted per flow. The reading is only as good as the records underneath it. That is one more argument for a target platform with evidence-grade logging. A server that writes activity to both file and database, as Sysax Multi Server does, turns "what happened to FLOW-041 last night" into a query rather than an expedition.
Rollback Triggers Decided in Advance
The worst time to decide whether to roll back is the moment you need to. At three in the morning, with a partner waiting and the team exhausted, two biases collide. One is sunk-cost reasoning ("we've come too far to go back"). The other is panic ("put it all back the way it was"). Both are wrong often enough that the decision should not be made then at all — only executed then. During the calm before cutover, write the triggers down as condition, response, and decider:
| Observation during soak | Pre-agreed response | Decided by |
|---|---|---|
| A critical flow fails and cannot be fixed within its recorded late tolerance | Per-flow rollback: re-enable the old job or endpoint for that flow only | Flow owner + migration lead |
| Several flows failing with one shared root cause in the new platform | Platform rollback: reverse the addressing flip, old platform resumes | Migration lead + service owner |
| Any content-integrity difference: wrong, truncated, or corrupted files | Freeze further waves; per-flow rollback of affected flows; investigate before anything else moves | Migration lead, no delegation |
| Elevated-but-coping: retries up, windows still met | No rollback; fix forward with a named owner and a review date | Daily soak review |
Notice the shape: most triggers point at per-flow rollback, because the register gave every flow its own reversible cutover. Whole-platform rollback is the rare, big lever — kept oiled, rarely pulled. And say it out loud to the team before go-live: a rollback executed on a trigger is not the migration failing. It is the safety system working exactly as designed, and the flow goes back across once the cause is fixed. The same decide-first habit, scaled down to everyday changes, is in change rollout and rollback. A rollback is the smoke alarm going off, not the house burning down.
Pair the trigger table with a runbook so the response is execution, not archaeology. Both levels fit on a card:
Per-flow rollback (target: under fifteen minutes)
1. Disable the flow's job or account on the NEW platform.
2. Re-enable its counterpart on the OLD platform (job or account
was left disabled-not-deleted at cutover).
3. If the flow accumulated state on the new platform, run the
flow's return-sync steps (written per flow, see below).
4. Confirm the next cycle end to end on the old platform.
5. Update the register row; note the trigger that fired.
Platform rollback (target: under one hour)
1. Migration lead and service owner confirm the trigger together.
2. Reverse the addressing: point the service name back at the old
platform (five minutes at the low TTL) or reverse the address
move.
3. Re-enable all old-platform accounts and jobs from the register.
4. Run the return-sync for flows with accumulated state.
5. Notify partners on the prepared "we have reverted" notice;
freshness checks confirm each flow as it resumes.
What Keeps Rollback Real
Rollback plans rot silently while everyone watches the new platform. Keep these true for the whole soak, and check them weekly like the backups they are:
- The old platform stays whole. Accounts are disabled per wave — as the cutover discipline requires — but disabled reversibly, one step from working. Nothing is deleted, credentials are intact, and jobs are preserved in their switched-off state.
- The paths back stay open. The DNS TTL stays low so the name can flip back in minutes, as set up in cutover strategies. Partners keep the old address in their allowlists ("add, don't replace" exists for this moment). Your own firewall keeps the old platform reachable.
- Identity does not expire mid-soak. A certificate on the old platform that lapses three weeks into the soak converts your five-minute rollback into a rollback-plus-certificate-emergency at the worst hour.
- The data gap has a plan. Every day of soak, files accumulate on the new platform that the old one has never seen. Rolling a flow back means bringing its recent state back with it — a reverse of the same seed-and-delta copying that fed the migration. Write the return-sync steps down per flow before they are needed; "we'll figure out the files" is where three-in-the-morning rollbacks go wrong.
- It has been rehearsed. Once, early in the soak, pick one low-stakes flow and actually roll it back and forward again. It is a fire drill with a stopwatch. The drill converts the plan from prose into a checklist someone has run. It finds the expired password on the old side while that discovery is still free.
The Point of No Return — Declared, Not Discovered
Rollback does not stay cheap forever, and pretending otherwise is its own hazard. At some point, reverting stops being "flip back" and becomes "a restore project." Enough state has accumulated on the new platform, partners are beginning to retire old rules, or the old hardware is needed elsewhere. The mature move is to declare that point rather than discover it. At the decommission gate, announce that same-day rollback has ended. Switch the on-call playbook from "revert" to "fix forward," and only then start dismantling. An undeclared point of no return means someone on call believes in an escape hatch that quietly stopped existing. I have been that someone, and the hatch turned out to be a painting of one.
Decommission Only After the Evidence Says So
Turning off the old platform is the migration's last irreversible act, so it gets a gate of its own — every line, not most of them:
Decommission gate -- all must be true:
[ ] Every register row is at "soaked," with its evidence pack linked.
[ ] The soak covered one full business cycle, including a month-end,
with the clean-day targets met per flow.
[ ] The old platform's logs show no legitimate traffic for the agreed
quiet period -- every late straggler chased and resolved, not
just observed.
[ ] Rare flows (quarterly, annual) have either run clean once on the
new platform or have a named owner watching their first live run.
[ ] Files remaining on the old platform have been reviewed: migrated,
archived per retention policy, or deliberately retired.
[ ] The point of no return has been declared to everyone on call.
Then shut down in stages, letting each stage disprove the previous one. First disable the listeners but leave the machine and its logs running for a final quiet period. A connection refused makes a straggler visible, whereas a vanished machine just makes them fail mysteriously. Archive the configuration, keys, and activity logs before anything is wiped. The old platform's records may need to answer questions for years, and your retention policy says exactly how long. Confirm the disks hold nothing unmigrated, with the checks from verifying nothing left behind. Release addresses and retire DNS names deliberately, keeping the service alias pointed at the new platform. The same last-listener discipline, told for protocol retirements, is in proving FTP is gone. The logs, not the project plan, are the arbiter of "nobody uses it anymore." The archive itself is a configuration backup by another name, and backing up transfer configuration says what belongs in it.
Remember: decommissioning is the only migration step you cannot roll back. It spends the escape hatch. That is why it comes last, why it waits for evidence instead of a date, and why the gate above is a checklist and not a feeling.
Wrapping Up: The Migration Ends With Evidence, Not Relief
Relief arrives the morning after cutover; the end arrives weeks later, when all the evidence is in place. Every flow has an evidence pack, the soak has crossed a month-end, and the stragglers are resolved. The old platform goes dark against logs that prove its silence. That is what this series has been building toward since the opening article: not a heroic weekend, but a sequence of boring, verifiable steps. The steps are inventory, build, port, coordinate, cut over, soak, decommission — each one gated on proof. If you are starting your own migration, begin at the register and keep the partners close. Keep rollback real until the evidence lets you retire it along with the old machine. Then you may say "looks like it all works," and mean it.
Frequently Asked Questions
How long should the soak period last?
How do we tell a rollback trigger from a bug we should just fix?
Who should have the authority to roll back at night?
Files arrived on the new platform during the soak. How does rollback handle them?
Do we have to keep the old server's hardware after decommissioning?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
