Failover Drills: Testing HA Before You Need It
A standby node that has never served traffic is a hypothesis. It probably has the right host key, and its certificate is probably current. The mirror job probably ran last night, and the runbook probably still matches the way the pair is wired. Every one of those "probably"s is a way for the first real failover to become the first real outage. That could happen at two in the morning, with a partner on the phone. The only way to turn a hypothesis into a fact is to fail over on purpose and watch what happens. Write it down, and fix what you saw.
This article is the practical guide to doing that. It covers the two kinds of drill, planned and simulated, and what to measure so that a drill produces numbers rather than impressions. It covers how to include an in-flight upload so you see what a partner would see. It provides a step-by-step drill runbook and a checklist. It covers the findings drills usually expose and how to fix them, and how to record results so the next drill starts from evidence. This article closes our High Availability for Transfer Services series, and it assumes the pair described in active-passive failover. Everything applies to a cluster as well.
Why an Untested Pair Is a Single Server With Extra Steps
Redundancy decays. The active node gets a new certificate and the standby does not. Someone adds a firewall rule on one machine. A patch changes a default on the node that was rebooted and not on the one that was not. The mirror job starts failing on a folder it cannot read and nobody reads the log. None of these produce a symptom, because the standby is never asked to do anything. So the pair keeps looking healthy right up until the moment it is needed. This slow divergence is called configuration drift, and a drill is the only routine event that surfaces it.
A failover drill is a deliberate, scheduled failover with observers, a stopwatch, and a written result. It is not a test environment exercise; it happens on the production pair, because the production pair is the thing whose behavior you need to know. Testing changes before they reach production is a different discipline, covered in our Testing and Staging Transfer Changes series. A drill tests whether the redundancy you already built still works.
How often? Once a quarter is the common answer, plus once after any change that touches either node. That includes a patch cycle, a certificate renewal, a host key rotation, or a change to the storage or the mirror job. The best estates fold the drill into the patch routine itself, so that every patch cycle is a drill with a purpose.
Two Kinds of Drill
A planned failover is the gentle version. In a maintenance window, you tell the pair to move the address, exactly as you would to patch a node. It proves the runbook, the identity every node presents, and the storage handover, with the least risk. A simulated failure is the honest version: you break something and watch whether the pair notices and recovers on its own. It proves the health check, the detection time, the automation, and the split-brain defenses. It carries more risk, because you are testing the parts that act without you. Do the planned kind first, and do the simulated kind only once the planned kind has passed twice in a row.
| Scenario | How to simulate it | What it really tests | Risk |
|---|---|---|---|
| Planned move | Lower priority or enter maintenance mode on the active node | Runbook, identity, storage handover, partner reconnect | Low |
| Service crash | Stop or kill the transfer service process | Health check depth, detection time | Low |
| Disk full | Create a file that fills the data volume | Whether the check does a real upload; shared-storage blind spot | Medium; on shared storage this breaks both nodes |
| Host loss | Power off the virtual machine or the server | Automation with no graceful shutdown; stale locks; partial files | Medium |
| Network loss | Disable the partner-facing network interface | Detection when the node is alive but unreachable | Medium |
| Heartbeat link loss | Disable only the heartbeat interface, both nodes still up | Split-brain defenses: quorum, fencing | High; have a hand on the power button |
Two of these deserve a warning. The disk-full scenario on shared storage fills the disk for both nodes. Run it only on a pair with separate storage, or simulate it by making the probe folder unwritable instead. The heartbeat-loss scenario can produce two active nodes, the very failure the design is meant to prevent. Run it with someone ready to power off a node by hand. Never run it on a pair whose split-brain defenses have not first been reviewed on paper. Those defenses are described in shared storage vs replication.
What to Measure
A drill without numbers is an anecdote. Decide before you start what you will measure, and give each item a name so that the next drill can measure the same thing. The core set is short:
- Time to detect: from the moment of failure to the moment the health check or cluster marks the node down. With a five-second check and three failures to trip, expect around fifteen seconds; a planned move detects instantly.
- Time to recover: from the moment of failure to the first successful partner transfer on the new node. This is your measured RTO, and it is the number partners care about. Take it from the synthetic partner's log, not from your own console.
- Files lost: files the old node's log shows as completed that are absent on the new node. This is your measured RPO. On shared storage it should be zero; with a mirror job it should be no more than the last interval's worth.
- Files duplicated: these may be files a downstream job already consumed and deleted on the old node that are present again on the new node. Or they may be files a partner uploaded twice because their retry did not notice the first attempt had landed. Our article on why duplicates happen explains both mechanisms.
- Sessions dropped and partner-visible errors: how many sessions were open at the moment of failure, and what the synthetic partner logged.
- Drift found: every difference between the nodes discovered during the drill, from a fingerprint mismatch to a missing firewall rule.
The diagram shows where each timing measurement starts and ends on a drill timeline.
A worked result
Suppose the drill stops the service on the active node at T+0. Node A's log shows the third failed check at T+14 seconds; node B's log shows the address arriving at T+15. The synthetic partner, retrying every thirty seconds, logs a reset at T+2, a failed reconnect at T+32, and a completed upload at T+63. Time to detect is fourteen seconds. Time to recover is sixty-three seconds, against a partner promise of one hour. Compare node A's "upload complete" lines for the two minutes before T+0 with node B's folder. That finds one file present on A and absent on B. It completed forty seconds before the trigger and the mirror had not run. Files lost, one, inside the two-minute RPO but a real file that someone must copy across when node A returns. Files duplicated, zero. Drift found: the standby's certificate expires in eleven days and the active node's was renewed last month. That last line is why you drill.
How In-Flight Transfers Behave in a Drill
Every drill should include a transfer that is in progress at the moment of failure, because that is the case partners will actually hit. Prepare a large probe file and start uploading it from the synthetic partner about a minute before the trigger:
# make a 3 GB probe file once, then start a slow upload from the synthetic partner dd if=/dev/urandom of=probe_3g.bin bs=1M count=3000 sftp -i probe_key -o StrictHostKeyChecking=yes probe@sftp.example.com <<'EOF' put probe_3g.bin /probe/probe_3g.bin.part rename /probe/probe_3g.bin.part /probe/probe_3g.bin EOF
Then trigger the failure and watch four things. First, the client: it should report a reset, not a hang. A client that hangs for minutes has no keepalive or timeout, which is a finding. Second, the partial file: on shared storage it is visible on the new node under its temporary name and nothing downstream touches it. With a mirror that excludes .part files it is absent from the new node entirely. Either is correct. A partial under its final name, or a downstream job that picks it up, is a finding. The fix is in temp names and atomic renames. Third, the retry: the client should reconnect, see the same host key, and either resume or start the file again. Note which, because that is what partners will experience. Fourth, the old node: when it comes back, the abandoned partial must be cleaned up, and the mirror must not copy it anywhere.
For FTP and FTPS, run the same test in passive mode. Confirm that the data connection for the retried transfer succeeds against the new node without any firewall change on the partner side. That is the proof that both nodes announce the same address and range. If a load balancer is involved, confirm that a retried transfer does not land on a node that has no idea about its passive port. That is the affinity problem explained in active-active and load-balanced clusters.
The Drill Runbook
Write the drill as a numbered runbook and follow it exactly, including the parts that feel unnecessary. The person running the steps is the driver. A second person, the observer, watches the synthetic partner and the logs and touches nothing. The scribe, who may be the observer in a small team, writes down every time and every surprise.
- One week before: schedule the window and send the maintenance notice using the wording from making redundancy invisible to partners. Confirm the synthetic partner has been running cleanly for at least a week.
- One day before: verify the host key fingerprints and certificate expiry on every node by its own address. Verify the mirror lag is inside the RPO. Take a fresh backup of configuration and data, because a drill is a change.
- Define the abort rule: "if partner logins have not succeeded on the new node within ten minutes, fail back and stop." Write it down before you start, so the decision is not made under pressure.
- At T minus five minutes: open the logs of both nodes and the synthetic partner side by side; confirm clocks agree; note which sessions are open.
- At T minus one minute: start the in-flight upload from the synthetic partner.
- At T+0: trigger the scenario, and say "trigger" out loud so the scribe records the time from a clock, not from memory.
- Observe without touching until the synthetic partner logs a successful transfer or the abort time arrives. Record time to detect, address move, and first successful transfer.
- Verify from outside with a real account: log in, upload, list, download, delete. Record the host key or certificate the client reports.
- Swap the mirror direction if the pair uses replication, and confirm the old job is stopped. This is the step most often forgotten.
- Watch for fifteen minutes: partner logins arriving, no authentication or permission errors, no lock errors on the storage, the abandoned partial handled correctly.
- Reconcile files: compare the old node's completed-upload log for the last RPO interval with the new node's folders. Copy anything missing; record the count.
- Decide on failback, and if you fail back, run steps six to eleven again in reverse. Leaving the roles swapped is a legitimate choice and halves the drill's disruption.
- Write the record the same day, while the surprises are fresh.
Remember: the driver never touches anything during step seven. The temptation to "help" the failover along is exactly how you fail to learn whether it works on its own. If it needs help, that is the finding.
The Drill Checklist
The runbook says what to do; the checklist says what must be true at the end. Copy it into the drill record and tick each line with evidence, a log excerpt or a screenshot, not a feeling.
IDENTITY [ ] Host key fingerprint on the new node matches the partner-published fingerprint [ ] Certificate on the new node is the same, unexpired, and chains correctly [ ] Synthetic partner connected with strict checking and no prompts STORAGE [ ] Yesterday's uploads visible on the new node in the same paths [ ] Files completed in the last RPO interval reconciled (count recorded) [ ] In-flight partial handled: temp name only, not processed downstream [ ] Mirror direction swapped and old job stopped (replication pairs) AUTOMATION [ ] Time to detect recorded and within design (e.g. under twenty seconds) [ ] Time to recover recorded and within the partner promise [ ] No flapping: address moved once and stayed [ ] Downstream jobs resumed against the new node without intervention PARTNER VIEW [ ] Synthetic partner: one reset, successful retry, no other errors [ ] FTP/FTPS passive transfer succeeded with no partner-side change EVIDENCE [ ] Log excerpts from both nodes and the synthetic partner attached [ ] Findings listed with an owner and a due date each [ ] Next drill scheduled
Fixing What the Drill Exposes
A drill that finds nothing is either a very mature pair or a shallow drill. Most find two or three of the following, and each has a straightforward fix and a straightforward owner:
- Host key or certificate drift. Copy the correct key or certificate to the drifted node. Then add the fingerprint and expiry comparison to the routine health checks so it cannot drift silently again.
- Mirror never swapped, or swapped late. Make the swap a scripted step with a single command. Make the old job refuse to run when its node does not hold the virtual IP.
- Health check too shallow. The service crash was detected but the disk-full scenario was not: upgrade the check to a real upload.
- Flapping. The address moved twice. Raise the "rise" threshold, or fix the marginal node the check was right about.
- Missing firewall rule or route on the standby. Add it, and put both nodes' rule sets under the same configuration management or mirror.
- Stale lock on shared storage. The new node could not write a file the old node had open. Find the storage setting that expires orphaned locks, and add a runbook step until it is fixed.
- Clocks disagree. Logs could not be lined up. Point both nodes at the same time source and include it in the health check.
- A partner job that does not retry. The synthetic partner reconnected; a real partner's job failed hard. Contact the partner with the log line and the retry wording from the onboarding standard.
- The runbook was wrong. A step referred to a command or path that no longer exists. Fix the runbook now, not at the next drill.
Treat the findings the way you would treat a real incident. List them, give each an owner and a due date, and re-drill once the fixes are in. The format of a blameless review works well even though nothing went wrong for a partner. Our article on running your own postmortem gives a template that needs almost no adaptation.
Recording Results So the Next Drill Starts From Evidence
The drill record is a short document kept with the runbook, one per drill, with the same headings every time. Those headings cover drill number, scenario, nodes and roles, and who drove and who observed. They cover each measurement with the log line it came from, the checklist with evidence, and the findings with owners. They include the date set for the next drill. Keep the numbers in a small table across drills so that a trend is visible. Time to recover creeping up from forty seconds to ninety is a warning that something changed. These records are also the evidence behind any availability promise you make to partners. "Recovery within one hour" backed by six drill records showing under two minutes is a different conversation from the same claim backed by a diagram.
The logs are the raw material, so make sure they survive the drill and are easy to line up. A transfer server can log to a file or a database on each node, as Sysax Multi Server does. That gives you the "first successful login on the new node" timestamp directly. A database shared by both nodes gives it in one query. Reading those lines is a skill to have before the drill rather than during it. Our reading transfer logs article is the primer. And a failover drill is not a restore drill. Rebuilding a node from backup, recovering keys from escrow, and the wider disaster scenario are exercised separately. Our Disaster Recovery for Transfer Workflows series describes that.
The Short Version
A pair you have never failed over is a pair you do not understand. Drill it: planned moves first, simulated failures once those pass, an in-flight upload every time, and a synthetic partner watching from outside. Measure time to detect, time to recover, files lost, and files duplicated. Tick a checklist with evidence, and fix the drift the drill exposes. Keep a record so the next drill, and the next partner conversation, starts from numbers. That is the whole of high availability in practice: the design from the first article in this series, made true by repetition.
Frequently Asked Questions
Is it really safe to run drills on the production pair?
How long should a drill take?
What if the drill fails and partners are affected?
Do I need to fail back after every drill?
What is the single most common thing a first drill finds?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
