Making Redundancy Invisible to Partners
A redundant transfer service is only as good as what the partner's script sees. That script was written once, months or years ago, by someone who is no longer available. It has one hostname, one host key fingerprint, one username, and one folder path hard-wired into it. It does not know that you have two nodes, or three, or that you moved an address last night. If any of those hard-wired facts stops being true during a failover, the partner's job fails. From their side your "highly available" service just went down.
This article is about the partner's side of the design: what their client actually checks on every connection. It covers how to make every node present exactly the same identity, and what to do about DNS. It covers which protocols need a session to stick to one node, and what a partner sees during the seconds of a failover. It covers how to test all of it from the partner's chair rather than your own. This article is part of our High Availability for Transfer Services series and applies equally to the active-passive pair and the active-active cluster.
What the Partner's Client Actually Checks
Every automated transfer job, whatever protocol it speaks, goes through the same short sequence before a single byte of file data moves. Each step compares something the client stored with something your server presents, and any mismatch stops the job. Listing them is the whole design brief for partner-transparent redundancy.
| What the client checks | Where it is stored on the partner side | What happens if a node differs |
|---|---|---|
| Hostname resolves to an address | Script or connection profile, plus their DNS cache | Connection refused or timeout until the cache expires |
| Address is allowed by their firewall | Their outbound firewall rules | Silent timeout; a ticket to their network team |
| SSH host key fingerprint (SFTP) | Their known_hosts file or client key cache | "Host key has changed" and a hard failure |
| Certificate name and chain (FTPS, HTTPS) | Their trust store, sometimes a pinned fingerprint | TLS verification failure |
| Username and credential | Their credential store | Authentication failure, possibly a lockout |
| Folder path and permissions | Their script | "No such file" or "permission denied" |
| Passive-mode address and port range (FTP, FTPS) | Their firewall rules for the data connection | Login works, every transfer hangs |
The design goal is that every row is identical whichever node answers. Each of the sections below takes one or two rows and shows how to make that true.
One Hostname, One Address
Partners must know exactly one name, and that name must resolve to exactly one address that is valid no matter which node is serving. In an active-passive pair the address is the virtual IP, the floating address that moves between nodes. In an active-active cluster it is the address the load balancer owns. Either way, the node's own addresses, 10.20.0.51 and 10.20.0.52, are never given to a partner. They never appear in an onboarding document or in a passive-mode reply. The moment a partner has a node address, they will allowlist it, script it, and be broken by the first failover.
The same rule applies in the other direction. If your service also connects outbound to partners, to push files or fetch them, the partner sees your source address. Their firewall allows only the one they were told. Make sure outbound connections from either node leave through the same fixed address, typically by source NAT at your edge firewall. That way, a failover of the pushing job does not change the address the partner sees.
DNS strategy
DNS is where good intentions go to expire slowly. The name sftp.example.com should have one A record pointing at the virtual IP, with a moderate time-to-live. That is the number of seconds a resolver may cache the answer. Five minutes is a reasonable value: short enough that a deliberate address change propagates within a coffee break. It is long enough that partners' resolvers are not hammering your DNS servers. Do not publish the node addresses under any name a partner might discover. Do not publish two A records for the service name hoping clients will "try the other one." Most clients try the first answer they get and report a failure.
Understand, too, that TTL is a suggestion the whole path does not always honor. Some corporate resolvers cache for longer than the record says. Some runtime environments cache a lookup for the life of the process. In those environments, a long-running partner job that resolved your name at nine in the morning keeps that answer until it is restarted. For this reason, same-site failover should never depend on DNS: the virtual IP moves, the name does not change, and no cache anywhere goes stale. DNS changes belong to cross-site recovery, where the second site necessarily has different addresses. That scenario is part of our Disaster Recovery for Transfer Workflows series, and its partner communication is a different exercise.
Identical Host Keys on Every Node
An SFTP server identifies itself with an SSH host key, a key pair generated once at installation. The client stores the public half, or its fingerprint, the first time it connects. On every later connection it refuses to proceed if the server presents a different key. A different key is exactly what an impersonating server would present. Our host keys and known_hosts article explains the mechanism. Here the operational rule is short: every node presents the same host key. Not equivalent keys, not keys of the same type, the same key.
That means copying the private host key file from the first node to every other node over an authenticated channel. It means setting the same restrictive file permissions and restarting the service. It also means making sure nothing regenerates the key later. That could be a rebuild from an image or an installer that helpfully creates a fresh key. Or it could be a well-meaning hardening script that "rotates" it on one node. Verify from outside, against each node's own address, and compare:
# fetch the host key each node presents and print its fingerprint
for node in 10.20.0.51 10.20.0.52 10.20.0.53; do
printf '%s ' "$node"
ssh-keyscan -t ed25519 -p 22 "$node" 2>/dev/null | ssh-keygen -lf -
done
10.20.0.51 256 SHA256:Qm9Vf3kLx0rP8sT1uWyZaBcDeFgHiJkLmNoPqRsTuVw 10.20.0.51 (ED25519)
10.20.0.52 256 SHA256:Qm9Vf3kLx0rP8sT1uWyZaBcDeFgHiJkLmNoPqRsTuVw 10.20.0.52 (ED25519)
10.20.0.53 256 SHA256:Qm9Vf3kLx0rP8sT1uWyZaBcDeFgHiJkLmNoPqRsTuVw 10.20.0.53 (ED25519)
Three identical fingerprints is the answer you want. Run the same loop for every key type the service offers, since a client may have stored the RSA fingerprint rather than the newer one. Put the loop in the health checks so that a drifted key is caught by you, within minutes, rather than by a partner. When the day comes to rotate the host key deliberately, rotate it on every node in the same maintenance window. Give partners the new fingerprint in advance. The procedure is in key rotation and inventory.
Identical Certificates on Every Node
FTPS and HTTPS clients check a certificate instead of a host key. The server presents one, and the client verifies that it was issued for the hostname it connected to. The client checks that it chains to an authority the client trusts, and that it has not expired. Some partners go further and pin the certificate's fingerprint. Either way, the rule is the same as for host keys. Install the same certificate and private key on every node, issued for the service hostname, with the same intermediate chain. Our installing and chaining certificates article covers the mechanics on one server. In a cluster, repeat them on each node and then verify from outside.
# fingerprint of the certificate each FTPS node presents (explicit FTPS on port 21)
for node in 10.20.0.51 10.20.0.52; do
printf '%s ' "$node"
openssl s_client -connect "$node":21 -starttls ftp -servername ftps.example.com </dev/null 2>/dev/null \
| openssl x509 -noout -fingerprint -sha256
done
Renewal is where clusters break. A certificate renewed on the active node and forgotten on the standby is a certificate that expires on the standby. The expiry will be discovered on the morning after a failover. Treat renewal as a cluster-wide task with a checklist. Monitor expiry on every node's own address, not just on the virtual IP, which only ever shows you the active node. The monitoring side is in certificate expiry monitoring.
Identical Accounts and Folder Views
A partner's username, password or public key, home folder, and permissions must be the same on every node. A change on one node must reach the others before the next failover. There are two ways to get there. Either every node reads the same account store, a directory service or a shared database, so there is nothing to replicate. Or the account configuration is replicated as part of the configuration mirror described in the active-passive article. That schedule must be short enough that a password reset at ten past two is on the standby by a quarter past. The second approach has a trap. A partner may change their own password on the active node and then be failed over to a standby that still has the old one. If partners can self-service credentials, replicate on change, not on a timer, or use a shared store. The wider lifecycle is in our partner credential lifecycle article.
Folder views must match just as exactly: the same path, the same drive letter or mount point, the same permissions. And, above all, the files must be the same. That is the storage question answered in shared storage vs replication. A partner whose script does cd /inbound/acme must find that folder on every node, and must find yesterday's upload still inside it.
Session Affinity Needs, Protocol by Protocol
Session affinity is the requirement that every connection belonging to one partner session reaches the same node. It only matters when several nodes serve at once, and it matters differently for each protocol:
- SFTP: none needed. One TCP connection carries the whole session, so whichever node accepts it keeps it. A partner opening several parallel sessions may land on several nodes, which is fine on consistent storage.
- FTP and FTPS: needed. The data connection for each transfer must reach the node that issued the passive reply. Give each node its own slice of the passive range and forward by port, or use source-address affinity. The reasons are in our article on FTP load balancers and proxies. Whatever the method, every node must announce the same external address, the virtual IP. A server may let you set the announced address and the passive range per node, as Sysax Multi Server does. That makes it a matter of typing the same values into each node's settings.
- HTTPS upload portals: usually needed. A browser or client logs in and receives a session token. If the next request lands on a node that has never seen that token, the user is logged out mid-upload. Either the balancer keeps each client on one node, or the nodes share their session store.
In an active-passive pair none of this applies, because only one node ever serves. Partners' passive-mode firewall rules keep working because the address and range are the same on the standby.
What a Partner Sees During a Failover
Here is a failover from the partner's chair, using the same twelve-second detection window as elsewhere in this series. Their scheduled job is uploading a file at ten past two in the morning.
Mar 14 02:10:04 Uploading orders_0314.csv (41% complete) Mar 14 02:10:19 ERROR: Connection reset by peer Mar 14 02:10:19 Transfer failed; will retry in 30 seconds (attempt 1 of 5) Mar 14 02:10:49 Connecting to sftp.example.com (10.20.0.50)... Mar 14 02:10:50 Host key matches known_hosts entry Mar 14 02:10:50 Authenticated as acme_out (public key) Mar 14 02:10:50 Uploading orders_0314.csv (0% complete) Mar 14 02:11:13 Transfer complete
Notice what did not happen. There was no "host key has changed" prompt, because the standby presented the same key. There was no authentication failure, because the account was there. There was no folder error, because the path existed. The only visible event was one reset connection and one retry, and the retry succeeded without a human. This is the outcome to design for, and it depends on the partner's side too. The client retries with a short pause. It uploads to a temporary name and renames on completion so the abandoned 41 percent can never be mistaken for a delivery. It pins the host key rather than accepting whatever it is shown. Those three habits belong in your onboarding standard. Our articles on retry strategies and backoff and temp names and atomic renames give the partner-facing wording. And partner exchange standards shows how to make them the default rather than a request.
Remember: a failover is invisible to a partner only if their client retries. A job that fails hard on the first reset will report an outage no matter how good your design is. Ask for retry behavior at onboarding, and test for it during drills.
Communicating Maintenance
Even a design that makes failover invisible deserves a maintenance notice, for two reasons. Some partners' monitoring will see the reset and want to know it was expected. And some partners have change-freeze windows during which any interruption, however brief, is a contractual event. Keep the notice short and concrete. Say when, in a stated time zone. Say what will happen ("connections may drop once and reconnect within a minute"). Say what will not change ("hostname, address, host key, certificate, and credentials are unchanged"). Say whom to contact. Send it far enough ahead for the partner to reschedule a job, typically a few business days. If your agreements define notice periods, follow them. Our article on partner SLAs and expectations shows how such clauses are usually written. It is worth writing the "no notice required for interruptions under one minute" clause into new agreements once your drills prove you can deliver it.
Testing From the Partner's Chair
Everything above can be verified from your own desk, and none of it should be. Your desk is inside your network, with your DNS and your firewall rules. A partner sits outside. Build a synthetic partner: a small job that runs from a network you do not control. It could run on a low-cost virtual machine at another provider or a machine at a branch office. The job behaves exactly like a real partner would.
- Resolve the service name with the public DNS, not an internal server, and record the address returned.
- Connect and verify identity. For SFTP, connect with strict host-key checking against a known_hosts entry you prepared in advance. For FTPS or HTTPS, verify the certificate against the public chain, with no "accept anything" flags.
- Log in with a dedicated probe account that has the same kind of folder and permissions a real partner has.
- Upload a file to a temporary name, rename it, download it back, compare, delete. For FTP and FTPS, use passive mode so the data connection crosses your firewall exactly as a partner's would.
- Record the time each step took and log a single pass or fail line, so a week of results is readable.
- Run it every few minutes, always, and keep its history. Then, during every failover drill, watch it: it is the only measurement of the failover a partner would agree with.
Run the same probe against each node's own address as well as the virtual IP, from inside the network. That way, a drifted host key or an expired certificate on the standby is found before the standby is needed. A scheduled-transfer tool with retry and reporting built in, such as Sysax FTP Automation, can run the external probe as an ordinary scheduled job. It can alert on failure, which is exactly the behavior a real partner's job would have. The next article is failover drills. It covers how to fit the probe into a full drill, and what to measure.
The Short Version
Partners connect to one name that resolves to one address. Every node behind it must present the same host key, the same certificate, the same accounts, and the same folders. That way, the only thing a partner ever sees of a failover is one reset connection followed by a successful retry. Same-site failover uses a virtual IP so DNS caches never matter. FTP and FTPS need session affinity in a cluster and SFTP does not. The entire design is verified by a synthetic partner running from outside your network. The pieces are described in the rest of this series. The pair is in active-passive failover, and the storage is in shared storage vs replication. The proof is in failover drills.
Frequently Asked Questions
Is it safe to use the same SSH host key on several servers?
Should I tell partners the addresses of the individual nodes?
Why can't I just lower the DNS TTL and switch the record on failover?
Will a partner's upload resume after a failover?
How do I know the standby's certificate is still valid if nobody ever connects to it?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
