Backing Up Transfer Configuration, Keys, and Jobs
"Is the transfer server backed up?" "Yes. Nightly. The job reports success." Both people in that exchange are telling the truth, and the truth has a hole in it. Somebody set the job up years ago, and it has run every night since. On the day you need it you discover what it contains: the operating system volume and the data drive. It contains nothing about the twelve scheduled tasks that ran on the automation host. It has none of the permissions that were painstakingly set on two hundred partner folders, or the passwords those tasks used. The backup was real. It just was not a backup of the service. A green tick is not evidence. It is a mood.
This article turns the scope list from Disaster Recovery Scope: More Than the Server into an actual backup job. It covers what to capture for each part of a transfer platform, in what form, how often, and where to put it. It ends with the two things that separate a backup from a hope. One is a manifest that lets you check the copy is complete. The other is a restore test that proves it can be used. This article is part of our Disaster Recovery for Transfer Workflows series. It is written so that a junior administrator can build the job from scratch on a quiet afternoon. That is the only kind of afternoon to build it on.
A Backup Is for Restoring, Not for Storing
Three words get muddled, so define them. A backup is a copy of data made so it can be restored later. A system image is a backup of an entire disk or machine, restorable as a whole — the thing you use for a bare-metal rebuild. An export is a copy of one component in a portable form. Examples are a scheduled task as an XML file, a certificate as a PFX file, or a registry key as a .reg file. The export can be imported onto a different machine without restoring everything else.
A good transfer DR backup uses all three. The image answers "how do I get a working server back fast?" The exports answer "how do I get this one thing back onto a server that already works?" — the far more common need, and the one image-only plans cannot serve. Plain copies answer "what were the files?"
Whatever its form, a backup is only worth what it restores. The recovery point objective (RPO) was set per flow in the scoping article. It is the maximum age of the newest usable copy. It decides how often each piece must be captured. The restore test at the end decides whether "usable" is true. Until then, "usable" is an adjective, not a finding.
The Backup List for a Transfer Platform
Work through the platform component by component, asking of each: where does it live, what form should the copy take, and what is easy to miss?
Server configuration
Listeners and ports, the passive port range and external address, TLS settings, connection limits, the banner, the root folder, logging destinations. Products store this in one of three places: a configuration folder (often under C:\ProgramData\ on Windows or /etc/ on Linux), the registry, or a database. Find out which — once — and record the exact path in the backup job. Products that run as a Windows service, Sysax Multi Server among them, document where their settings live in the manual; do not guess. If the product offers a configuration export, use it as well as copying the raw files. An export is what you want when restoring one setting onto a healthy server.
For registry-backed settings, the export is one command, run as an administrator:
reg export "HKLM\SOFTWARE\ExampleVendor\TransferServer" D:\Backups\transferserver.reg /y
reg export writes the key and everything under it to a text file that reg import can load on the restored machine. The flag /y overwrites an existing file without prompting. Easy to miss: settings under the service account's user hive rather than the machine hive.
The account store
If accounts are defined inside the transfer product, they are part of its configuration. But check whether password hashes and per-user home folders are in the same place. Some products keep the user list in one file and credentials in another. If accounts come from a directory, you are not backing them up here. In that case, you are recording the dependency and the group names the product maps to. That lets the restored server be pointed at the directory again.
Folder tree and permissions
The directory structure under the transfer root — one folder per partner, inbox and outbox under each — is easy to recreate. The access control lists on those folders are not: they encode months of "partner A may write but not list, partner B may read only." On Windows, icacls saves them to a text file:
icacls D:\Transfer /save D:\Backups\transfer-acls.txt /T /C
/T walks the whole tree, and /C continues past errors instead of stopping at the first folder you cannot read. On restore, the file is applied with icacls D:\ /restore D:\Backups\transfer-acls.txt. Note that the restore is run against the parent of the saved folder, because the saved names are relative to it. On Linux, getfacl -R /srv/transfer > transfer-acls.txt and setfacl --restore=transfer-acls.txt do the same job. Easy to miss: share permissions, which are separate from NTFS permissions, and folder ownership.
Job definitions, scripts, and connection profiles
Every scheduled transfer is three things: a schedule, a script or job definition, and a connection profile holding host, port, account, and credential. They often live on a separate automation host, so this part of the backup runs there, not on the transfer server. On Windows, Task Scheduler tasks export cleanly to XML:
schtasks /Query /TN "\Transfer\PayrollPush" /XML > D:\Backups\tasks\PayrollPush.xml
The XML holds the trigger, the action, the account the task runs as, and its settings. What it does not hold is that account's password: on restore, schtasks /Create /XML PayrollPush.xml /TN "\Transfer\PayrollPush" /RU svc_transfer /RP * will prompt for it. That is a feature, but it means the service account passwords must be recoverable from your password manager or secrets vault. That is exactly where job credentials storage says they belong. Scripts and profiles are plain files under a folder — copy the folder. If the automation product keeps its own job store, find it the same way you found the server configuration. On Linux, crontab -l -u transfer > crontab-transfer.txt captures the schedule for one user; the scripts are wherever the crontab points.
Keys and certificates
Host keys, TLS certificates with their private keys, client keys, partner public keys, PGP keyrings, and tokens. These are backed up on a different schedule (on every change). They are stored with different protection (encrypted, access-controlled, split custody), and restored at a specific point in the runbook. They get their own article: Protecting Keys and Certificates for Recovery. The one rule that belongs here: do not let them ride along, unencrypted, inside the general configuration backup. A configuration backup is read by many people; a key escrow is opened by two.
In-flight data folders
The inbox, outbox, staging, and archive folders are the only part of the platform where the RPO clock ticks by the minute. They are also ordinary files, which means your existing backup system is the right tool. The only decisions are frequency and what to include. Inbound folders deserve the most frequent capture, because their contents cannot be regenerated locally. If partners keep originals, a day-old copy plus the activity log may be enough. Outbound folders are usually regenerable from the upstream system. Archive folders are large and change slowly. Match each folder to its flow's RPO, and write the mapping down. Unwritten, it lives in one head, and heads take holidays.
Logs
Activity logs answer "what arrived after the last backup?" during a disaster and prove you handled it correctly afterwards. Shipping them off the server continuously is cheaper than any RPO you could achieve for the data itself. It is the basis of the catch-up conversation with partners. Our article on centralizing transfer logs covers the mechanics.
The operating system
Finally, the image. Windows Server Backup, a built-in, can take one from the command line once the feature is installed:
wbadmin start backup -backupTarget:\\backup01\wsb\sftp01 -allCritical -quiet wbadmin get versions -backupTarget:\\backup01\wsb\sftp01
-allCritical includes every volume the operating system needs, which makes the result usable for a bare-metal recovery from installation media. The second command lists what is there. Your backup system probably does this already with a nicer interface. The point is that the image is one item on the list, and the items above are separate from it. That is because an image is slow to restore, all-or-nothing, and taken at most nightly. Restoring an image to recover one setting is moving house to fetch a book.
How Often, and Where
The table sets a frequency and a form for each item. Adjust the frequencies to your own RPOs; the shape of the table is what matters.
| Item | Form | How often | Where |
|---|---|---|---|
| Server configuration, account store | Copy of config tree + product export + registry export | Nightly, and after every change | Dated snapshot folder, then offsite |
| Folder tree and permissions | icacls /save file |
Nightly | With the config snapshot |
| Jobs, scripts, profiles, schedules | Task XML exports + copy of scripts folder | Nightly, and after every change | With the config snapshot (taken on the job host) |
| Keys, certificates, tokens | Encrypted escrow archive | On every change or rotation | Separate: offline copy plus offsite copy |
| Inbound data folders | File backup or storage snapshot | Per the flow's RPO — often hourly | Backup system, offsite copy |
| Outbound and archive folders | File backup | Nightly | Backup system |
| Activity logs | Shipped to a collector | Continuously | Off the server |
| Operating system | System image | Nightly or weekly | Backup system, offsite copy |
"Where" follows one old rule: three copies, on two kinds of storage, one of them somewhere else. The snapshot folder on the server's own disk is a convenience for restoring a single setting — not a backup, because it dies with the server. The copy on your backup system is the working backup. The offsite copy is what DR actually uses, and it must be reachable from the recovery site when the primary site is gone. A scheduled task in Sysax FTP Automation is one straightforward way to push each night's dated snapshot folder to an offsite SFTP server. It can notify you by email if the push fails.
Remember: configuration backups contain secrets — connection profiles, password hashes, sometimes plaintext credentials in old scripts. Encrypt the archive before it leaves the server, restrict who can read the backup share, and keep the real key material out of it altogether.
The Backup Job, Concretely
The script below is a complete nightly job for a Windows transfer platform. It builds a dated snapshot folder, captures every item above except the operating system image and the key escrow, and finishes with a manifest. Adjust the paths; keep the structure.
# backup-transfer-config.ps1 — nightly snapshot of the transfer platform
$stamp = Get-Date -Format "yyyyMMdd"
$dest = "D:\Backups\transfer\$stamp"
New-Item -ItemType Directory -Path $dest -Force | Out-Null
# 1. Product configuration tree, keeping timestamps and NTFS security
robocopy "C:\ProgramData\TransferServer" "$dest\config" /E /COPY:DATSO /R:2 /W:5 /NP /LOG:"$dest\robocopy-config.log"
if ($LASTEXITCODE -ge 8) { throw "robocopy (config) failed: exit code $LASTEXITCODE" }
# 2. Registry-backed settings, if the product uses them
reg export "HKLM\SOFTWARE\ExampleVendor\TransferServer" "$dest\transferserver.reg" /y | Out-Null
# 3. Permissions on the data tree (ACLs only, not the files)
icacls "D:\Transfer" /save "$dest\transfer-acls.txt" /T /C | Out-Null
# 4. Scheduled tasks under the \Transfer\ folder, one XML file each
New-Item -ItemType Directory -Path "$dest\tasks" -Force | Out-Null
Get-ScheduledTask -TaskPath "\Transfer\" | ForEach-Object {
Export-ScheduledTask -TaskName $_.TaskName -TaskPath $_.TaskPath |
Out-File "$dest\tasks\$($_.TaskName).xml" -Encoding unicode
}
# 5. Scripts and connection profiles, excluding their log files
robocopy "D:\Jobs" "$dest\jobs" /E /COPY:DAT /XF *.log /R:2 /W:5 /NP /LOG:"$dest\robocopy-jobs.log"
if ($LASTEXITCODE -ge 8) { throw "robocopy (jobs) failed: exit code $LASTEXITCODE" }
# 6. Manifest: one SHA-256 line per file, relative paths, sha256sum format
Get-ChildItem -Path $dest -Recurse -File |
Where-Object { $_.Name -ne "MANIFEST.sha256" } |
Get-FileHash -Algorithm SHA256 |
ForEach-Object { "{0} {1}" -f $_.Hash.ToLower(), $_.Path.Substring($dest.Length + 1) } |
Out-File "$dest\MANIFEST.sha256" -Encoding ascii
Write-Output "Snapshot complete: $dest"
What each step does, and why it is written that way:
- Step 1 uses
robocopy /Eto copy the tree including empty folders./COPY:DATSOcopies data, attributes, timestamps, security descriptors, and owner — a plain copy would lose the last two./R:2 /W:5retries a locked file twice, five seconds apart, instead of the default of a million retries. The flag/NPkeeps progress percentages out of the log. Robocopy's exit code is a bitmask; anything below 8 means success, so the script only fails on 8 or higher. Our robocopy for migrations article explains the flags in depth. - Step 3 captures permissions separately, so the tree's ACLs can be restored even if the files come back from a different source.
- Step 4 exports every task under one Task Scheduler folder; keeping transfer jobs in their own folder is what makes this a one-liner. Run this step on the job host if that is a different machine.
- Step 6 writes the manifest in the two-space format
sha256sumuses, so a Linux recovery host can verify it withsha256sum -c MANIFEST.sha256and no special tooling.
Schedule it under an account that can read the configuration tree, and have it fail loudly. A backup job whose failures nobody sees is the most dangerous object in the plan. I have inherited two of those and, in an earlier job, written one. The Linux shape is the same. Use rsync -a for the trees, crontab -l for schedules, and getfacl -R for permissions. Use find . -type f -print0 | sort -z | xargs -0 sha256sum > MANIFEST.sha256 for the manifest.
The Manifest: Knowing the Copy Is Complete
A manifest is a list of what a backup should contain, with a hash of each file. A hash is a short fingerprint computed from a file's contents; change one byte and the fingerprint changes. The manifest answers two questions the backup job's "success" message cannot: is every expected file present, and is each one intact? The generated file looks like this:
3b7f1c...e920 config\server.conf a91d44...0c31 config\users\users.db 5e02ab...77f4 transferserver.reg c4d9e0...b118 transfer-acls.txt 0f6a3d...45de tasks\PayrollPush.xml 0f6a3d...45de tasks\ClaimsPull.xml 7c11b2...9a06 jobs\payroll\push.ps1 7c11b2...9a06 jobs\payroll\payroll.conf
A manifest catches two things nothing else does. Silent omission: if the task export produced zero files because the Task Scheduler folder was renamed, the manifest has no tasks\ lines. Comparing it with yesterday's shows that omission. Silent corruption in transit: verifying the offsite copy against the manifest proves every byte arrived. The general technique is in checksum files and manifests. For DR, add one habit — keep a copy of the manifest outside the snapshot. That lets you tell whether the snapshot itself is complete. A manifest inside the snapshot vouches for itself, which is not vouching.
Bluewater Bank saw the value of that habit within a month of adopting it. An administrator tidied the Task Scheduler library and moved the transfer jobs from \Transfer\ into \Transfer\Prod\, which was neater. The nightly backup kept reporting success, because exporting zero tasks is not an error. The Monday comparison of manifests showed every tasks\ line had vanished. The export path was corrected before lunch, and the tidy-up became a change ticket, retroactively. Nothing was lost. The manifest was the reason nothing was lost.
The Restore Test: A Backup Versus a Hope
Everything above produces files. None of it proves those files can become a working service. The restore test does. Take the newest snapshot, apply it to a scratch machine that is not production, and see whether the result behaves like the original. If you already have a staging environment, it is the natural scratch machine. A small, regular test beats a large, rare one — monthly, on a throwaway virtual machine, in under an hour. The first one we ran was humbling, and I mean that as a recommendation.
The sequence for a configuration-level test:
- Build or clone a scratch VM with the operating system and the transfer product installed, on an isolated network with no route to partners.
- Copy the snapshot folder across and verify it against the manifest. Any mismatch stops the test — that is a finding, not an inconvenience.
- Stop the transfer service. Copy the configuration tree back to its live location. Import the registry file with
reg import. Apply the ACL file withicacls /restoreagainst the parent folder. - Import one scheduled task from its XML and supply the service account password from the vault. Leave it disabled.
- Start the service. Confirm it listens on the expected ports, that the user list matches production, and that a test account can log in and see its home folder.
- Write down how long each step took and everything that surprised you.
The surprises are the point. A typical first-run finding is that the configuration references a certificate thumbprint that is not on the scratch machine. That is expected — keys are restored separately, and the test shows exactly where the runbook must do it. Another finding might be a script with a hard-coded drive letter the scratch VM lacks. Or the task import fails because the service account does not exist locally. Each is a plan defect found for free. The full-scale version of this exercise, restoring the whole platform against the clock, is Restore Drills: Proving the Plan Works.
Rule of thumb: a backup that has never been restored is a hope. A backup restored once, a long time ago, is a memory. Only a backup restored recently, on a machine that was not production, from the copy you would actually use, is a backup.
Putting It Together
A transfer platform's backup is not one job but a set. It includes an image for the machine, exports for the components, and plain copies for the data. It includes a shipped stream for the logs and a separately protected escrow for keys. Each has its own frequency, driven by the RPO of the flows that depend on it. Each has its own place — the offsite copy being the one DR actually uses. A manifest says the set is complete; a restore test says it works. Both are skipped more often than either deserves.
From here, the key escrow deserves its own reading in Protecting Keys and Certificates for Recovery. The order in which all of these pieces go back onto a rebuilt server is the subject of The Transfer Disaster Recovery Runbook. If the scheduled-task side of your platform is the least familiar part, Task Scheduler for transfers is the companion read.
Frequently Asked Questions
Isn't a nightly system image enough to back up a transfer server?
Why doesn't my exported scheduled task include its password?
How do I know the backup actually contains everything?
How often should I back up transfer configuration?
Should keys and certificates go in the same backup as the configuration?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
