Writing Robust Python Transfer Scripts
There are two versions of every transfer script. The first works: you run it at your desk, files move, everyone is pleased. The second version runs at two in the morning while the network flickers, a partner's server reboots mid-session, and nobody is watching. It either handles all of that gracefully or becomes the mystery you reconstruct from fragments the next day. The distance between those two versions is not cleverness. It is a small set of disciplines, applied consistently.
This article is those disciplines for Python: a failure model to design against and exception handling that catches narrowly instead of hopefully. It covers timeouts on everything that can wait forever, retries that back off and give up, and logging that answers questions. It covers exit codes a scheduler can act on, and configuration kept out of the code. Each one is a few lines; together they are the difference between a demo and a production job. This is part of our Python Automation series and applies equally to the SFTP and FTP/FTPS scripts built earlier in it.
Start with a Failure Model
Robustness starts before any try is typed, with a question: what exactly can go wrong here? For a transfer job the honest list is short and knowable. Name resolution can fail. The connection can time out or be refused. Authentication can be rejected. The remote directory can be missing. The disk — either end — can fill. A file can be half-written when you read it. And the same file can arrive twice.
Every entry on that list falls into one of two classes, and the classification drives everything else in this article. A transient failure is weather: a timeout, a reset connection, a server momentarily too busy. Waiting and retrying is the correct response. A permanent failure is geology: a wrong password, a missing directory, a malformed file. No number of retries fixes geology — retrying just delays the alert, floods the log, and sometimes makes things actively worse. The full taxonomy, with real error messages sorted into each bucket, is the opening article of our retry and error handling series. The working version fits in a table:
| Failure | Class | Right response |
|---|---|---|
| Timeout, connection reset, host briefly unreachable | Transient | Retry with backoff, capped attempts |
| Authentication rejected | Permanent | Stop and alert — retries can trigger lockouts |
| Remote directory missing | Permanent | Stop — configuration or server change |
| Disk full, either side | Permanent (until fixed) | Stop and alert loudly |
| One file of many fails | Either | Record it, continue, report a partial failure |
| Same file delivered again | Design question | Make reruns safe — see idempotency |
Exceptions Done Right
Python reports failures as exceptions — typed objects that interrupt execution until something catches them. The type is the information: a timeout error and an authentication error are different classes, which is exactly what lets code treat weather and geology differently. Three rules produce exception handling you can trust.
Catch narrowly. An except clause should name the exceptions it expects and knows how to handle. That means the transfer library's error classes, OSError for file and network trouble — and nothing more. The notorious anti-pattern is the bare except:, which catches everything, including the exception raised when someone tries to stop the script. Its slightly more respectable cousin except Exception still converts every surprise — a typo three functions down, a missing config key — into "transfer failed, retrying." That buries the real bug. Reserve broad catches for one place: the top of the program, where the goal is to log whatever happened and exit nonzero.
Catch where you can act. If a function cannot do anything useful about a failure, it should not catch it — let it travel upward to the code that can decide. In a transfer job that usually means catching per-file errors inside the file loop (record, continue) and connection errors in main() (distinct exit code, stop).
Use the whole statement. try has four parts, and transfer scripts benefit from all of them:
results = {}
for path in batch:
try:
upload(sftp, path) # only the risky operation
except TRANSIENT_ERRORS as exc:
results[path.name] = f"transient: {exc}"
except PERMANENT_ERRORS as exc:
results[path.name] = f"permanent: {exc}"
else:
results[path.name] = "ok" # runs only when no exception
archive(path) # success-only follow-up work
# ...meanwhile, at the level that owns the connection:
try:
run_batch()
finally:
client.close() # runs no matter what happened
The else block holds work that must happen only on success — moving the source file to an archive, updating a ledger. This work is kept outside the try so its own bugs are never mislabeled as transfer failures. The finally block holds cleanup that must happen regardless: closing connections, removing temp files, releasing locks.
One reading skill completes the section. When an uncaught exception does escape, Python prints a traceback — the chain of calls that led to the failure. Read it from the bottom up. The last line names the exception and its message; the lines above it walk backward through the code that was executing. The bottom two or three lines answer "what failed and where" for almost every transfer-script incident. That is why the top-level handler shown later logs the full traceback rather than just the message.
Timeouts Everywhere
The default behavior of most network code is to wait forever, and forever is the worst possible failure mode for an unattended job. It produces no error, no log line, and no exit code. You just have a script that is still "running" at breakfast, blocking tonight's run behind it. The fix is mechanical. Every constructor in this series takes a timeout, so give every one a value. That includes FTP(host, timeout=30) and its FTPS sibling, and the SSH client's timeout=, banner_timeout=, and auth_timeout=. For any library that forgets to offer one, socket.setdefaulttimeout(30) sets a process-wide backstop before the first connection is made.
Per-operation timeouts still allow a job to crawl — a hundred files, each taking just under the limit — so give the whole run a deadline too:
import time
deadline = time.monotonic() + 25 * 60 # this run gets 25 minutes, total
for path in batch:
if time.monotonic() > deadline:
log.error("deadline exceeded with %d files left", remaining_count)
break
transfer_one(path)
time.monotonic() is the right clock here — it only moves forward, immune to the wall-clock adjustments that make timestamp math lie. Checking between files keeps the logic simple and the stop clean: no file is abandoned midway, and the temp-name pattern protects the one in flight. Belt and braces: many schedulers can also kill a task that exceeds a time limit. Set that limit too, generously, as the backstop for the day your deadline logic itself has the bug.
Retries with Backoff
Now the transient column of the failure model gets its machinery. The rules of a well-behaved retry: retry only transient failures. Wait longer after each attempt (exponential backoff). Add randomness (jitter) so a fleet of jobs does not hammer a recovering server in unison. Cap the attempts, and when the cap is reached, fail loudly. As a helper:
import random
import time
TRANSIENT_ERRORS = (TimeoutError, ConnectionError, EOFError)
# extend per protocol, e.g. ftplib.error_temp or the SSH library's SSHException
def with_retries(operation, attempts=4, base_delay=2.0):
"""Call operation(); on transient failure, back off and try again."""
for attempt in range(1, attempts + 1):
try:
return operation()
except TRANSIENT_ERRORS as exc:
if attempt == attempts:
raise # out of budget: fail loudly
delay = base_delay * 2 ** (attempt - 1)
delay += random.uniform(0, delay / 2) # jitter
log.warning("attempt %d/%d failed (%s); retrying in %.0fs",
attempt, attempts, exc, delay)
time.sleep(delay)
Walk the numbers: with a base delay of two seconds, waits run roughly 2, 4, then 8 seconds plus jitter. That is three retries inside about twenty seconds, which rides out the vast majority of network blips without stretching the job. Anything a twenty-second window cannot ride out is an outage. The correct handler for an outage is the next scheduled run, not a loop that retries into the night. That is also why the budget stays small: a job that retries forever is indistinguishable from a job that hung.
Note what the helper retries: operation, one step — a single upload, one listing call — not the entire job. Retrying at the smallest sensible unit means attempt three does not redo work attempts one and two finished. And notice what it never retries: nothing in PERMANENT_ERRORS. Hammering a server with a rejected password does not fix the password. It can trip the account-lockout defenses described in authentication failures and lockouts. That can turn a five-minute credential fix into a locked service account at 2 a.m.
Remember: a retry is a bet that time fixes the problem. Timeouts and resets — take the bet, briefly. Rejected credentials, missing paths, full disks — fold immediately and alert a human.
Logging Without Ceremony
When the 2 a.m. run fails, the log is the only witness. print() is not logging — under a scheduler its output lands wherever standard output happens to point, without timestamps or severity. The standard library's logging module fixes all of that in four lines of setup. Ignore its reputation for complexity, which comes from features you do not need yet:
import logging
logging.basicConfig(
filename="/var/log/jobs/nightly-push.log",
level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s",
datefmt="%b %d %H:%M:%S",
)
log = logging.getLogger("nightly-push")
Call it once at the top of main(), then log everywhere with log.info(...), log.warning(...), log.error(...). What to record is a short list. Record the start of each run with its parameters (never its secrets) and each file's outcome with sizes. Record every retry with its reason, and a final summary line with totals. The wider discipline of which events deserve a line is covered in what to log; a good run's output reads like a story told by someone terse:
Mar 14 02:10:01 INFO starting nightly-push: 14 files queued Mar 14 02:10:09 WARNING attempt 1/4 failed (timed out); retrying in 2s Mar 14 02:10:14 INFO uploaded orders.csv (184320 bytes) Mar 14 02:10:19 ERROR permanent failure for legacy.csv: 550 No such directory Mar 14 02:10:26 INFO finished: 13 uploaded, 1 failed
Two habits complete the picture. First, never log credentials — not passwords, not keys, not full connection strings. Logs get copied into tickets and chat threads, and secrets in logs outlive every rotation. Second, remember your log is only the client's testimony. The server keeps its own account of the same session. A server such as Sysax Multi Server writes activity logs of every login and file operation. When the two accounts disagree, the comparison itself is usually the diagnosis.
Exit Codes the Scheduler Can Read
The scheduler that launches your script cannot read its log. It sees exactly one datum when the process ends: the exit code, an integer where zero means success and anything else means failure. Cron decides whether to send mail based on it; Task Scheduler records it as the Last Run Result; monitoring dashboards color green or red by it. A script that swallows its errors and exits zero is invisible to every one of those layers — the failure simply does not exist. (How schedulers consume these codes is part of our scheduled jobs series.)
The pattern that makes codes deliberate: main() returns them, one place maps them, and a top-level catch guarantees even the unexpected exits nonzero:
import sys
# 0 success | 1 partial: some files failed | 2 bad config or usage
# 3 authentication refused | 4 endpoint unreachable
def run():
try:
sys.exit(main())
except Exception:
log.exception("unhandled error") # full traceback into the log
sys.exit(10)
if __name__ == "__main__":
run()
Keep the map tiny — five or six values, documented in a comment at the top — and keep it stable, because monitoring rules will grow around it. The one unforgivable code is a dishonest zero.
Configuration Out of Code
Every hard-coded hostname is a future code change wearing a disguise. Robust scripts split three things that beginners fuse together. Code is logic, in version control. Configuration — hosts, paths, patterns, schedules — lives in a config file or arguments, so the same script serves many flows. Secrets live in neither. They come from the environment set by the scheduler, from key files with tight permissions, or from a credential store, following the rules in service account hygiene. A config file that contains a password is a secret with bad storage habits.
The last piece of the discipline is a rehearsal switch. Give every transfer script a --dry-run flag that reports what the run would do — which files, where to — and changes nothing:
import argparse
parser = argparse.ArgumentParser(description="Nightly SFTP push")
parser.add_argument("--config", default="/etc/jobs/nightly-push.ini",
help="path to the job's config file")
parser.add_argument("--dry-run", action="store_true",
help="log what would transfer; change nothing")
args = parser.parse_args()
A dry run is how you test a config change at 3 p.m. instead of discovering it at 3 a.m. It turns the first deployment of any new flow into a rehearsal instead of a premiere. The full worked version of this pattern — config file, CLI, dry-run, and packaging — is the subject of building a small transfer utility.
The Robustness Checklist
Before a Python transfer script earns a schedule, walk it past this list:
- Every network constructor has an explicit timeout; the run has an overall deadline.
- Every
exceptnames specific exceptions — the only broad catch is at the top, and it exits nonzero. - Transient failures retry with backoff, jitter, and a small attempt cap; permanent failures stop and alert.
- Connections close in
finally; temp files are cleaned or clearly named. - Uploads and downloads land under temp names and are renamed only when complete.
- Logging is configured with timestamps and levels; secrets never appear in it.
- The run ends with a summary line: attempted, succeeded, failed.
- Exit codes are documented and honest — partial failure is not a zero.
- Configuration is in a file or arguments; secrets are in neither code nor config.
- A
--dry-runflag exists, and someone has actually used it against production config.
It is fair to notice that this list is also a build-versus-configure decision in disguise. Every item is code you must write and own. For standard flows on Windows, a tool like Sysax FTP Automation ships the retry and error handling, scheduling, folder monitoring, and email notification layers. These are configuration rather than code. The checklist earns its keep on the flows whose logic is genuinely yours.
Where to Go Next
Robustness is a habit, not a feature — ten small disciplines that turn "it worked when I ran it" into "it runs." Apply them to the scripts from the SFTP and FTP/FTPS articles. Then take the next two steps. The article on building a small transfer utility assembles these pieces into a tool a coworker can run. The article on deploying and scheduling Python transfer jobs gets the result running unattended without the environment surprises that undo everything this article built.
Frequently Asked Questions
Is it ever OK to catch every exception?
How many retry attempts should a transfer job make?
Do I really need the logging module for a fifty-line script?
What exit code should a partial failure return?
Where do the credentials go if not in the script?
Should the script send its own email alerts when it fails?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
