Home › Topics › Python Automation › SFTP in Python

SFTP Transfers in Python, Step by Step

SFTP is the default protocol for secure, scripted file movement: it encrypts everything, authenticates with SSH keys, and needs exactly one firewall-friendly connection. So the first surprise for anyone scripting it in Python is that the standard library — famous for including batteries — has no SFTP support at all. The second surprise is subtler: the most popular shortcut for getting a script connected quietly disables the security check that makes SSH trustworthy in the first place.

This article walks the whole path, step by step. It covers which library fills the gap and how to install it without making a mess. It covers how to handle the host key question honestly in a job that runs with nobody watching. It covers how to authenticate with a key instead of a password, and the upload and download operations themselves. It includes a complete small script that puts every piece together. This article is part of our Python Automation series. If you are still weighing whether the job belongs in Python at all, the series opener on when Python beats shell scripts settles that first.

The Standard Library Gap, and the Library That Fills It

Python ships ftplib for FTP and FTPS — covered in the companion article — but nothing for SSH. SFTP is not a standalone protocol: it runs as a subsystem inside an SSH connection, inheriting SSH's encryption and authentication. (If that layering is fuzzy, our explainer on how SFTP works is the five-minute prerequisite.) Implementing SSH is a serious cryptographic undertaking, which is why it lives outside the standard library.

The long-standing answer in the Python world is Paramiko, a pure-Python SSH implementation that has been the de-facto standard for many years. Several friendlier wrapper libraries exist on top of it, but this article drives it directly, for a reason worth stating. The wrappers tend to hide exactly the decisions — host keys, authentication sources, timeouts — that an unattended transfer job must get right. Once you can make those decisions deliberately, use whatever convenience layer your team prefers.

Install it into a virtual environment — a private, per-job copy of the Python package space. Avoid installing it into the system-wide interpreter, where an unrelated upgrade could break your job later:

python3 -m venv /opt/jobs/reports/.venv
/opt/jobs/reports/.venv/bin/python -m pip install paramiko
/opt/jobs/reports/.venv/bin/python -c "import paramiko; print('ok')"

The third line is a smoke test: if it prints ok, the library is importable by that specific interpreter. Everything about environments, pinning, and invoking this interpreter from a scheduler is covered in deploying and scheduling Python transfer jobs. For now, the one rule is to run your script with the environment's own python, not the bare system one.

First Contact: The Host Key Decision

Before any authentication happens, SSH performs a check in the other direction: the server proves its identity to you. Every SSH server has a host key — a cryptographic keypair whose public half identifies that machine. Your client compares the key the server presents against a local record of known servers, the known_hosts file. A match means you are talking to the same machine as last time. A mismatch — or an absent entry — means you might be talking to an impostor performing a man-in-the-middle attack. In that case, everything you send, credentials included, would flow through the attacker. The full story is in host keys and known_hosts.

Paramiko's SSHClient makes this decision explicit through a missing host key policy — what to do when a server is not in the loaded known_hosts data:

  • RejectPolicy (the default): refuse to connect to unknown servers, raising an exception. This is the correct behavior for unattended jobs.
  • AutoAddPolicy: accept and record whatever key the server presents, first time, no questions. Understand what that means: on first contact, your script trusts whoever answers the network — which is precisely the moment an attacker would answer. Nearly every SFTP tutorial on the internet pastes this line in. Use it, at most, against a throwaway lab box; never in a production or scheduled job.
  • WarningPolicy: log a warning and connect anyway — the same risk as AutoAddPolicy with better-documented regret.

The unattended-job pattern that follows from this: keep a small, job-owned known_hosts file containing exactly the servers this job talks to, load it, and leave the default reject behavior in place.

import paramiko

client = paramiko.SSHClient()
client.load_host_keys("/opt/jobs/reports/known_hosts")
# no set_missing_host_key_policy() call: the default rejects unknown servers

Two loading methods exist and differ usefully. load_system_host_keys() reads the calling user's ~/.ssh/known_hosts as a read-only trust source. load_host_keys(path) reads a file the client treats as its own — a better fit for jobs. The file lives with the job, travels with it between machines, and does not depend on which account happens to run it.

How does the server's key get into that file honestly? Not by connecting and hoping. One option is to copy the entry from a machine that already talks to the server. Another is to obtain the key's fingerprint from the server's administrator and verify it during one interactive ssh or sftp session. Or — if you run the server yourself — read it straight from the server's configuration. A fingerprint is a short hash of the key — a compact string two people can compare over the phone or a ticket. So "verify" genuinely means checking that the fingerprint your client displays matches the one the server's owner published. One deliberate verification, then every future run checks automatically. And if a run ever fails with a host key mismatch, stop the job and investigate. The mismatch means either a server rebuild nobody announced or the attack the check exists to catch.

Remember: AutoAddPolicy does not "fix" host key errors — it turns off the only check that protects your credentials from an impostor server. Unattended jobs verify a known_hosts file and reject everything else.

Authenticating with a Key, Not a Password

With the server's identity settled, it is your script's turn to prove itself. For unattended jobs, that should mean an SSH key, not a password. That means no secret typed into a scheduler field, no password expiry breaking the job at quarter-end. It also means easy revocation of exactly one credential when the job is retired. Generate a dedicated keypair for the job — a modern type such as Ed25519. Have the server's administrator authorize its public half for a dedicated service account that owns this flow and nothing else. The reasoning and the hygiene rules live in generating and storing SSH keys and service account hygiene.

The connect call, written for determinism:

client.connect(
    "sftp.partner.example.com",
    port=22,                       # SFTP rides the normal SSH port
    username="reports-job",        # the dedicated service account
    key_filename="/opt/jobs/reports/key_ed25519",
    allow_agent=False,             # ignore any SSH agent
    look_for_keys=False,           # do not hunt ~/.ssh for other keys
    timeout=15,                    # seconds to establish the TCP connection
    banner_timeout=15,             # seconds to receive the SSH banner
    auth_timeout=15,               # seconds for authentication to finish
)

Each parameter earns its place. key_filename points at the private key file — readable only by the account that runs the job. The two False flags matter more than they look. By default the library will also try any running SSH agent and any default keys in the user's ~/.ssh. That means a job can accidentally work at your desk using your key and then fail — or worse, keep working as you — in production. Turning both off makes the job authenticate the same way everywhere. The three timeouts cover the three phases of getting connected. Without them, a silent firewall drop leaves the script hanging until the scheduler's patience, not yours, runs out. If the key is protected by a passphrase, add passphrase= — sourced from somewhere safer than the script text, a topic the robustness article takes further.

Uploading, Downloading, and Looking Around

An authenticated client opens an SFTP session with client.open_sftp(), which returns the object that does the actual file work. The essential operations:

sftp = client.open_sftp()

sftp.put("/data/outbox/report.csv", "/inbox/report.csv.part", confirm=True)
sftp.rename("/inbox/report.csv.part", "/inbox/report.csv")

sftp.get("/outbox/batch.csv", "/data/incoming/batch.csv.part")

for entry in sftp.listdir_attr("/outbox"):
    print(entry.filename, entry.st_size, entry.st_mtime)

sftp.mkdir("/inbox/archive")       # raises OSError if it already exists
sftp.remove("/outbox/batch.csv")   # delete a remote file
sftp.close()

Three details in that block repay attention. First, confirm=True makes put() stat the uploaded file afterward and raise an error if the size does not match what was sent. That is a cheap completeness check. (Sizes are not proof of intact content; when that matters, compare checksums as described in hashing explained.)

Second, the upload goes to a temporary name and is then renamed. This is the temp-name-and-rename pattern. The receiving side sees report.csv appear only as a complete file, never as a half-written one being consumed by an eager watcher. The protocol's rename typically refuses to overwrite an existing target on most servers, which is harmless here because the final name is new. That is exactly the behavior you want if a duplicate would otherwise be clobbered silently.

Third, listdir_attr() beats listdir() for automation: it returns attribute records — name, size, modification time — in one round trip. That lets the script filter by pattern and skip files that are still growing without a second call per file.

Two practical notes on the transfer calls themselves. They stream: put() and get() copy in chunks, so a file larger than memory is no problem and needs no special handling. And both accept a callback argument — a function called with bytes-transferred-so-far and total bytes. That is how you add progress reporting for very large files. It is also how you add a log line every few hundred megabytes so a long transfer is visibly alive rather than silently hung. For routine job files, skip it; the completion log line is enough.

A Complete Small Script

Here is everything so far assembled into a real job. It fetches every CSV file from a partner's outbox, lands each one atomically, logs what happened, and exits with a code the scheduler can read.

import logging
import os
import sys
from pathlib import Path

import paramiko

HOST = "sftp.partner.example.com"
USER = "reports-job"                          # dedicated service account
KEY_FILE = "/opt/jobs/reports/key_ed25519"
KNOWN_HOSTS = "/opt/jobs/reports/known_hosts"
REMOTE_DIR = "/outbox"
LOCAL_DIR = Path("/data/incoming")

log = logging.getLogger("reports-pull")

def main():
    logging.basicConfig(level=logging.INFO,
                        format="%(asctime)s %(levelname)s %(message)s",
                        datefmt="%b %d %H:%M:%S")
    client = paramiko.SSHClient()
    client.load_host_keys(KNOWN_HOSTS)        # default policy rejects unknowns
    try:
        client.connect(HOST, port=22, username=USER, key_filename=KEY_FILE,
                       allow_agent=False, look_for_keys=False,
                       timeout=15, banner_timeout=15, auth_timeout=15)
    except paramiko.AuthenticationException:
        log.error("authentication rejected for %s@%s", USER, HOST)
        return 3
    except (paramiko.SSHException, OSError) as exc:
        log.error("could not connect to %s: %s", HOST, exc)
        return 4

    failures = 0
    try:
        sftp = client.open_sftp()
        for entry in sftp.listdir_attr(REMOTE_DIR):
            if not entry.filename.endswith(".csv"):
                continue
            final = LOCAL_DIR / entry.filename
            part = final.with_name(final.name + ".part")
            try:
                sftp.get(f"{REMOTE_DIR}/{entry.filename}", str(part))
                if part.stat().st_size != entry.st_size:
                    raise OSError("size mismatch after download")
                os.replace(part, final)       # atomic move to the final name
                log.info("fetched %s (%d bytes)", entry.filename, entry.st_size)
            except OSError as exc:
                failures += 1
                log.error("failed %s: %s", entry.filename, exc)
    finally:
        client.close()                        # always runs, even on a crash

    return 1 if failures else 0

if __name__ == "__main__":
    sys.exit(main())

The shape is as important as the calls. Connection problems and per-file problems are handled separately, because they mean different things. A refused connection fails the whole run with a distinct exit code. One bad file is counted and the loop continues to the next. The finally block guarantees the SSH connection closes no matter which path the run took. Leaked connections are how a flaky job slowly exhausts a server's session limit. (The client also works as a context manager — with paramiko.SSHClient() as client: — which closes on exit the same way. The explicit finally is used here because it keeps the cleanup visible while you are learning what needs cleaning.) And the exit codes — 0 clean, 1 some files failed, 3 auth, 4 unreachable — give the scheduler and your monitoring something honest to react to.

Run it with the environment's interpreter, and the log tells the story of each run in a form worth keeping:

$ /opt/jobs/reports/.venv/bin/python reports_pull.py
Mar 14 02:10:02 INFO fetched invoices.csv (184320 bytes)
Mar 14 02:10:05 INFO fetched orders.csv (92160 bytes)
Mar 14 02:10:11 ERROR failed legacy.csv: size mismatch after download
$ echo $?
1

That final echo $? prints the exit code the script returned — the same value cron or Task Scheduler will see. Checking it by hand once, before the job ever runs unattended, confirms the contract between your script and its scheduler actually holds.

The Errors You Will Meet

When the script misbehaves, the exception class is your first diagnostic. The common ones:

Exception Usual meaning Retry it?
AuthenticationException Wrong account, key not authorized, or wrong key file No — fix the credential; retries can trigger lockouts
BadHostKeyException Server's key differs from known_hosts No — investigate; rebuild or man-in-the-middle
SSHException Connection or protocol trouble mid-session Often — usually transient
Timeout errors Dead host, firewall drop, congested network Yes — with backoff and an attempt cap
OSError from file operations Missing remote path, permissions, full disk Rarely — the cause persists until fixed

The pattern behind the table: distinguish failures that might pass (network weather) from failures that will not (wrong credentials, missing directories). Retry only the first kind, and make the second kind loud. That discipline — with backoff, jitter, and attempt budgets — is the subject of our retry and error handling series. Its Python implementation is worked through in writing robust Python transfer scripts.

Test Against a Server You Control

Do not point a half-written script at a partner's production server; their intrusion logs — and their patience — deserve better. Stand up an SFTP server you control, aim the script at it, and iterate freely. On Windows, Sysax Multi Server serves SFTP alongside FTPS, FTP, and HTTPS. Its activity logging is the underrated half of the exercise. The server-side log shows every session, login, and file operation your script performed, which settles "did it actually upload?" questions with evidence instead of guesswork. Watching your own script from the server's side once will teach you more about SFTP than a week of client-side debugging.

Testing is also the honest moment to re-ask a question from the start of this series: does this flow need custom code at all? Suppose the job has revealed itself to be a standard scheduled upload or download with no special logic. A configurable tool such as Sysax FTP Automation already speaks SFTP with scheduling, folder monitoring, and email notifications built in. In that case, nobody has to maintain a script. Write Python for the flows with logic worth owning.

Where to Go Next

You can now connect with verified host keys, authenticate with a dedicated key, move files atomically, and exit with codes a scheduler understands. That is the complete skeleton of a trustworthy SFTP job. From here, you have three directions. Harden the skeleton with retries, timeouts, and real logging in writing robust Python transfer scripts. Or grow it into a configurable tool a coworker can run in building a small transfer utility. Or cover the legacy endpoints with FTP and FTPS in Python.

Frequently Asked Questions

Why doesn't Python include SFTP in the standard library?
SFTP runs inside SSH, and a full SSH implementation is a large, security-critical project with its own release rhythm. Keeping it outside the standard library lets it ship fixes on its own schedule. The practical consequence is simply one dependency to install and pin per job.
Is AutoAddPolicy ever acceptable?
Only against a disposable test box where an impostor would cost you nothing. It trusts whatever key the network presents on first contact, which is exactly the opening a man-in-the-middle needs. Production and scheduled jobs should load a known_hosts file and keep the default reject behavior.
How do I get a server's host key into known_hosts safely?
Verify it once, out of band. Compare the fingerprint with the server's administrator or copy the entry from a machine that already trusts the server. Or read it from the server's own configuration if you run it. After that single verification, every future run checks the key automatically.
Can the job just use my personal SSH key?
It will work, and it is still the wrong move. The job breaks when you leave or rotate your key, and every transfer is logged as you. Give each job a dedicated service account and its own key, so access can be granted, audited, and revoked independently.
Does SFTP need extra firewall ports the way FTP does?
No. SFTP runs entirely inside one SSH connection on a single port — normally port 22 — with no separate data connections to negotiate. That is a large part of why it became the default protocol for automated transfers through firewalls.
Can put() handle files bigger than the machine's memory?
Yes. Both put() and get() stream the file in chunks rather than loading it whole, so file size is limited by disk and patience, not RAM. For very large files, pass a callback to log progress so a long transfer is distinguishable from a hung one.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.