Home › Topics › Large Files › Resumability

Designing Large Transfers to Resume

Most transfer protocols can resume an interrupted file. Most large transfers still restart from zero when they fail. The gap between those two facts is design: resume is a capability the protocol offers. Whether an unattended job actually uses it depends on a dozen small decisions about names, state, verification, and retries that nobody made deliberately. The job uploads to a temporary name, the retry looks for the final name, finds nothing, and starts again. The source file was regenerated between attempts, the resume appends new bytes onto old ones, and the result is corrupt. Each of these is a design flaw, not a network problem.

This article is about making those decisions on purpose. It does not explain how resume works inside each protocol — restart markers in FTP, offsets in SFTP, byte ranges in HTTP. Our resume and checkpoint restart series does that. Here the subject is the transfer around the protocol. We cover how temporary names interact with resume, how to detect a changed source, and what a job needs to remember between attempts. We explain how to make a re-run safe no matter what state the last run left behind and how to verify a resumed file. We also cover the retry loop that does all of this without a person. It is part of our Large File Strategies series.

Four Things a Resumable Transfer Needs

Resume means continuing an interrupted transfer from the byte where it stopped instead of from the beginning. For that to happen in an unattended job, four things have to be true at once:

  1. The protocol supports it. FTP, FTPS, and SFTP can resume uploads and downloads. HTTP can resume downloads through byte ranges, and uploads only when the server offers a chunked or resumable upload scheme. rsync resumes with its partial-file options. A streaming pipeline — compress-and-pipe over SSH, for example — cannot resume at all, because there is no file on either side to continue.
  2. The client asks for it. Resume is never the default. Every client has a flag or setting, and a scripted transfer has to pass it explicitly on every run.
  3. The server permits it. Some servers disable resumed uploads, some disable them for particular accounts or folders, and some silently overwrite when asked to append. Test it with a deliberately interrupted small transfer before trusting it on a large one; the resume series has the checks for each server type.
  4. The design cooperates. Names, state, verification, and retries all have to be arranged so that the second attempt can find and continue what the first left behind. This is the part that is usually missing, and it is the rest of this article.

Temporary Names and Resume

The standard defense against a half-written file being picked up by whatever consumes it is a temporary name. Upload as export.bin.part, then rename to export.bin only when the transfer completes. That way, the final name never exists in an incomplete state. The full argument is in temp names and atomic renames, and every large transfer should use it. But temporary names and resume interact, and the interaction has to be designed.

The first rule: resume targets the temporary name. After an interruption, the partial file on the server is called export.bin.part. A retry that uploads to export.bin with resume enabled finds no such file, starts from byte zero, and — worse — leaves the old partial lying around forever. The retry must ask the server for the size of export.bin.part, continue that file, and rename only after the whole thing is confirmed. Some clients and servers generate their own temporary names by appending a random suffix or writing into a hidden folder. If such a tool drives the transfer, find out what it does on an interrupted run before relying on it. A temporary name the next run cannot predict is a temporary name that cannot be resumed.

The second rule: the rename is the last step, after verification. A rename immediately after the last byte is sent is one step too early. The transfer may have completed with the wrong length, and now the wrong file has the right name. Verify first, rename second. The rename itself is instant on the same volume and cannot be interrupted halfway, which is exactly why it is the right signal of completion.

The third rule concerns whatever is watching the receiving folder. A watcher that reacts to *.part files, or to any file that stops growing for thirty seconds, will pick up a paused partial during the retry delay. Configure the watcher to ignore the temporary pattern outright, so that only the rename triggers it. The arrival contract article covers the receiving side of that agreement.

The Modified-Source Problem

Here is the failure that turns resume from a convenience into a hazard. Attempt one sends 28 GB of a 40 GB export and dies. Overnight, the export job runs again and regenerates the file with today's data. Attempt two, the next morning, asks the server how much of export.bin.part it has — 28 GB — and appends bytes 28 GB onward from the new file. The result has the right length, passes a size check, and is 28 GB of yesterday followed by 12 GB of today. Nothing in the protocol can detect this; the bytes were delivered exactly as requested.

The defense is to record what the source looked like when the transfer started, and refuse to resume if it has changed. Three cheap facts identify a source: its size in bytes, its modification time, and a hash of its first few megabytes. The hash catches a regenerated file of identical size. Before resuming, compare all three with the live source. If any differ, the partial on the server is worthless: delete it and start from zero. That is the only correct outcome, and a job that does it automatically never ships a stitched-together file.

Remember: a resume is only valid if the source is byte-for-byte what it was when the partial was written. Size, modification time, and a head hash cost nothing to record and compare. A job that resumes without checking them will, one day, deliver a file that is half yesterday and half today. That file will pass every size check on the way.

The State File

The three source facts, the names in play, and where the transfer got to all need to live somewhere between attempts. A state file — a small text file beside the job — is the simplest place. It is the large-file cousin of the manifest used in chunked transfers, and it carries everything the next run needs to decide what to do:

job: export-push
source: /srv/export/export.bin
source_size: 40000000000
source_mtime: Mar 14 01:52
source_head_sha256: 5b7d0c…91ae   (first 64 MB)
remote_tmp: /inbound/export.bin.part
remote_final: /inbound/export.bin
started: Mar 14 02:10
attempts: 1
last_confirmed_bytes: 27917287424
status: interrupted

Write it when the transfer starts, update it after every attempt, and mark it done after the final verification and rename. If the job is invoked again and finds a state file with status: done for a source that has not changed, it has nothing to do. That single check is what makes the job safe to run twice, which is the property the next section is about.

Idempotent Re-Runs

An idempotent job is one you can run any number of times and get the same result as running it once. A second run after success does no harm. A run after a failure finishes the work rather than duplicating or corrupting it. The general idea is explained in idempotency in plain words. For a resumable large transfer it comes down to inspecting the situation at the start of every run and choosing one of five actions. The diagram below shows the decision.

Decision flow for an idempotent resumable transfer job. The job inspects the state file, the source, and the remote temporary file, then takes one of five actions: skip because already done, start fresh, resume from the remote size, verify and rename, or delete the remote partial and restart because the source changed or the remote is larger than the source.

The five branches in words. If the state file says done and the source is unchanged, exit — nothing to do. If there is no remote partial, start a fresh upload to the temporary name. If the remote partial is smaller than the source and the source is unchanged, resume from the remote size. If the remote partial is exactly the source size, skip the transfer and go straight to verification. A previous attempt may have finished sending and died before the rename. And if the source has changed, or the remote partial is somehow larger than the source, delete the remote partial and start fresh. A larger partial is a sure sign of an earlier stitched resume or a stale file. Every branch ends in the same place: verify, rename, record. Any failure updates the state file and returns to the top on the next run.

The Commands That Resume

The design above works with any client that can resume; three long-standing open-source tools show the shape. Each uploads to the temporary name, resumes if a partial exists, and renames only on success.

# curl: -C - asks the server for the partial's size and continues from there;
# the -Q command runs on the server only after a successful transfer
$ curl -sS -u alex: --key ~/.ssh/id_ed25519 -C - -T /srv/export/export.bin \
    -Q "-rename /inbound/export.bin.part /inbound/export.bin" \
    sftp://10.20.30.40/inbound/export.bin.part

# lftp: put -c continues an existing remote file; mv renames it afterward
$ lftp -e "put -c /srv/export/export.bin -o export.bin.part; mv export.bin.part export.bin; bye" \
    sftp://alex@10.20.30.40/inbound/

# rsync: --partial keeps the partial on interruption, --append-verify continues it
# next time and checks the whole file's checksum at the end
$ rsync --partial --append-verify --timeout=120 /srv/export/export.bin \
    alex@10.20.30.40:/srv/inbound/export.bin.part
$ ssh alex@10.20.30.40 mv /srv/inbound/export.bin.part /srv/inbound/export.bin

The rsync variant deserves a note. Plain --append trusts the existing remote bytes without checking them, which is exactly the modified-source hazard. --append-verify includes the existing bytes in a whole-file checksum after the transfer and re-sends the entire file from scratch if it does not match. This is a built-in version of the source check above, at the price of reading the whole file on both ends. Our rsync flags that matter article covers the surrounding options. Windows includes curl.exe, so the first form works unchanged from PowerShell. Scripted SFTP clients on Windows have their own resume switches, and PowerShell SFTP scripting shows where they live.

Verifying After a Resume

A resumed file deserves more suspicion than a file sent in one piece, because the join point is a place where things go wrong. A server may have reported a partial's size before its last buffer was flushed to disk. A resume may have started one byte off, or a client may have appended to the wrong file. Two checks catch all of these.

Size first. Compare the remote file's size with the source's. It is instant, and a mismatch means something is wrong before you spend minutes hashing. Then hash. Compute a hash of the whole remote file. Either ask the server, if it offers a hashing command, or read the file on the server side. Then compare the hash with the source's hash. A size match with a hash mismatch is the signature of a bad join or a changed source. The only fix is to delete the remote file and start over; do not try to repair it. Computing hashes on 40 GB without doubling the transfer time is its own subject, covered in verifying large transfers. The verification checks specific to resumed transfers, protocol by protocol, are in the resume series.

The Unattended Pattern

All of the above is useless unless it runs without a person. The pattern is a loop: attempt, and on failure record state, wait, and attempt again. Set a limit, so a permanent failure eventually raises an alarm instead of retrying forever. Because every attempt starts by inspecting the situation, the loop body is simply "run the job." The job decides whether that means start, resume, verify, or skip.

#!/bin/bash
# Retry a resumable upload up to six times with a growing delay.
# push-export.sh inspects the state file and does the right thing on each call.
for attempt in 1 2 3 4 5 6; do
    if /opt/jobs/push-export.sh; then
        exit 0
    fi
    echo "$(date '+%b %d %H:%M') export-push attempt $attempt failed, waiting" >&2
    sleep $(( 60 * attempt ))
done
echo "$(date '+%b %d %H:%M') export-push gave up after 6 attempts" >&2
exit 1

The delay grows with each attempt — one minute, then two, up to six. That way, a far end that is briefly down is not hammered, and a longer outage is given time to end. The reasoning behind the schedule and the limit is in retry strategies and backoff. What the loop produces in practice looks like this in the job's log:

Mar 14 02:10:03 export-push attempt=1 remote=0 source=40000000000 action=start
Mar 14 02:31:47 export-push attempt=1 error: connection reset by peer at 27917287424
Mar 14 02:32:47 export-push attempt=2 remote=27917287424 source=40000000000 action=resume
Mar 14 02:41:58 export-push attempt=2 sent=12082712576 size=ok sha256=ok renamed=export.bin
Mar 14 02:41:58 export-push status=done attempts=2

Read it against the running example of this series — a 40 GB file on a 200 megabit link, about 22 MB per second. The first attempt ran twenty-one minutes and died at 27.9 GB. One minute later the second attempt found the partial, confirmed the source had not changed, resumed, and sent the remaining 12 GB in nine minutes. Total time: thirty-two minutes, against thirty for an uninterrupted run — and against fifty-one for a restart from zero. Nobody was awake. That is the whole point.

Scheduled-transfer tools package this loop. Sysax FTP Automation, for example, has built-in retry and error handling for its scheduled jobs. So the attempt-wait-attempt cycle and the failure alert are configuration rather than script. The resume-or-restart decision remains a matter of how the transfer is designed, exactly as above.

Chunks as Resume Points

One more design option combines the two large-file tactics. If the file is split into chunks, each chunk boundary is a natural resume point that needs no protocol support at all. In that design, the job records which chunks the receiver has confirmed, and a retry simply sends the ones that are missing. This is the right approach when the far end cannot resume. It has a side benefit — the per-chunk hashes in the manifest give you verification for free at every boundary. Where the far end can resume, a chunked transfer whose pieces also resume is about as robust as a network transfer gets. In that case, an interruption costs a few seconds within one chunk, and a corrupt piece is identified by name.

The Version to Keep in Your Head

Resume is a protocol feature; a transfer that actually resumes unattended is a design. Upload to a temporary name and make the retry target that same name. Record the source's size, modification time, and head hash, and refuse to resume onto anything that has changed. Keep a state file so each run knows what the last one did. Start every run by inspecting the situation and choosing skip, start, resume, verify, or delete-and-restart. That makes the job safe to run any number of times. Verify size and hash before the rename, never after. Wrap it all in a retry loop with a growing delay and a limit.

Continue with verifying large transfers without doubling the time for the hashing side. Continue with choosing protocols and settings for multi-gigabyte moves if you are still deciding which protocol will carry the file. For a wider view of recovery design — what to do when the failure is not the network but the far end, the credentials, or the file itself — see designing for recovery.

Frequently Asked Questions

My client supports resume, so why does my scheduled job always restart from zero?
Usually because the retry does not target the file the first attempt left behind. If the job uploads to a temporary name and the retry looks for the final name, or a tool generates a different temporary name each run, there is nothing to resume. Make the temporary name fixed and predictable, and point the retry at it.
Is it dangerous to resume a file that was regenerated between attempts?
Yes. The resume appends new bytes onto old ones and produces a file with the right length and wrong contents, which passes a size check. Record the source's size, modification time, and a hash of its first few megabytes when the transfer starts. Start over if any of them change.
Should the file be renamed to its final name as soon as the last byte is sent?
No — verify first. Compare size, then hash, and rename only when both match. A rename before verification means a wrong file can carry the right name, and anything watching the folder will consume it.
What is the difference between rsync --append and --append-verify?
--append continues from the remote file's length and trusts the existing bytes without checking them. --append-verify does the same but then verifies a checksum of the whole file, including the old bytes. It re-sends the file from scratch if the checksum does not match. Use --append-verify for anything that matters.
How many times should an unattended job retry before giving up?
Enough to outlast a typical short outage, with a growing delay between attempts, but few enough that a permanent problem raises an alert within your deadline. Five or six attempts over twenty to thirty minutes is a common starting point; adjust it to the job's cutoff time.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.