Exactly-Once Thinking for File Pipelines
At some point, someone with sign-off authority will ask you a version of this question. "Can you guarantee that every file is processed exactly once — no losses, no duplicates?" It sounds like a reasonable requirement, the kind you should be able to promise with enough care and budget. And the honest answer is strange enough that it is worth learning to give well. No one can guarantee exactly-once delivery over a real network. And yes, you can still guarantee the result they actually want.
That answer comes from distributed-systems theory, where "exactly-once" is a famous impossibility with a famous workaround. The theory travels surprisingly well. A nightly SFTP batch job and a global message queue turn out to have the same anatomy and the same escape hatch. This article — part of our duplicate detection and idempotency series — translates the idea into file-pipeline terms. It explains why the naive guarantee is unachievable, and the two-mechanism combination that achieves the same outcome. It covers where each mechanism physically lives in a file flow. It also covers the long list of things you get to stop worrying about once the pieces are in place.
Three Possible Promises, and the One You Can't Make
Start with vocabulary. Suppose a sender moves a file to a receiver over a network that can fail at any moment. There are exactly three delivery promises the sender's logic can aim for:
- At-most-once: send each file one time, never retry. No duplicates, ever — but any failure, including a lost confirmation, means the file may simply not arrive, and nobody re-sends it.
- At-least-once: send, and keep re-sending until a confirmation arrives. Nothing is ever lost — but a lost confirmation triggers a re-send of a file that already made it, so duplicates are possible by design.
- Exactly-once: every file arrives once and only once. The one everybody wants.
Here is why the third promise is not available at the delivery level. Picture the final moments of an upload: the last byte arrives, the server stores the file, and the server sends back its success confirmation. Now suppose that confirmation is lost — a dropped connection in the closing handshake, a timeout, a crashed session. The sender is left with silence, and silence is ambiguous. It looks identical whether the upload failed or the upload succeeded and only the receipt died. The sender must now act on incomplete information. If it declines to retry, it risks losing a file (that is at-most-once). If it retries, it risks duplicating one (at-least-once). What it cannot do is know. No cleverness in the retry logic can manufacture the missing knowledge, because any further confirmation can be lost the same way. We walked through this exact scenario, timeline and all, in where duplicate files actually come from. Here it graduates from anecdote to law. Over an unreliable channel, a sender can never be certain the other side received something. So "delivered exactly once" is not a promise anyone can keep at the wire level.
If the abstraction feels slippery, a postal analogy pins it down. You mail a contract and ask for a signed return receipt. Weeks pass; no receipt. Did the contract get lost, or did the receipt get lost on its way back? From your mailbox, the two situations are indistinguishable. Mailing a letter that asks "did you get my letter?" just creates another letter whose fate you cannot know. At some point you either give up (and maybe the contract is lost) or send a second copy (and maybe the recipient now has two). Every transfer protocol, however sophisticated, is having this same correspondence at machine speed.
This has a practical corollary for anyone reading vendor datasheets: when a product advertises exactly-once anything, look for the fine print. What is really on offer — in every serious system — is the two-part construction this article is about. That is not a criticism; it is the correct engineering. It just means the guarantee lives partly in your half of the system, whether or not you bought the product.
Pick Your Failure: A Missing File or an Extra Copy
Since perfect delivery is off the table, the design question becomes: which imperfect promise do you build on? Compare the two honest options by their failure modes:
| Promise | Failure mode | How you find out | How you fix it |
|---|---|---|---|
| At-most-once | A file silently never arrives | You don't, until a reconciliation or an angry call | Manual hunt: which file, from when, is it still available? |
| At-least-once | A file arrives twice | The duplicate is sitting right there, detectable by name and hash | Automatically: detect and skip the extra copy |
Framed this way, the choice stops feeling like a dilemma. A missing file is an invisible failure — nothing exists to alert on, and discovery usually happens downstream, late, and embarrassingly. An extra copy is a visible failure. The evidence is a file you can inspect, compare, and discard, and a machine can do all three. Reliability engineering has a strong preference for failures that announce themselves. So the professional consensus, for file pipelines as much as message systems, is: build on at-least-once, then neutralize the duplicates. Retry until confirmed — the discipline our retry and error handling series covers in depth. Treat the resulting extra copies as a handled, expected event rather than an incident.
The Escape Hatch: Move the Guarantee Up a Level
Now the central idea, the one worth writing on the whiteboard. Exactly-once delivery is impossible — but delivery was never the point. Nobody's finance team cares how many times the bytes of settle_YYYYMMDD.csv crossed the network. They care that its rows landed in the settlement table once. The requirement is about effect, not transit. Effects happen on your side of the wall, where you have something the network can never give the sender. You have complete knowledge of what has already been processed.
So you split the guarantee into two mechanisms, each placed where it can actually be enforced:
- At-least-once delivery on the transfer leg: retries, resends, generous timeouts — make sure every file gets there, accepting that "there" may receive some files more than once.
- Idempotent processing goes on the receiving leg: a duplicate gate in front of every irreversible step. That gate is the processed-files ledger with its atomic claim, from detecting duplicates. So no matter how many copies arrive, each file's content takes effect exactly once.
The combination is called effectively exactly-once, and the qualifier "effectively" is not weasel-wording — it is precision. Transfers may repeat; outcomes do not. The diagram shows the two mechanisms in position, and where duplicates are allowed to exist versus where they die.
Notice the asymmetry of effort. The delivery leg gets to be simple and aggressive — retry freely, resend on doubt. That is because nothing on that leg is irreversible. All the care concentrates at one narrow gate on the processing leg. There, the ledger's atomic claim decides, once per file identity, whether the effect happens. One careful gate is much easier to build and audit than a whole pipeline of half-careful steps. If the gate's mechanics are not fresh in your mind, the foundation article on idempotency in plain words is the place to start.
Where Each Mechanism Lives in a Real Flow
Mapping the theory onto an actual nightly job, leg by leg:
The delivery leg — the sender's half
- Retries with backoff on every transfer, until a confirmation or a human-alerting give-up. In practice this is scheduler configuration. A tool like Sysax FTP Automation runs scheduled transfer tasks with retry and error handling built in. It emails you when a transfer keeps failing. That is the at-least-once half of the construction, configured rather than coded.
- Stable file naming, so that when a retry does re-deliver, the copy collides recognizably with the original instead of arriving as a stranger.
- Post-upload verification belongs where the stakes justify it. After uploading, list the remote file and compare sizes, or fetch a hash if the server side can produce one. Those are the patterns in verifying transfers end to end. Verification shrinks the uncertainty window and catches truncated uploads at the cheapest moment.
The processing leg — the receiver's half
- Completeness first: never process a file still being written. Atomic renames, settle checks, and the rest of the arrival discipline live in our partial-file safety series. They are the precondition for everything below.
- The ledger gate immediately before the irreversible step: hash the completed file, attempt the atomic claim, process on success, skip-and-log on duplicate.
- Guarded side effects: notifications and downstream triggers keyed to the file's identity, so a reprocessed or duplicate file cannot re-fire them.
Split this way, each half is testable on its own. You can prove the delivery leg delivers (kill connections mid-transfer and watch retries recover). You can separately prove the gate holds (feed the same file in twice and watch the second copy get skipped). When both tests pass, the composed guarantee follows. No step in the middle needs to be perfect, because the design never asked it to be.
When the two halves belong to different companies
In partner flows, the legs often straddle an organizational boundary. The partner owns the sender and its retries; you own the intake and the gate. Two ground rules keep that split healthy. First, the receiver always owns deduplication. You cannot audit the partner's retry logic, and they cannot see your processing history, so the gate can only live on your side. Never accept "we promise we only send once" as a control. Second, say out loud, ideally in the onboarding notes, that resends are welcome. A partner who knows duplicates are handled will resend promptly when something looks wrong. They will not sit on a suspected failure for a day while they investigate. The gate does not just protect your data; it buys the whole flow permission to recover fast.
Receipts Shrink the Problem; They Cannot Close It
A natural objection: "What about acknowledgments? If the receiver confirms each file, doesn't that give us exactly-once?" Receipts are genuinely valuable. These include protocol-level confirmations, verified uploads, and application-level receipts like the signed delivery notifications used in B2B flows (see MDN proof of delivery). All narrow the window in which a sender is uncertain. Narrower uncertainty means fewer unnecessary resends, which means fewer duplicates to skip. Well-run pipelines use them.
But look at where the receipt travels: across the same unreliable network, in the other direction. A receipt can be lost exactly the way the original confirmation was lost. Then the sender is back in the ambiguous silence, and the safe response is still to resend. Receipts reduce the frequency of duplicates; only the idempotent gate removes their harm. That ordering deserves a rule of thumb:
Remember: receipts and verification are optimizations; the ledger gate is correctness. Build the gate first — it makes every other failure recoverable. Add receipts second — they make recoveries rarer. A pipeline with receipts but no gate still double-processes; a pipeline with the gate but no receipts merely retries more than it needs to.
What You Can Stop Worrying About
The quiet payoff of exactly-once thinking is subtraction. Once the two mechanisms are in place, several chronic anxieties simply leave the room:
- Stop tuning retries timidly. Teams without a duplicate gate keep retry counts low out of fear. With the gate, retries are free of side effects, so you can set them by delivery need rather than by dread. That is the transient failure math in retry and error handling.
- Stop adjudicating ambiguous runs by hand. "The job says failed but the file looks like it arrived — should I rerun?" With the gate: yes. Always yes. The rerun is safe by construction, and the 2 a.m. judgment call disappears.
- Stop chasing the perfect network. Flaky partner links and overloaded VPNs stop being correctness problems and become mere performance problems.
- Stop treating duplicate arrivals as incidents. A skipped duplicate is the system working. Log it, count it, and review the count weekly. A sudden spike usually points at a new delivery path or a retry storm. The taxonomy in where duplicates come from will help you name it.
- Redirect the worry that remains. The failure mode still standing is absence — the file that never arrived because a sender gave up permanently. That is what freshness monitoring is for: alert when an expected file has not appeared by its deadline. That discipline is covered in our transfer job monitoring series. Exactly-once thinking moves your alerting budget from "too many copies" (machine-handled) to "zero copies" (worth waking someone).
One File's Eventful Journey
To make the construction concrete, trace a single file through a bad night. The client uploads settle_YYYYMMDD.csv; the upload completes but the confirmation is lost; the client retries; the second copy arrives and is skipped. Here is how that night should read in the logs — client first, then the receiving pipeline:
# sender's job log Mar 14 02:10:31 UPLOAD settle_YYYYMMDD.csv start (attempt 1) Mar 14 02:12:05 ERROR timeout waiting for server reply — marking attempt failed Mar 14 02:13:05 UPLOAD settle_YYYYMMDD.csv start (attempt 2, retry 1 of 3) Mar 14 02:14:42 OK transfer confirmed (attempt 2) # receiving pipeline log Mar 14 02:12:01 ARRIVED settle_YYYYMMDD.csv size=48211 sha256=9f86d081... Mar 14 02:12:02 CLAIM ok — processing; run_0042 Mar 14 02:12:19 LOADED 1240 rows; ledger result=loaded Mar 14 02:14:42 ARRIVED settle_YYYYMMDD.csv size=48211 sha256=9f86d081... Mar 14 02:14:42 SKIP duplicate of run_0042 — moved to duplicates/
Read the two halves against each other and the whole article is in miniature. The sender's log tells an honest story of failure and recovery; it believes attempt 1 failed. The receiver's log knows better: attempt 1 delivered the file, which was claimed and loaded. Attempt 2's copy hit the gate and bounced. Note that the receiving server's own transfer log independently corroborates the arrivals. A server like Sysax Multi Server records every transfer to its log file and database. So both uploads appear there with timestamps and source addresses. The question "did the retry really deliver a second copy?" has a queryable answer. Three records, one consistent story, zero double-loaded rows: effectively exactly-once, working as designed.
The Version to Tell a Colleague
Exactly-once delivery is impossible to promise over a real network. A sender can never distinguish "it failed" from "it worked and the confirmation was lost." So it must either under-deliver or over-deliver, and over-delivering is the safe choice. The professional construction is two mechanisms: at-least-once delivery (retry until confirmed) plus an idempotent processing gate (an atomic ledger claim before anything irreversible). Together they give effectively exactly-once outcomes: files may travel more than once, but each takes effect once. Receipts and verification make duplicates rarer; the gate makes them harmless; monitoring watches for the one failure left, the file that never came.
For the gate's construction details, read detecting duplicates: names, sizes, hashes, and ledgers. To watch a real pipeline acquire the whole construction after an incident, finish the series with the worked example. It includes the ledger, hash checks, and deliberate rerun mode.
Frequently Asked Questions
Why is exactly-once delivery impossible? It sounds solvable.
Is "effectively exactly-once" just marketing language?
Should I choose at-most-once anywhere?
If I verify every upload with a hash, do I still need the ledger?
What should alert a human in an exactly-once pipeline?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
