The Partial-File Safety Checklist for Any Flow
The rest of this series taught the mechanics one at a time. It explained why half-written files get read and how renames and markers stop it at the source. It explained how settle checks defend a consumer left on its own, and what an in-flight transfer failure leaves behind. This closing article is the working document. It is the piece you open when someone says "is the invoicing flow safe?" and you need a real answer by the end of the day.
It compresses the series into three tools. One is a control matrix that shows every defense, who owns it, and where it fails. Another is a copyable audit checklist. The third is a worked audit of one realistic, risky flow taken from exposed to safe, with the retrofits ranked by effort. This article is the capstone of our Partial-File Safety series. It is designed to be used with a specific flow in mind — pick one before you read on.
Three Seats at the Table
Every file handoff has the same three seats, and partial-file safety is decided by what each seat does:
- The producer — whoever writes or delivers the file — can make incompleteness invisible. Write under a temp name and rename atomically on completion, or announce completion with a marker or control file. Producer controls are the strongest, because they convert "probably done" into "declared done."
- The consumer — whoever picks the file up — can refuse to be rushed. Match final names only, hold every arrival behind a settle check, and verify against expected sizes or hashes. Process idempotently so the rare miss is recoverable. Consumer controls are the ones you can always deploy, because they need nobody's permission.
- Both sides together own the contract and the evidence. That includes the written agreement on names, ordering, and cleanup, and the logging that records what actually happened. That also includes the sweeps that catch debris — including the abandoned partials left by transfers that died midway.
A flow is in good shape when at least one real control sits in each seat. A flow is in trouble when everything depends on a single seat — usually the consumer — or, worse, on timing luck.
Why frame this as an audit at all? Because almost nobody gets to design these flows fresh. File handoffs accrete. A script written for one purpose gains a watcher, the watcher gains a second consumer, and the partner changes upload tools. Five years later the person responsible for the flow is not the person who built any part of it. Auditing means recovering the facts instead of trusting the folklore that came with the handover. Find what actually writes, what actually triggers, and what actually happens when a transfer dies.
The Control Matrix
Here is the whole series in one table. "Effort" assumes you are retrofitting an existing flow, not building fresh.
| Control | Seat | Protects against | Effort | Where it fails |
|---|---|---|---|---|
| Temp name + atomic rename | Producer | Consumers ever seeing an in-progress file | Low | Temp on a different filesystem — the rename silently becomes a copy |
| Upload to temp remote name, rename on server | Producer | Live and dead uploads sitting under final names | Low | Upload accounts without rename permission |
| Marker file written after the data | Producer | Early triggers; incomplete multi-file sets | Low | Ordering violations: parallel or name-sorted uploads sending the marker early |
| Control file with size, count, or hash | Producer | Truncation the sender's own tooling never noticed | Medium | Formats nobody agreed on; control data the consumer never checks |
| Final-name pattern filtering | Consumer | Processing temp names and foreign debris | Trivial | Catch-all patterns that match every file in the folder |
| Settle check (size stability + first-seen age) | Consumer | Early reads when the producer offers no signal | Low-Medium | Producer pauses longer than the window; abandoned partials look settled |
| Verify before processing (size/count/hash) | Consumer | Truncated, corrupted, and hybrid files of every origin | Medium | Needs sender-side evidence to compare against |
| Claim-by-move, then idempotent processing | Consumer | Double pickups; unrecoverable damage from the rare miss | Medium-High | Legacy imports that cannot safely re-run take real redesign |
| Sweeps for stale partials and orphan markers | Both | Old debris being processed later; stuck flows nobody notices | Low | Silent deletion that destroys the evidence an incident needs |
| Session logging, retained and consulted | Both | Blind diagnosis; truncations nobody can prove | Low | Logs that exist but are never read catch nothing |
| Written arrival contract with the sender | Both | Drift, mismatched assumptions, undocumented conventions | Low | Contracts nobody re-verifies after a tooling change |
Remember: no single row of that table is sufficient on its own. The resilient shape is one control from each seat: a producer signal, a consumer gate with verification, and shared operational hygiene. Two layers catch what the third misses.
The Audit Checklist
The matrix answers "what exists?"; the checklist answers "what does this flow have?" Print it, or paste it into the flow's runbook, one copy per flow:
PARTIAL-FILE SAFETY AUDIT -- one sheet per flow Flow: ____________ Producer: ____________ Consumer: ____________ PRODUCER SIDE [ ] Files appear under final names only when complete (rename, or server temp-name behavior) [ ] Temp/staging location is on the same filesystem as the destination [ ] Multi-file sets are gated by a marker or manifest, written last [ ] Markers are sent as a separate step, only after data transfer confirms success [ ] Critical deliveries carry control data: expected size, count, or hash CONSUMER SIDE [ ] Trigger pattern matches final names only; temp names are excluded [ ] A marker gate or settle check stands between "noticed" and "processed" [ ] The settle window exceeds the producer's longest observed pause [ ] File age is measured from first-seen, not from preserved timestamps [ ] Deliveries are verified against expected size/count/hash before processing [ ] Files are claimed (moved) before processing, and reprocessing is safe BOTH SIDES / OPERATIONS [ ] The killed-transfer test has been run against this flow, results documented [ ] The slow-writer test has been run against this consumer, results documented [ ] Stale partials and orphan markers are swept, quarantined, and alerted on [ ] Transfer sessions are logged, retained, and actually consulted on failure [ ] The arrival contract is written down and re-checked after tooling changes
Scoring is deliberately blunt. Any unchecked box in the producer section means the consumer section must carry the flow. Unchecked boxes in both sections means the flow works on luck, and the ticket history usually proves it.
Running the Audit on a Real Flow
Facts first, opinions second. An audit is four short investigations:
- Establish who writes and how. What tool produces or uploads the file? Does it rename, or write straight to the final name? Does anything document it — or does someone have to test? When the producer is a partner, this question goes in an email; the answer goes in the contract.
- Establish what triggers the consumer. Find the actual pattern and the actual gate. "The watcher picks it up" is not an answer; "the monitor matches
*.csvand launches the import with no settle step" is. The workflow anatomy to compare against is in our watch folders and event-driven transfers series. - Run the two tests. For the slow-writer test, write a file into the watched folder over a full minute and see whether the consumer pounces. For the killed-transfer test, kill an upload midway and see what remains, under what name, and what the log says. Both are ten-minute experiments, and they convert every "should be fine" into a fact.
- Read the evidence trail. Does the server log sessions, and are the logs retained and reachable? A server such as Sysax Multi Server records every upload session to file and database. If your estate has that evidence, the audit checks whether anyone consults it. If it has nothing, that absence is itself a finding.
One scope note before you start: audit the whole journey, not just the first hop. A flow whose intake is immaculate can still forward files onward — to an archive, a second system, another partner. Every onward hop is a fresh producer-consumer handoff with its own race. The checklist applies at each hop where a process picks up files that another process delivers. Most real flows turn out to contain two or three such seams. The incident usually lives in the one nobody thought of as a handoff.
Worked Example: The Nightly Partner Batch, Made Safe
A composite of flows we have all inherited. A partner uploads three CSV extracts every night over SFTP into a drop folder on your server. A folder monitor watches the drop and launches the import as files arrive. The import loads straight into a reporting database. Month-end runs are big and slow; the flow has a history of "corrupt file" tickets that resolve themselves by morning.
What the audit found
- Uploads land directly under final names — the server writes them in place as bytes arrive, and dead uploads stay there. (Killed-transfer test: a truncated file with a plausible name survived indefinitely.)
- The monitor matched every new file and fired per file — the import for file one started while files two and three were still uploading. (Slow-writer test: the consumer pounced eleven seconds into a sixty-second write.)
- No settle check, no verification, no control data — the import trusted whatever it opened.
- The import was not rerun-safe: reprocessing a delivery doubled rows, so recovery from any incident meant manual database surgery.
- Server session logs existed but had never been consulted; nobody had ever compared a "corrupt file" ticket against the matching session's completion record.
The retrofit, ranked by effort
- Same day, consumer only: tighten the trigger to
*.csvand add a settle step sized from measured transfer times. Claim files into a working folder before importing. In a folder-monitoring harness like Sysax FTP Automation this is configuration work. The monitor watches for arriving files and runs the task chain, with retry, error handling, and an email when something will not settle. - One email, producer's five minutes: ask the partner to add a fourth upload — an empty
feed_YYYYMMDD.donesent after the three data files complete. The trigger then moves to the marker, and per-file races disappear along with the batch-incompleteness problem. - One sprint, both sides: upgrade the marker to a control file listing each file's byte size and hash. Verify all three files against it before the import starts, quarantining mismatches with an alert. The verification habits are exactly those from verifying transfers end to end.
- The larger project: make the import idempotent. Load through a staging table keyed on the delivery date so a rerun replaces instead of duplicates. Follow the patterns in our duplicate detection and idempotency series. This is the layer that turns any residual miss from an incident into a retry.
The flow afterward
Partner uploads three files in any order, then the control file, last. The monitor fires on the marker only. The job verifies sizes and hashes, claims the set, and runs the idempotent import. It then archives the files, removes the marker, and mails a one-line success summary. A killed upload now leaves either an unreferenced data file (swept and alerted at noon) or a delivery that fails verification loudly at 02:15. The session log answers, in one lookup, exactly how many bytes the failed transfer delivered. The month-end "corrupt file" tickets stopped, because the race they described no longer exists.
Honesty about the residual risk, because there always is some: the partner could still send the marker early if their tooling changes. That ordering lives in the written contract now, and the session log makes spot-checking it a two-minute lookup. But the ordering is a promise rather than a mechanism. And verification is only as good as the control file's own accuracy. Neither residue justifies more engineering for this flow; both justify the periodic re-check that the next section turns into a habit.
Ranking Retrofits When You Cannot Do Everything
Most estates have more risky flows than project capacity. The ordering that pays best, as a rule:
- Consumer gates first. Pattern filtering, settle checks, claim-by-move: cheap, entirely under your control, and they cut the everyday race immediately — even before any partner conversation happens.
- The contract second. A marker or temp-name convention costs the sender minutes and converts your inference into their declaration. Most partners say yes; the ones who cannot tell you which consumer defenses must stay load-bearing.
- Verification third. Control totals and hashes close the failure modes gates cannot see — abandoned partials, hybrids, silent truncation.
- Idempotency last but scheduled. It is the most engineering effort and the layer that makes every other layer's rare failure survivable. Flows feeding financial or customer-facing data justify it first.
Apply the same order across flows: fix the flow whose partial file costs the most, not the one easiest to fix. A half-written internal log copy is a nuisance; a half-loaded billing extract is a bad quarter.
And harvest the free wins while the projects queue. Three of the matrix's rows cost an afternoon across an entire estate. Tighten every watcher's pattern to final names only. Schedule one sweep that quarantines and reports stale .part files and orphan markers everywhere. Confirm that transfer logging is switched on and retained. None of them requires touching a producer, a partner, or an import. Each one converts a class of silent failure into a visible one, which is most of the battle in this subject.
Keeping It Safe as Things Change
Partial-file safety decays, because it lives in the seams between systems and the seams are what change. Three habits preserve it:
- Re-verify after every tooling change. A sender upgrading their client, a server migration, a new monitor — any of these can silently change upload ordering, temp-name behavior, or timestamp handling. Re-run the two tests; it is twenty minutes.
- Make new flows inherit the checklist. The cheapest audit is the one done at design time. Put the checklist into the template every new integration starts from, next to naming and retry and error handling standards.
- Watch the leading indicators. Settle-check resets climbing, first-attempt failures that succeed on retry, stale partials appearing in sweeps — each is the race being survived, not absent. Treat them as smoke, and read the logs while the trail is warm.
The Series in One Breath
A file is not a fact until its writer is finished, and nothing about a filename says so. So producers should make completion explicit with atomic renames or markers. Consumers should gate on settlement and verify before trusting. Everyone should assume transfers can die midway. Each flow deserves the audit this page turns into an afternoon's work. Pick your riskiest flow, print the checklist, and start with the boxes that cost the least.
Frequently Asked Questions
Which single control gives the most protection for the least effort?
Do I really need controls in all three seats?
How often should a flow be re-audited?
What if the producer is a partner who will not change anything?
What is the fastest sign an existing flow is at risk?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
