Marker and Control Files: Signaling Done
The core problem of this series is a missing bit of information: nothing about a file's presence says whether its writer is finished. The atomic-rename pattern smuggles that bit through the filename itself — the final name appears only when the file is complete. But renames cannot solve everything. They cannot make a five-file batch appear as one unit. They are not available to every sender's tooling. And sometimes the consumer needs more than "done" — it needs "done, and here is how many records to expect."
The second producer-side pattern covers those cases: send the "done" signal as a small second file. This article explains the marker-file pattern in full. It covers the ordering rule that makes it trustworthy and the flavors from empty flag to checksum-bearing control file. It covers the cleanup discipline that keeps folders sane. It also covers the orphan scenarios you must design against before they design your incident reports. It is part of our Partial-File Safety series.
The Done-File Pattern in Plain Words
A marker file is also called a done file, ready file, flag file, or sentinel. It is a second file whose only job is to announce that a first file is complete. The producer writes the data file — feed_YYYYMMDD.csv — and, once it is entirely written and closed, creates the marker — feed_YYYYMMDD.done. The consumer ignores data files entirely and watches for markers. When the marker appears, the data file it names is, by contract, complete and safe to read.
Notice what moved: the "am I finished?" question is no longer answered by inspecting the data file. As the opening article showed, the data file cannot answer it. Instead, the answer comes from the existence of an object the producer only creates after finishing. The data file may take forty minutes to arrive; nobody cares, because nobody is watching it. The marker takes a fraction of a second to create, so its own write window is negligible.
That is the entire pattern. Its power is its portability: it needs no rename support, no shared filesystem, no special server behavior. It just needs the ability to create one more small file, which every producer on earth has.
The everyday analogy is the delivery note. Goods arrive at a warehouse over hours, pallet by pallet. Nobody starts unpacking until the driver hands over the signed note that says the shipment is complete. The note is small, it comes last, and its arrival — not the sight of pallets — is what starts the work. Everything else in this article is the discipline that keeps the note honest.
Ordering: The One Rule That Makes Markers Work
The marker's meaning comes entirely from when it is created. The rule: the marker is written after the data file is completely written, closed, and — for uploads — confirmed transferred. The full producer sequence:
- Write or upload the data file completely.
- Confirm success — the local write closed cleanly, or the transfer command returned success.
- Only then create the marker, as a separate, second operation.
Every marker-pattern failure in the wild traces back to a violation of that ordering, and the violations are sneakier than they sound:
- Parallel uploads. A sender's client is configured to transfer several files at once. So the 2 KB marker finishes long before the 2 GB data file it vouches for. The consumer fires, reads a partial file, and the marker told it to.
- Batch transfers with sorted ordering. A sender queues the whole folder in one job, and the tool uploads in name order. Whether the marker goes first or last now depends on alphabetical accident. The name
feed_YYYYMMDD.donehappens to sort after.csv, but rename the marker_readyand it sorts first. A contract that depends on sort order is not a contract. - Marker written early "because it's tiny." A script creates the marker up front intending to fill it in later, or touches it before the data write "to reserve the name." The consumer does not read intentions.
The fix in each case is the same: the marker travels alone, in a second step that only runs after the first step's success is confirmed. In a script, that is a second command after checking the first one's exit code. In a transfer tool, it is a second task chained on success of the first. Any decent automation harness is built to express that shape, including the job chaining in a tool like Sysax FTP Automation.
One refinement for the cautious: the marker is a file too, and can be read mid-write like any other. Keep it empty or a single line so the exposure is microscopic. Or close the loop entirely by writing the marker with the temp-name-and-rename pattern. That costs one extra line and makes the signal itself race-free.
Remember: a marker is a promise about the past — "the data file was complete when I was created." Anything that lets the marker exist before that moment (parallel uploads, sorted batches, touch-first scripts) turns the promise into a lie the consumer will act on.
Marker Flavors: From Empty Flag to Control File
Markers differ in how much they say. The empty flavor answers one question; the richer flavors — usually called control files — answer the next questions the consumer was going to ask anyway:
| Flavor | Typically contains | What the consumer can now verify | Notes |
|---|---|---|---|
Empty marker (.done, .ok, .ready) |
Nothing — existence is the message | That the producer declared the write finished | Simplest to produce; proves completion of the write, not correctness of the content |
Control file with totals (.ctl) |
Byte size, record count, business date token | That the data file on disk matches the size and count the sender intended | Catches truncated transfers the empty marker would bless |
| Checksum or manifest file | Hashes, and for batches a list of member files with sizes | Bit-for-bit integrity, and completeness of a whole multi-file set | The strongest form — see our guide to checksum files and manifests |
| Trigger file with instructions | Routing hints or processing parameters for the batch | Completion, plus what to do next | Validate its contents like any untrusted input before acting on it |
The step up from empty marker to control file is worth taking for any flow that matters, because it closes the pattern's one blind spot. An empty marker proves the producer finished its procedure. But if the producer's own upload was silently truncated and its tooling did not notice, the marker still gets written. A control file carrying the expected size or hash lets the consumer verify the bytes rather than trust the ceremony. The wider discipline is covered in verifying transfers end to end.
When Markers Beat Renames — and When They Do Not
Reach for markers instead of (or on top of) renames when:
- The delivery is a set, not a file. A rename commits one file at a time; nothing makes twelve files appear atomically. One marker — ideally a manifest listing all twelve names and sizes — gates the entire batch, so the consumer starts only when the set is whole.
- The sender's tooling cannot rename. Constrained upload clients, appliances with fixed firmware behavior, and locked-down drop accounts that may create files but not rename them can all still create one extra file.
- The handoff crosses filesystems or systems. Where a move cannot be atomic — different volumes, staged copies between servers — the done bit can no longer ride on the name. In those cases, it rides in a marker instead.
- The consumer needs metadata anyway. If counts, dates, or hashes must travel with the data, a control file carries the completion signal and the verification data in one object.
- Auditors and operators need visible evidence. A marker is a plain, human-readable artifact of "the sender declared this complete" — easier to point at than an invisible rename.
Prefer the plain rename when a flow is one file at a time and you control the producer. It has no second object to create, order, clean up, or orphan — fewer moving parts, nothing to drift. And the two patterns combine naturally: rename each data file into place as it completes, then drop one marker when the whole batch is present. Belt, then suspenders.
The Consumer's Half: Gate, Verify, Consume
On the consuming side the pattern inverts the watch: trigger on markers, never on data. A watcher pointed at *.done could be a script, a scheduled sweep, or a folder-monitoring tool such as Sysax FTP Automation. It watches for arriving files that match the marker pattern and runs the import task when one lands. Such a watcher simply cannot fire early, because the trigger object does not exist until the producer says so. The consumer then derives the data filename from the marker name, verifies whatever the marker lets it verify, and only then processes. A skeleton of the whole gate in shell:
# Trigger on markers only; derive, verify, claim, process
for marker in /data/inbox/*.done; do
[ -e "$marker" ] || continue # no markers this sweep
data="${marker%.done}.csv" # exact-name derivation
if [ ! -e "$data" ]; then # marker without data: park it
mv -- "$marker" /data/error/ && alert "orphan marker $marker"
continue
fi
expected=$(cat "$marker") # control file carries byte count
actual=$(stat -c %s "$data")
if [ "$expected" != "$actual" ]; then # declared size must match disk
mv -- "$data" "$marker" /data/error/ && alert "size mismatch $data"
continue
fi
mv -- "$data" "$marker" /data/work/ # claim the pair, then process
import /data/work/"$(basename "$data")" && archive_pair "$data" "$marker"
done
Two details in that sketch carry most of the safety. The derivation must be exact. One character of drift between Feed_YYYYMMDD.csv and feed_YYYYMMDD.csv, or a case-sensitivity mismatch between systems, and every delivery becomes an orphan. And the consumer claims the pair by moving it to a working folder (same volume, so the move is atomic) before processing. That way, a re-scan during a slow import does not pick the same delivery up twice. The surrounding workflow — inbox, working, done, and error areas — is the anatomy our watch folders and event-driven transfers series covers in depth.
Cleanup Discipline: Who Deletes What, and When
A marker's life must end, and the end deserves as much design as the beginning. The consumer owns cleanup — it is the only party that knows processing succeeded. The tidy lifecycle is: claim the pair, process the data, archive the data, remove the marker. Keep that order, with the marker's removal as the final receipt.
Crash windows make the order matter. Remove the marker before processing and a crash mid-import strands a claimed data file with no signal pointing at it. Recovery then requires someone to notice a lonely file in the working folder. Remove it after, and a crash between processing and cleanup leaves a marker that will re-fire on the next sweep. The delivery then gets processed twice. There is no ordering that eliminates both windows; that is a small instance of a deep distributed-systems truth. The practical resolution is to accept possible re-fires and make processing safe to repeat. That is idempotent, in the vocabulary of our duplicate detection and idempotency series. The reason: "processed twice, harmlessly" is a far better failure than "never processed, silently."
Producers need one cleanup rule of their own: never reuse a marker name for a new delivery until the old pair is gone. Datestamped, per-delivery names — feed_YYYYMMDD_HHMMSS.done rather than an eternal feed.done — sidestep the whole class of stale-signal bugs. They pay the usual dividends that make naming conventions worth designing once and enforcing forever.
Orphans: Markers Without Data, Data Without Markers
Design for the two mismatch states before they occur, because both will:
- Data without a marker is normal for a while — the data file lands first by design. It becomes a problem with age. A data file still unaccompanied after your longest plausible producer cycle means the producer died between the two steps, or the marker went to the wrong place. Alert on age, not on existence. The file itself is likely complete, but nothing proves it. Do not be tempted to process it anyway without the verification steps you would demand of any unsignaled file.
- A marker without data is never normal — it means the contract broke. The cause might be a naming mismatch, a cleanup bug that removed data but not marker, or a sender that shipped the receipt without the goods. Park the marker in the error folder immediately and alert; do not leave it in the inbox to confuse the next sweep.
Both cases belong in the same operational machinery as every other intake failure. That means quarantine moves, alerts a human reads, and a periodic sweep for anything stuck. These are patterns that our retry and error handling series builds out. Orphans are not embarrassments; they are the pattern working, converting silent timing bugs into visible tickets.
Writing the Convention Down
The marker pattern is a two-party protocol, and undocumented protocols rot. One page is enough, and it should answer exactly these questions:
- Marker naming: data file
feed_YYYYMMDD.csvpairs withfeed_YYYYMMDD.done— exact case, exact extension. - Ordering: marker is sent only after the data transfer completes successfully, as a separate step — never in the same parallel batch.
- Contents: empty, or byte count and record count, one per line (agree the exact format).
- Cleanup: consumer removes the pair after successful processing; sender never resends a marker without its data file.
- Failure signaling: what the sender does if generation fails after upload started (send nothing, or send an agreed
.errfile — decide, and write it down).
Then verify the contract is being honored, occasionally and after every sender-side change. Server-side transfer logs are the ground truth. A server like Sysax Multi Server logs every upload session to file and to a database. So you can check that the marker's upload genuinely started after the data file's upload finished. Timestamps settle the argument that neither side's assurances can.
The Signal, Recapped
A marker file moves the "done" bit into an object that only exists once it is true — provided four conditions hold. The ordering rule is sacred, and the names derive exactly. The pair's lifecycle has one owner, and orphans are parked loudly instead of ignored. Use markers where renames cannot reach: multi-file sets, constrained senders, cross-system handoffs, and any flow where the consumer should verify counts or hashes before trusting its luck. From here, settle checks covers the consumer left alone with neither renames nor markers. The series capstone, the partial-file safety checklist, puts every control — this one included — into a single audit matrix.
Frequently Asked Questions
Should the marker file be sent first or last?
What should go inside a marker file?
What do I do with a marker whose data file is missing?
Are markers better than atomic renames?
Who is responsible for deleting the marker?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
