Size-Stability and Settle Checks: The Consumer's Defense
The first two fixes in this series live on the producer's side of the handoff. Write under a temp name and rename when done, or announce completion with a marker file. Both require the producer to cooperate. Frequently, it will not. The sender is a partner whose tooling is fixed, or an appliance with firmware behavior nobody can change. Or it is a legacy application that has written straight to its final filename for years and always will.
What remains is the consumer's defense: watch the file itself and infer, from what you can observe, that its writer is probably finished. That family of techniques — settle checks — is the subject of this article. It explains how size-stability polling works, how to tune it for real senders, and where minimum-age rules and open-handle checks fit. It also covers, told honestly, the cases where all of them can be fooled. This is part of our Partial-File Safety series, and it is the article to reach for when you cannot change the other side.
The Problem from the Consumer's Chair
From the consuming side of a folder, a file offers exactly four observable facts. These are its name, its size, its timestamps, and — sometimes — whether any process currently holds it open. None of these says "finished." The producer's intent is simply not recorded anywhere you can read; that is the root problem this whole series exists to solve.
A settle check turns those observable facts into an inference: if the file has stopped changing for long enough, its writer has probably finished. The word "probably" is doing real work in that sentence, and this article will not pretend otherwise. Producer-side signals are proofs — a rename or a marker is the producer saying "done." Settle checks are evidence — the file holding still while you watched. Good engineering treats them accordingly: as the primary gate only when nothing better is available, and always backed by verification.
It helps to see how this relates to the watcher that noticed the file in the first place. A folder watcher — whether it polls or receives filesystem events — answers the question "did something change?" A settle check answers the opposite question: "has it stopped changing?" The first starts the clock; the second decides when the clock has run long enough. Watchers alone fire at the worst possible moment, the beginning of the write. A settle check inserted between the event and the action moves the trigger to the end of it.
The Size-Stability Check
The workhorse of the family: sample the file's size repeatedly, and declare it settled only after several consecutive samples agree. One growing file, observed every fifteen seconds:
02:10:15 size 12,582,912 changed - keep waiting, reset the count 02:10:30 size 26,214,400 changed - keep waiting, reset the count 02:10:45 size 44,040,192 changed - keep waiting, reset the count 02:11:00 size 50,331,648 stable x1 02:11:15 size 50,331,648 stable x2 02:11:30 size 50,331,648 stable x3 - settled: verify, then process
Three ideas are packed in there. First, stability is measured in consecutive agreements, not one lucky pair — a single unchanged reading can be a producer catching its breath. Second, any change resets the count: the file must earn its way out by holding still for the whole required stretch. Third, settling does not mean processing — it means the file has qualified for verification. Whatever completeness evidence you have (expected size, record count, checksum, a sane trailer line) gets checked before the import runs. Watching the modified time alongside the size costs nothing and catches the rare writer that changes bytes without growing the file.
One honesty note for folders that live on a network share: the size you sample may itself be slightly stale. Client machines cache directory metadata for a short while. So two "identical" readings a few seconds apart can both be answered from the same cached snapshot while the file quietly grows on the server. The cure is the same knob you already have — longer poll intervals. That puts consecutive samples far enough apart that the cache has genuinely refreshed between them. Shares deserve WAN-grade settings even when the wire is fast.
The diagram below shows the same logic as a decision loop. That is the shape you will implement with a shell script, a scheduled job, or a setting in a monitoring tool.
The red exit matters as much as the green one. A file that never settles within your maximum wait is not a scheduling nuisance — it is a finding. That could be a stalled transfer, a runaway writer, or a sender behaving in a way your parameters do not expect. Park the file in an error folder and tell a human, the discipline our retry and error handling series builds out.
A Worked Algorithm with Tunable Parameters
Here is the loop as a small, copyable script. It is deliberately parameter-first: the four numbers at the top are the whole tuning surface.
#!/bin/sh
# settle.sh FILE -- exit 0 once FILE has held still long enough
FILE="$1"
POLL_SECS=15 # gap between size samples
STABLE_ROUNDS=3 # consecutive unchanged samples required
MIN_AGE_SECS=60 # never release a file younger than this
MAX_WAIT_SECS=1800 # give up, park, and alert after this long
first_seen=$(date +%s)
stable=0; last=-1
while :; do
size=$(stat -c %s "$FILE" 2>/dev/null) || exit 2 # vanished: renamed or removed
if [ "$size" = "$last" ]; then
stable=$((stable + 1))
else
stable=0; last=$size
fi
age=$(( $(date +%s) - first_seen ))
if [ "$stable" -ge "$STABLE_ROUNDS" ] && [ "$age" -ge "$MIN_AGE_SECS" ]; then
exit 0 # settled: caller verifies, then processes
fi
[ "$age" -ge "$MAX_WAIT_SECS" ] && exit 3 # never settled: caller parks + alerts
sleep "$POLL_SECS"
done
Note that age is measured from first seen — the moment your watcher noticed the file — not from the file's own modified time. The next section explains why that choice is deliberate. What each knob trades away:
| Parameter | What it controls | Set too low | Set too high | Starting point |
|---|---|---|---|---|
POLL_SECS |
Gap between samples | Short producer pauses read as "settled" | Every file waits longer than needed | 10-30 s local, 30-60 s over a WAN |
STABLE_ROUNDS |
Agreements required in a row | One pause fools the check | Added latency on every delivery | 2-4 |
MIN_AGE_SECS |
Floor under fast small files | A tiny file races through mid-stall | Small files pay a flat delay | About one typical transfer duration |
MAX_WAIT_SECS |
When waiting becomes an incident | False alarms on slow days | Stuck files linger unnoticed | 3x your longest normal transfer |
The one relationship to memorize: POLL_SECS × STABLE_ROUNDS must exceed the longest pause your producer ever takes mid-write. That product is the stillness you demand. A producer that can stall for ninety seconds while its source database thinks must face a settle window longer than ninety seconds. Otherwise, it will slip a truncated file through.
Minimum-Age Rules — and the Timestamp Trap
The simpler cousin of size polling is the minimum-age rule: only pick up files whose age exceeds some threshold. It is cheap, easy to bolt onto an existing sweep ("process files older than ten minutes"), and often good enough for flows with predictable, quick writes.
The trap is what "age" is computed from. The obvious source — the file's modified time — lies to you in two common ways. First, many transfer and copy tools preserve the source file's timestamps. The file lands with an mtime from hours or days ago. It looks ancient the instant it arrives, and sails past any age rule while still growing. Second, clocks differ: an mtime stamped by another machine is only as good as that machine's clock. The robust choice is the one the script above makes — age from when you first saw the file. Record that first-seen time in your own ledger, on your own clock. Modified-time stability remains useful as a supplementary signal; modified-time age is the part that misleads.
Remember: a settle check proves stillness, not completeness. A file can hold perfectly still because its writer finished — or because its writer is stalled, throttled, or dead. Stillness qualifies a file for verification; it never replaces it.
Open-Handle Checks: What They Can and Cannot Tell You
The third observable is whether anything currently has the file open. On Unix-family systems the long-standing tool is lsof ("list open files"). On Windows, the idiomatic probe is attempting to open the file exclusively and seeing whether the system refuses:
# Unix: is any local process holding the file open?
lsof -- /data/inbox/feed_YYYYMMDD.csv && echo "in use right now"
# Windows PowerShell: try an exclusive open; a sharing violation means "in use"
try { $h=[System.IO.File]::Open('D:\inbox\feed_YYYYMMDD.csv','Open','Read','None'); $h.Close(); 'no local writer at this instant' }
catch { 'locked - a writer holds it' }
Used honestly, these are one-way signals. A positive result is strong evidence the file is still in use: hold off. A positive result means the file is open, or the exclusive open fails with a sharing violation. A negative result proves almost nothing, for two reasons:
- The check is point-in-time. It answers "is it open at this instant?" A producer that writes in open-append-close bursts is closed between every burst; probe in a gap and the file looks free. The writer can reopen a millisecond after your check succeeds.
- The check only sees the machine it runs on. When the folder is a network share, the writer's handle lives on the file server or on another client. Your local
lsofenumerates local processes; it cannot see a writer two machines away. Server-side tools can, but only if you can run them on the server.
The practical role that falls out: use handle checks as a veto, not a gate. Let size stability qualify the file, then let a positive in-use signal cancel the release. Never let a clean handle check shortcut the settle window.
The Honest Failure Cases
Settle checks make partial reads rare. These are the ways they still fail, so you can decide which ones your flow must additionally defend against:
- The long pause. The producer stalls for longer than
POLL_SECS × STABLE_ROUNDS— a slow upstream query, a saturated link, a garbage-collection stall — then resumes. The file settles, releases, and grows again after your import started. Defense: size the window from measured behavior (next section), and verify before processing. - The abandoned partial. A transfer dies midway and never resumes; the file sits at a stable partial size forever. To a settle check, "abandoned" and "finished" are identical — both hold still. This is the failure mode settle checks are structurally blind to. The defense is completeness evidence — an expected size or count from a control file, a checksum, a manifest — checked before processing. That is exactly what our guides to end-to-end verification and reliable transfer integrity lay out. The network-failure side of this story is the next article, when the transfer itself dies midway.
- The file that never settles. Append-forever files — live logs, files a producer holds open and trickles into all day — will ride the loop to
MAX_WAITevery time. That is not a tuning problem; it is the wrong pattern for that flow. Such sources need a producer-side handoff (rotate, then deliver) instead of a consumer-side wait. - Format-blind release. The settle check passes, the size even matches expectations, but the content is garbage. No settle check sees inside the file. Structural validation — parseable header, sane row counts, expected trailer — belongs in the pipeline right after release. That territory is covered by our pre- and post-processing series.
Tuning for Real Sender Behavior
Every number above should come from measurement, not folklore. Two sources tell you how your senders actually behave:
- Server-side transfer logs. If deliveries arrive through your own transfer server, its session log records when each upload started, when it finished, and how many bytes moved. That is the exact distribution you need for
MIN_AGEandMAX_WAIT. A server like Sysax Multi Server writes every upload session to file and to a database. So pulling "longest transfer this quarter" is a query, not an archaeology project. - Your own settle logs. Log every release with its wait time and every reset with its cause. A flow whose files routinely need one reset is healthy. A flow with five resets per file is telling you its producer pauses, and your window should grow.
Tune per flow, not globally. The partner on a fast private line and the one trickling over a congested link deserve different windows. A single global setting will be wrong for both. Keep the parameters where operators can see and adjust them — a config block, not constants buried in code.
For multi-file deliveries, settle the batch, not just each file. A sender that drops five related files over ten minutes has not finished when the first file settles. The simple generalization is folder quiescence. Require that no file in the delivery set has changed, and no new file has appeared, for the full settle window before the batch job starts. If the sender can be persuaded to add a marker for the set, so much the better. That is the cleaner solution from the marker-file article. The settle window then only has to protect against the marker arriving early.
Finally, give the loop a harness. A settle check is a policy. Something still has to watch the folder, run the check, launch the import on success, and park the file on timeout. It must retry sensibly and email someone when a file is stuck. That operational shell is what folder-monitoring automation provides. Sysax FTP Automation, for instance, watches for arriving files, runs tasks against them, and brings retry, error handling, and email notifications along. So your settle-and-verify step runs inside machinery that already knows how to fail loudly. The broader workflow design around watchers — debouncing, arrival contracts, error folders — is its own subject, covered in our watch folders and event-driven transfers series.
Layered, Not Lonely
The settle check earns its keep as the consumer's independent line of defense — and it works best as one layer among several. If you can get even partial producer cooperation, take it. A temp-name convention or a marker file turns inference into proof and lets the settle check fade into a backstop. Whatever the producer does, verify completeness before processing and make processing safe to repeat. Then a settle check that releases one truncated file a year becomes a nuisance caught by verification, instead of a quiet data corruption discovered by your customers.
Read next: when the transfer itself dies midway, for the failure mode that produces the abandoned partials settle checks cannot see. Also read the series capstone checklist that puts producer and consumer defenses side by side for any flow you need to audit.
Frequently Asked Questions
How long should a consumer wait before processing a new file?
Can I rely on an open-handle check instead of size polling?
What if a file never stops growing?
Is checking the modified time as good as checking the size?
Can a settle check ever pass on an incomplete file?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
