Home › Topics › Multi-Site Distribution › Distribution Proof

Proving Every Site Got the Right Version

At half past eight the operations manager asks the only question that matters: "do all the branches have today's prices?" There are two ways to answer. One is a feeling — the jobs ran, nothing obviously failed, probably yes. The other is a report: forty rows, one per site, each naming the release the site verified and when, with any exceptions at the top. Distribution estates run on feelings for years, right up until the morning a feeling is wrong at the worst possible branch.

This article builds the second answer. It is a design, not a product. It has four small mechanisms — hashes, a release manifest, version markers, and receipts — plus one report that reads them. It has alerting that names stragglers while there is still time to act. Every piece is buildable with a transfer server's logs, a few text files, and modest scripting, and each piece earns its place independently. This article is part of our Multi-Site Distribution series and pairs closely with the state-tracking design from hub designs for branch distribution.

The Claim, Taken Apart

"Every site got the right version" sounds like one claim. It is four, and each needs its own evidence:

  1. Complete and uncorrupted — every file of the release arrived at the site, whole, with the bytes the hub published. Evidence: hashes checked against a manifest, at the site.
  2. The right version — what the site holds is the release you meant, not last week's, not a half-superseded mixture. Evidence: version markers.
  3. In time — it was there before the deadline that made anyone care. Evidence: timestamps on the receipt and in the hub's logs.
  4. In use — the site actually switched to it, rather than holding it unactivated in a folder. Evidence: the site's own state marker, written after activation.

Notice where the truth lives: at the site. The hub can prove it sent; only the site can prove it has. That is why every layer below pushes the checking outward to the branch and then brings a small piece of evidence back. That is the shape of every honest verification design. It is the reason "the transfer job reported success" will never be enough on its own. A job success is a claim about a conversation, not about the disk at branch-017.

How much of this machinery a flow deserves is proportional to what a wrong answer costs — the "failure meaning" line from the census sheet in the problem-shapes article. A training-video drop that can be a day late at a site can live on layers one and two and a weekly glance. A price file whose absence stops trading, or a recall notice with legal weight, deserves the full stack, receipts and alert ladder included. Build the layers in order and stop where the flow's stakes stop; the design below is a menu with a sensible sequence, not an all-or-nothing rite.

Layer One: Integrity — the Bytes Are Right

The foundation is per-file integrity. A hash is a short fingerprint computed from a file's contents. The same bytes always produce the same fingerprint. Any change — truncation, corruption, the wrong file under the right name — produces a different one. (New to the idea? The friendly full story is in hashing explained.) Size comparison catches gross truncation cheaply; a hash comparison catches everything else. The hub computes hashes at publish time; the site recomputes after download and compares. Match: the bytes are right. Mismatch: delete, retry, and if it persists, alert — never activate a file that failed its hash.

Two habits make integrity checking effortless instead of a chore. First, hash once at publish, not per site — the release's hashes are facts about the release, computed one time and shipped along with it. Second, make the site-side check part of the download job itself, so verification cannot be forgotten independently of the transfer. A pulled release that fails verification should look, to the rest of the machinery, exactly like a release that never arrived.

Layer Two: The Release Manifest

Per-file hashes need a container, and that container solves a second problem for free. A manifest is a small text file, generated at publish time, listing every file in the release with its size and hash:

hub:/distribution/prices/releases/rel_YYYYMMDD/MANIFEST.txt

release: rel_YYYYMMDD
flow: nightly-prices
files: 3
------------------------------------------------------------
prices_YYYYMMDD.csv        4188220   sha256  9f3c62d0a1...
promos_YYYYMMDD.csv         811034   sha256  4b0e77c19d...
notes_YYYYMMDD.txt            2210   sha256  d21f80aa3e...
------------------------------------------------------------
published: Mar 14 01:38 hub time

The site's job downloads the release, then verifies every line of the manifest: file present, size matches, hash matches. Only a fully verified set counts as delivered. The manifest is written into the release folder last — after every content file is in place. So its presence doubles as a completeness signal. A puller that arrives mid-publish and finds no manifest simply comes back later, immune to the half-published-release trap. Formats, tooling, and the command-line side of manifests are covered in depth in checksum files and manifests. The design point here is that the manifest is the release's identity card, and everything downstream refers to it.

Layer Three: Version Markers

Hashes prove bytes; version markers prove versions. Two tiny files carry the whole scheme:

  • The hub's current marker. This is a one-line file — current.txt containing rel_YYYYMMDD — sitting beside the release folders. It is rewritten (atomically, and only after the release and its manifest are complete) to point at each new release. It is the single authoritative statement of what "current" means, and every site and every report reads it rather than guessing from folder names.
  • The site's state marker. After a branch has downloaded, verified, and atomically swapped the release into use, its job writes its own one-line marker. The marker says "I am on rel_YYYYMMDD, verified Mar 14 02:41" and goes into a known local path. This is the site's testimony about itself, written only at the end of a fully successful sequence, so its content is trustworthy by construction.

Comparing the two markers answers the version question site by site. The hub says current is X; branch-014 says it is on X — current. Branch-031 says it is on W — stale, and you know exactly how stale, because W names its own vintage. Marker files as a coordination idiom — why they must be written last, why atomically, what belongs in them — get a full treatment in marker and control files.

Remember: the marker is written after the swap, never before, and written atomically. A marker that goes down before verification finishes is a forged receipt. It will one day claim currency for a site holding a corrupt half-release. The report built on that marker will lie with confidence.

Layer Four: The Receipt Comes Home

The site now knows its own truth; the hub needs to hear it. The receipt is the site's state marker traveling home: after activation, the branch job uploads a small file into its own folder on the hub.

hub:/distribution/prices/receipts/branch-014/rel_YYYYMMDD.rcpt

site: branch-014
flow: nightly-prices
release: rel_YYYYMMDD
manifest: verified, 3 of 3 files ok
downloaded: Mar 14 02:29 to Mar 14 02:38 site time
activated:  Mar 14 02:41 site time
status: current

Receipts make the hub's evidence positive. Instead of inferring health from an absence of errors, the hub holds an affirmative statement per site per release. The statement is timestamped from the only machine that could know. The transfer server's own records corroborate the receipts. On a hub running Sysax Multi Server, each branch authenticates as its own account with write access only to its own receipts folder. The server's activity log is written to file or to a database. It independently records which account downloaded which release files and uploaded which receipt, minute by minute. Receipt plus matching log line is delivery evidence of a quality that ends arguments. In the Pro and Enterprise editions, event triggers can act on the upload itself. They can fire your fold-receipts-into-state step the moment a receipt lands, rather than waiting for the next sweep.

In push designs the receipt is still worth having, but the pusher's own task result carries more of the weight. The hub-side task verified the upload and logged the outcome. The trade is described in the hub-designs article; either way, what feeds the next layer is one trustworthy record per site per release.

Plan for the awkward case where the delivery succeeded but the receipt could not travel. The branch verified and activated, then its link died before the upload. The site is healthy; the evidence is missing. Treat "verified locally, receipt pending" as its own state. The branch job retries the receipt upload on its next connection. The hub's report shows the site as unconfirmed rather than failed, with the server's download log as circumstantial support in the meantime. The distinction sounds fussy until the first morning someone nearly re-pushes a release onto forty healthy branches. This happens because a handful of receipts were stuck behind a flaky link.

The Distribution Report

Now the layer everyone actually sees. The distribution report is a small program that runs after the window closes (and on demand). It reads the hub's current marker, the receipts, and the per-site state records. It emits one row per site:

DISTRIBUTION REPORT — nightly-prices — Mar 14, generated 06:30 hub time
hub current release: rel_YYYYMMDD          sites: 40

EXCEPTIONS FIRST
site        expected        reported        verified      status
branch-031  rel_YYYYMMDD    (previous rel)  --            STALE: retrying, 3rd attempt
branch-007  rel_YYYYMMDD    --              --            ABSENT: no connection this window

ALL CLEAR (38 sites)
branch-001  rel_YYYYMMDD    rel_YYYYMMDD    Mar 14 02:12  current
branch-002  rel_YYYYMMDD    rel_YYYYMMDD    Mar 14 02:14  current
...
summary: 38 current / 1 stale / 1 absent   deadline: 90 min away

Three editorial rules make the report worth reading. Exceptions first — the healthy thirty-eight are a footnote; the two problems are the news. Absence is a status, not a blank — a site that never connected gets a row saying so. That row is produced by comparing the site roster against who reported, using the expected-versus-actual technique from freshness checks and expected files. And every row names its evidence — release identifiers and timestamps, not adjectives. That way, any row can be traced to a receipt and a log line when someone asks.

Be clear-eyed about what this is: something you build. Transfer servers and automation tools hand you the raw material — activity logs, task results, notification hooks. But the forty-row answer with your sites, your releases, and your deadlines is a modest script you own. It is a loop over the roster, a read of markers and receipts, a sort, and an output. Text in an email is a fine first version. When the estate wants history and charts, the presentation options are surveyed in transfer dashboards and reporting.

Keep the reports, not just the latest one. A dated copy per window turns the report into a history. The history answers questions the single morning snapshot cannot. Branch-031 stale three windows running is a pattern with a cause — a dying disk, a broken scheduler, a link that changed class. One stale morning is just weather. A monthly skim of exception rows per site is the cheapest capacity-planning and hardware-failure radar a distribution estate gets. It costs one folder of small text files.

Straggler Alerting: Before It Matters, Not After

The report tells the morning story; alerting exists so the story has a happy ending. The straggler may be late, failing, or absent, in the vocabulary of the thin-links article. It should surface while the night still has margin to spend, on a ladder that matches urgency to noise:

  1. Grace, silently. Inside the window, retries and late waves are the plan working, not news. No alert fires because a thin-class site is being slow within its budget.
  2. Notify at window close. Any site not verified when its window shuts goes to the operations channel and the morning report. Include its state, its retry count, and the minutes of margin left. On the push side, task failures raise their own hand. Automation built on Sysax FTP Automation sends email notifications when a task exhausts its retries. That makes the exhausted-budget moment loud without any extra plumbing.
  3. Escalate on deadline risk. When remaining margin drops below the time a fix realistically takes — a site still absent an hour before opening — a person gets woken. The alert names the site, the release, the last evidence seen, and the local deadline. An alert that requires research before action is half an alert.

Two disciplines keep the ladder trustworthy. Alert on evidence, not vibes: the trigger is "no verified receipt for branch-007 and margin below sixty minutes," never "the job looked slow." And watch the watcher. The report generator and the absence check are themselves jobs that can silently die. So give them the same heartbeat treatment described in monitoring the monitoring. A verification layer that fails silently is worse than none, because it spends its credibility precisely when the fleet needs it.

Re-Verification: Trust, but Re-Hash Occasionally

Everything so far verifies a release at delivery time. Disks and people keep acting after delivery. A file gets deleted "to free space." A site machine is restored from an old image and silently rolls back to a stale release while its marker still claims currency. A slowly failing drive corrupts a file that verified perfectly in spring. Release-night machinery never notices any of this, because nothing transferred.

The countermeasure is periodic re-verification: a scheduled site-side job — weekly is plenty for most estates. It re-checks the local release against its manifest (presence and sizes always; full re-hash on a sampling basis or after any restore). It re-writes the site marker with a fresh verified-at timestamp. Stale verified-at values then surface in the same morning report through the same plumbing. The restored-from-backup site betrays itself within a week instead of at the worst possible audit. The cost is a few minutes of disk reading per site per week. The payoff is that your report describes the present, not the night everything last went well.

When Someone Asks Months Later

The same machinery that answers "are we current today?" answers the harder question that arrives long afterward: "prove branch-022 had the recall notice before trading on the fourteenth." The answer is an evidence pack assembled from what you already keep. It includes the release's manifest (what the files were, hash by hash), and the hub's current marker history (what was official when). It includes branch-022's receipt (downloaded, verified, activated, timestamped), and the server's activity log rows corroborating the session. Retain receipts and activity logs for as long as such questions can arrive — that is a retention-policy decision, not a technical one. With those records retained, the pack takes minutes to produce. The general craft of delivery evidence, including when signatures and stronger non-repudiation are worth adding, is covered in proof-of-delivery patterns.

The Design on One Page

Publish with hashes in a manifest, written last. Point at the release with an atomic current marker. Verify at the site, swap atomically, and only then write the site marker. Send the receipt home to a per-site folder the site alone can write. Fold receipts and logs into per-site state; read the state into an exceptions-first report after each window; alert on evidence up a ladder that respects margin. Retain what you would need to prove it later. Six sentences, four small files, one report — and the operations manager's question gets a forty-row answer instead of a feeling. The worked forty-branch design shows all of it installed in a realistic estate, straggler drama included.

Frequently Asked Questions

Why isn't a successful transfer job enough proof of delivery?
Because the job's success describes a network conversation, not the disk at the site. Corruption, a later local deletion, a half-applied release, or files delivered but never activated all hide behind "job succeeded." Proof needs site-side verification against a manifest, plus a receipt sent back.
What goes in a release manifest?
One line per file — name, size, and hash — plus the release identifier and a publish timestamp. The hub generates it once at publish time and writes it into the release folder last. So its presence also signals that the release is complete and safe to fetch.
What is a version marker and why write it last?
A one-line file naming a release: the hub's marker says what is current, and each site's marker says what it runs. Each is written atomically and only after everything it vouches for is finished. A marker written early can claim success for a broken state, and every later report inherits the lie.
What should a distribution report show?
One row per site comparing the expected release against what the site reported, with verification timestamps and a status — current, stale, failed, or absent. Exceptions belong at the top, absent sites get explicit rows, and every row should trace back to a receipt and a log entry.
How do I catch the site that never even connected?
With an absence check: after the window closes, compare the full site roster against the receipts and log entries that actually arrived. A site with no evidence gets an "absent" row and enters the alert ladder. Error-based monitoring alone misses this case completely, because nothing failed — nothing ran.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.