Home › Topics › Multi-Site Distribution › Worked Design

A Worked Multi-Site Distribution Design, End to End

The earlier articles in this series each built one layer: mapping the problem, choosing a hub design, picking job machinery over replication, surviving thin links, and proving delivery. This article assembles all of it into a single, complete design for a realistic estate — a forty-branch retailer with two distribution flows. Just as importantly, it shows the reasoning behind every choice. Read it the way you would sit in on a design review: not to copy the design wholesale, but to watch requirements turn into decisions. That lets you run the same moves on your own estate.

Nothing here is exotic. The whole build uses a hardened transfer server, scheduled jobs, text-file markers, and a modest reporting script. That is deliberate: a distribution estate is something the whole team must be able to operate at three in the morning. This article closes our Multi-Site Distribution series.

The Estate and the Two Promises

The company runs forty branches — branch-001 through branch-040 — spread across three timezones. Each has a small Windows machine in the back office and a broadband link of varying quality. Twenty-eight branches are on solid business-grade lines, and nine are on thin shop-grade connections shared with the tills. Three are on fragile links that drop for minutes at a time most nights. The hub is a Windows server in the head-office data center with a generous uplink. Two flows must be distributed:

  • The nightly price file. The pricing system exports prices_YYYYMMDD.csv — a few megabytes — about an hour after midnight, hub time. Every branch must be using the new prices when it opens, local time. A branch that opens on stale prices is an incident with a name attached. Shape: content publish, latest wins; a branch that misses a night must jump straight to the newest file.
  • The weekly content set. Roughly two gigabytes of shelf-label artwork, planograms, and short training clips, published Friday evening, and required to be in use fleet-wide at Monday opening. Shape: versioned release; completeness and exactness matter more than speed, and the previous set must remain available for rollback.

Two flows, two very different centers of gravity — the price file is small and urgent, the content set is large and patient. So they get separate jobs, separate windows, and separate alerting throughout. They share only the hub and the machinery patterns.

For honesty's sake, here is what the design replaced, because it is what most estates actually run. There was a copying script on the pricing server with the branch list pasted into it. There was a shared folder some branches mapped over the WAN and some did not. There was a headoffice inbox where branch managers asked whether the new artwork had "gone out yet." It worked, in the sense that most branches had most files most mornings — and nobody could name the exceptions until customers did. Every choice below exists to make the exceptions nameable before opening time.

The Shape of the Design

Five decisions, stated up front, define the architecture. Branches pull from the hub on schedules. The hub publishes into an immutable release layout with atomic markers. Pull slots are derived from each branch's local quiet hours, in waves by link class. Every delivery is verified at the branch against a manifest and confirmed by a receipt. A morning report reads the evidence and names exceptions. The diagram shows the whole machine at a glance.

Overview of the forty-branch design. A hub server on the left holds a release area with current-version markers and a receipts area. Three timezone groups of branches on the right pull releases in waves and send receipts back. One branch is marked offline and pulls the current release when it returns. A timeline along the bottom shows the eastern, central, and western pull windows staggered through the night.

The Hub

The hub runs Sysax Multi Server as an SFTP endpoint. Each branch gets its own account — branch-014 logs in as itself, nothing shared. Per-area permissions give read-only access to both flows' release areas, and write access to exactly one folder, its own receipts directory. Branches with static addresses also get IP allow rules, so a stolen credential alone is not enough to connect from elsewhere. Activity logging goes to a database as well as to file. That makes every later question — did branch-014 fetch the manifest, when, from which address — a query instead of a log-trawl. The folder layout is the series' standard release pattern, doubled:

/distribution/
  prices/
    releases/
      rel_YYYYMMDD/
        prices_YYYYMMDD.csv
        MANIFEST.txt          <- written last within the folder
    current.txt               <- names the live release; atomic swap
    receipts/
      branch-001/ ... branch-040/
  content/
    releases/
      rel_YYYYMMDD/           <- ~2 GB: artwork, planograms, clips
        ...files...
        MANIFEST.txt
    current.txt
    previous.txt              <- names the prior set kept for rollback
    receipts/
      branch-001/ ... branch-040/

Publishing is automated and boring, which is the highest compliment a publish step can earn. When the pricing system drops its export into a landing folder, a folder-monitoring task in Sysax FTP Automation picks it up and runs the sequence. It creates the new release folder and moves the file in under its datestamped name (the sortable-name discipline from datestamp formats that sort). It computes hashes, writes MANIFEST.txt last, then rewrites current.txt atomically. If the export never arrives by its expected time, the monitor's silence trips an alert. The night's first straggler check is actually aimed at the producer, not the branches.

Two sizing checks close out the hub design. Protocol: SFTP, because one encrypted port through forty branch firewalls is the whole firewall conversation, fleet-wide. Capacity: the price flow is trivial — forty transfers of a few megabytes spread over hours. The content flow's worst case — a from-scratch weekend where every branch pulls the full two gigabytes — is eighty gigabytes across three nights. That is well inside a modest data-center uplink once waves spread it. An ordinary weekend moves a fraction of that, because the mirror-style pulls skip unchanged files. The wave offsets are themselves the concurrency control. That is one more reason the schedule lives in a runbook table rather than in forty separately-remembered branch configurations. Doing this arithmetic before go-live is the difference between a wave plan and a wave guess.

The Branch Job

Every branch runs the same pull job, differing only in schedule and retry profile. It fires from the Windows scheduler (the setup pattern is in task scheduler for transfers). Estates that prefer configuration over scripting install the automation client at each branch instead. The sequence is identical either way:

BRANCH PULL SEQUENCE — nightly-prices
1. Connect to hub as branch-0NN; read prices/current.txt
2. Compare with local marker — already current? write receipt
   footer "verified, no change" and stop (idempotent re-run)
3. Download rel folder to local staging under temp names
4. Verify every MANIFEST line: present, size, hash
5. Atomic swap: staged release -> live folder; keep previous
   release locally as prices.prev for instant rollback
6. Write local marker: "on rel_YYYYMMDD, verified 02:41 local"
7. Upload receipt to receipts/branch-0NN/ on the hub
On any failure: retry with backoff within the window
(fragile class: resume partial downloads); when the retry
budget is spent, stop and let the hub-side check flag us.

Step two is what makes the whole estate self-healing. The job can run as often as you like — hourly all night, if the window needs it. That is because a branch that is already current does nothing but confirm. That idempotence is also the catch-up mechanism. A branch that was powered off simply runs the same sequence when it returns, finds itself behind, and fetches the current release. It never fetches the missed ones, because prices are latest-wins. Retry profiles follow link class, using the backoff shapes from retry strategies and backoff. Solid branches get three brisk attempts. Fragile branches get patient, spaced attempts with resume enabled, because their enemy is the mid-transfer drop, not the refusal.

The job's security posture is worth a sentence, because it is the blast-radius story for the whole estate. The branch account can read releases and write one receipts folder — nothing else. A fully compromised branch machine can therefore fetch content every branch already receives and forge its own receipts. It cannot touch another branch's evidence. Above all, it cannot modify a release, because no branch account anywhere has write access to the release area. The one write path into releases is the hub-side publish step, which answers only to the pricing landing folder. Distribution flows outward; trust never does.

The Wave Schedule

The price file is small, so waves here are not about hub bandwidth. They are about margin management and not letting forty connections land on the same minute. Slots derive from local quiet hours. Fragile and thin links go first in each group, giving their retries the longest runway. Jitter of a few minutes separates branches within a wave. The plan fits on one table, and a copy lives in the on-call runbook:

Wave Branches Link classes First pull (local) Re-runs until Margin at opening
East-1 5 (incl. 2 fragile) fragile, thin 02:00 06:30, hourly about five hours
East-2 8 solid 02:30 06:30, hourly about six hours
Central-1 6 (incl. 1 fragile) fragile, thin 02:00 06:30, hourly about five hours
Central-2 9 solid 02:30 06:30, hourly about six hours
West-1 4 (thin) thin 02:00 06:30, hourly about five hours
West-2 8 solid 02:30 06:30, hourly about six hours

Because slots are local, the groups' windows land an hour apart in hub time. The hub serves a rolling wave that follows the map westward, never the whole fleet at once. The hourly re-runs are the same idempotent job; on a good night every branch is done on its first pass and the re-runs are heartbeats.

The content set plays a different game with the same pieces. Two gigabytes over a thin link at roughly a gigabyte an hour is a two-hour clean transfer. That is comfortable inside one night, but with no appetite for drama, the design spends the whole weekend. Mirror-style pulls run Friday and Saturday nights in the same wave pattern. Only changed files move, which most weeks trims the set by half or more. Sunday night exists purely for stragglers and re-verification, and activation is a separate step. Branch software switches to the new set only when the local marker says complete and verified. Every branch reaches that state days before Monday. Delivery sweats all weekend precisely so that activation never sweats at all.

Verification and the Morning Report

Verification is the full stack from proving every site got the right version, installed twice. Manifests are hashed at publish, with branch-side verification before the swap and markers written after it. Receipts are uploaded to per-branch folders, and the server's database log corroborates every session. A report job on the hub runs at each group's window close for the price flow. It is an absence check comparing the group's roster against receipts. A fleet-wide report lands in the operations mailbox before the earliest opening, exceptions first. It lists any branch stale, failed, or absent, with its retry count, last evidence, and minutes of margin. Sunday evening, the content flow gets the same treatment ahead of Monday.

Alerting follows the margin ladder. Nothing pages while re-runs still have hours of runway. Window-close exceptions go to the operations channel. A branch still unverified when its margin drops under ninety minutes pages the on-call, naming the branch, the release, and the local opening time. The report generator itself is watched by a heartbeat, because a verification layer that dies silently is worse than none.

A typical bad-but-fine morning reads like this in the report. Thirty-eight branches are current on their first pass. Branch-031 is current on its third attempt just before four, after two mid-transfer drops. That is noted, not alarming: it is the third slow night this month, with a link review queued. Branch-023 is absent all window, flagged at close, and escalated when margin thinned past five. That is resolved as a site power cut, self-healed by the mid-morning re-run. Nobody guessed, nobody drove anywhere, and the whole story was written by the machinery before anyone asked. That is the texture the design is buying.

The Failure Playbook

A design is judged by its bad nights. Here are the ones this estate has planned for, and what happens in each:

  • Fragile link drops mid-download. The transfer resumes from its offset on the next attempt. Backoff spaces the attempts. The wave's early slot means even three drops leave hours of margin. Nobody is woken. The morning report shows a later verified-at and a higher attempt count — data for the link-class review, not an incident.
  • Branch powered off all night. The absence check flags branch-023 at window close; operations see it in the channel; the escalation fires only if margin runs low. When power returns at ten, the next scheduled run pulls the current release, verifies, swaps, and files its receipt. The report's history shows one absent night, self-healed, no human in the loop.
  • The pricing export is late. The publish monitor alerts at the expected-by time — aimed at the pricing team, not the branches. Branch jobs keep re-running harmlessly against the old current.txt. When the export lands at four, publish completes. The hourly re-runs then sweep the fleet current with the margin that early slots and hourly cadence bought. Past a defined point — margin under an hour for any group — it becomes a business incident with a business decision: open on old prices or delay.
  • A bad price file ships. Roll forward: the corrected export publishes as a new release, and current.txt flips. The hourly re-runs then replace the bad file fleet-wide within the hour. True rollback exists too — flip current.txt back, and branches also hold prices.prev locally for an on-the-spot reversal while the wire catches up.
  • Receipt stuck behind a dead link. The branch is verified and trading correctly but unconfirmed at the hub. The report says "unconfirmed — download logged, receipt pending." The branch job re-sends the receipt on its next connection, and nobody re-pushes anything at a healthy fleet.

Remember: every entry in this playbook was decided in an afternoon meeting, not at four in the morning. That is the entire value of a worked design — the bad nights are rehearsed while nobody is tired, and the on-call inherits decisions instead of dilemmas.

Why These Choices: The Reasoning Trail

  • Pull, not push — the fleet has forty branches with flaky uptime and consumer-grade links. Pull means one inbound door at the hub, one credential per branch, and offline branches that heal themselves. The full argument is in hub designs for branch distribution.
  • Explicit jobs, not replication — both flows have deadlines and need per-site proof; the price flow needs a fleet moment and the content flow needs release exactness. Continuous convergence offers neither, as replication vs scheduled transfer lays out.
  • Waves from local hours, fragile first — margin is the currency that buys quiet nights, and timezones hand it out free; the tactics come from distributing over thin and unreliable links.
  • Manifests, markers, receipts, report — because "the jobs ran" is not "the branches are current," and the operations manager's question deserves rows, not feelings — the whole evidence stack from the verification article.
  • Idempotent hourly re-runs — one mechanism serves as delivery, retry, catch-up, and heartbeat. Every special case the design does not need is a page of runbook nobody has to write.

Adapting It

The design flexes along the axes from the series opener. Ten branches: keep everything, shrink to two waves, and the report fits on a phone screen. With two hundred branches, waves get bigger and more numerous. The hub's uplink and session limits get real arithmetic, and a regional relay may earn its keep for a distant cluster. If branch machines are too locked-down to run jobs, flip to hub-side push tasks with the same release layout, verification, and report. The evidence design survives the direction change intact. Different payloads — configuration packs, per-site data feeds, signage media — slot into the same release-marker-receipt pattern with different sizes and windows. What should not flex: the immutable releases, the atomic markers, the site-side verification, and the exceptions-first report. Those four are the design; everything else is sizing.

Frequently Asked Questions

Why do the branches pull instead of the hub pushing?
This fleet has flaky links and machines that go offline. Pull gives each branch one outbound connection, one local credential, and automatic self-catch-up on return. The hub needs just one hardened inbound endpoint. Push becomes preferable mainly when branch machines cannot run their own jobs.
Why run the pull job hourly when the file only changes once a night?
Because the job is idempotent — a current branch checks the marker and stops. The hourly cadence then triples as retry, catch-up for returning branches, and a heartbeat, and it absorbs a late-publishing producer without any special handling.
How does the design roll back a bad price file?
Preferably forward: publish a corrected release and flip the current marker, and the hourly re-runs replace the bad file across the fleet within the hour. For instant local relief, each branch also keeps the previous release on disk and can swap back while the corrected version distributes.
Would this design work over FTP or FTPS instead of SFTP?
The architecture — releases, markers, manifests, receipts, waves — is protocol-independent, and the hub server speaks FTPS and HTTPS as well. SFTP was chosen here for its single-port firewall story across forty branch networks. Whatever protocol you pick, keep it encrypted and keep one protocol fleet-wide for sanity.
What would you build first if starting from nothing?
The release layout with manifests and markers, and one branch's pull-verify-swap-receipt job. That vertical slice proves the whole pattern; scaling it to forty branches is then configuration and scheduling, and the report script comes naturally once receipts exist to read.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.