Distributing Over Thin and Unreliable Site Links
Every multi-site design works beautifully between the hub and the well-connected pilot branch down the road. Then it meets the real estate. There is the branch on shop-grade broadband shared with the card terminals, and the site whose router reboots itself most nights. There is the depot at the end of a rural line that moves data at a tenth of the speed the provider's paperwork promised. Distribution designs are not judged by their best links. They are judged by whether the worst link in the fleet still gets its files, night after night, without a human coaxing it.
This article is the survival kit for that worst link. You will learn to inventory what your links can actually do. You will run the window arithmetic that says whether a payload can fit a night at all. You will learn to shrink what you send with delta thinking, and survive mid-transfer drops with resume and retry. You will learn to spend timezones as a scheduling resource. You will monitor each site's delivery as its own promise rather than as a line in a fleet average. This article is part of our Multi-Site Distribution series, and it assumes the hub-and-branch vocabulary of the earlier articles.
Know Your Links Before You Plan Around Them
Thin-link design starts with an unglamorous act: writing down what each site's connection really is. Not the advertised speed — the observed one. Four facts per site are enough, and a test transfer at a representative hour collects most of them:
- Usable throughput, measured by transferring a realistically sized file during the window you intend to use. The number that matters is what a transfer actually achieves at two in the morning, which routinely differs from the daytime figure in either direction.
- Sharing — what else uses the link, and when. A branch link that carries payment traffic and cloud backups has hours you must avoid regardless of raw speed.
- Stability pattern — does the link drop, and when? Many consumer-grade connections renegotiate nightly at a provider-chosen hour; some sites lose power to cleaning crews and timers. The pattern is knowable if anyone looks.
- Local quiet hours — when the site's own business least needs the link, stated in the site's local time.
With those facts, sort the fleet into a handful of link classes instead of pretending forty sites are forty special cases. Three classes cover most estates:
| Link class | Typical profile | Design consequences |
|---|---|---|
| Solid | Business-grade line, stable, generous headroom | Any slot works; schedule these last, they will catch up regardless |
| Thin | Slow but steady; shared with business traffic by day | Off-hours only; deltas and compression matter; long windows; early slots |
| Fragile | Slow and drops mid-transfer; nightly renegotiations; outages | Everything Thin needs, plus resume support, patient retries, and alerting tuned for lateness rather than panic |
Classes keep the design honest and the configuration manageable. Schedules, retry policies, and alert thresholds get set per class, with genuine one-off exceptions documented by name. When a site changes provider, it changes class, and everything downstream follows.
The Window Arithmetic
Every thin-link plan stands or falls on one line of arithmetic: transfer time equals payload size divided by usable throughput — plus margin. Do it in round numbers, per link class, before promising anything. If the weekly content set is around two gigabytes and a thin-class link moves roughly a gigabyte an hour at night, that is about two hours of clean transfer. In that case, add half again for retries and slow patches and you need a three-hour window. If the site's local quiet hours run from closing until early morning, the window exists. If the payload were ten times larger, it would not, and no scheduling cleverness will manufacture the missing hours.
When the arithmetic fails, you have exactly three honest levers, in order of preference:
- Shrink the payload — deltas, compression, and splitting content by urgency (the next section).
- Stretch the window — start earlier, finish later, or spread a big release across several nights, delivering ahead of the day it must take effect. This works because delivery and activation are separable. Bytes can arrive Tuesday through Thursday and switch on Friday. That distinction is introduced in the problem-shapes article. The distinction becomes pure gold on thin links.
- Change the transport — for the rare payload so large no window fits, stop pretending the link is the tool. Techniques for genuinely massive movements, including shipping storage instead of bits, live in our moving massive datasets series.
Remember: a payload that does not fit the window on the thinnest link is a design error, not an operations problem. It will not fix itself with retries — it will fail politely every night until someone does the arithmetic that should have been done first.
Send Less: Delta Thinking
The cheapest byte to move over a thin link is the byte you never send. Delta thinking is a hierarchy, and the biggest wins sit at the top where the machinery is simplest:
First, skip unchanged files. A mirror-style task compares the release area against what the site already holds — by name, size, and timestamp — and transfers only what differs. For a weekly content set where most images carry over from last week, this alone routinely cuts the transfer by more than half. It is also the easiest tactic to adopt, since mirror and sync tasks are standard equipment in automation tools. In Sysax FTP Automation they are generated by a wizard rather than scripted. That makes rolling the pattern across a fleet of branch jobs a configuration exercise instead of a programming project.
Second, compress what you do send. Text-heavy payloads — price files, catalogs, configuration — often shrink to a fraction of their size. Compress once at the hub, not forty times at the source, and weigh the site-side unpack step into the design. The tradeoffs, including the files that refuse to shrink, are covered in compression in transfer pipelines.
Third, bundle crowds of small files. Thousands of tiny files pay per-file overhead — a round trip or several each — that can dwarf the actual data on a high-latency link. Packing them into a single archive turns a chattering hour into a quiet few minutes. If your content set is image-heavy, this is frequently the difference between fitting the window and not.
Last, block-level deltas for large files that change slightly. When a single big file is mostly the same as yesterday's — a database export, a large catalog — rsync-style tools transfer only the changed blocks. The algorithm and its WAN behavior are their own subject, covered in rsync over WAN links. The design point here is to reach for block-level deltas only after the simpler tiers. This approach earns its complexity only when large-and-slightly-changed is genuinely your shape.
Survive the Drop: Resume and Retry
Thin links are usually also flaky links, and a drop mid-transfer is where naive designs quietly starve. Picture a fragile-class site whose router renegotiates every ninety minutes, receiving a set that takes two hours. A transfer that always restarts from byte zero will fail at the renegotiation, restart, fail again — forever, while consuming the link all night. The fix is resume: continuing an interrupted transfer from where it stopped rather than from the beginning. The common protocols support restarting from an offset, with mechanics and caveats detailed in protocol-level partial transfers. The design rule is simply that fragile-class sites must use a client and server pair that actually exercise resume. Your biggest single files deserve a deliberate test: kill the transfer partway, watch it resume, confirm the result verifies intact.
Resume handles the interrupted transfer; retry handles the failed attempt. Give every site job a retry policy with backoff — wait briefly, try again, wait longer, escalate. The approach is described in retry strategies and backoff. Give the site job a retry budget bounded by the window. After the third failure or an hour before deadline, whichever comes first, stop burning the link and raise a flag instead. Automation tools carry this logic natively. Retry counts, delays, and error handling are exactly the sort of thing you configure once per link class in a tool like Sysax FTP Automation. You also configure an email notification to fire when the budget is exhausted. That way, a bad night announces itself instead of surfacing at opening time.
One safety rule ties the section together: a site must never activate a file it has not finished receiving. Downloads land under a temporary name and are renamed into place only after size and checksum agree. That is the atomic-swap pattern from temp names and atomic renames. With that in place, drops and retries can mangle the night's schedule all they like; they can never put half a price file in front of the tills.
Stagger: Spending Time Instead of Bandwidth
Bandwidth on thin links is fixed; time is negotiable. Staggering — assigning different sites different slots — converts hub capacity you do not have into hours you do. Estates that span timezones get a gift with it. Express every site's window in local terms ("from an hour after closing until two hours before opening"). Translate to hub time, and the fleet spreads itself. The easternmost sites' quiet hours begin while the westernmost are still trading, so their transfers finish before the western windows even open. The hub's night becomes a rolling wave that follows the map, rather than one crushing spike at midnight.
Within each timezone group, apply the wave discipline from hub designs for branch distribution. Put fragile links in the earliest slots. They need the most runway for retries, and their stragglers overlap harmlessly with later waves. Put solid links last. Add jitter of a few minutes between starts so no wave lands as a synchronized herd. And schedule the whole plan against the deadline, not against midnight. The question each slot answers is "how much margin does this site have if tonight goes badly?" A fragile site whose window closes ten minutes before its deadline has no plan; it has a coin flip.
The Deliberately Slow Design
Some payloads deserve to be slow. When the weekly content set strains the thinnest quarter of the fleet, the mature move is often to stop trying to deliver it in one night at all. Spread delivery across three: a third of the set — or a third of the fleet — each night. Activate on the morning after the last. Every night's transfer shrinks to a comfortable fit, retries have room to breathe, and the deadline stops being a cliff. The cost is bookkeeping: sites now hold releases in transit. So the release layout must mark completeness explicitly with a manifest and a final marker file. That is exactly the machinery that the verification article builds, so nothing activates a set that is only two-thirds arrived.
The same logic scales down gracefully. An urgent two-megabyte price file and a lazy two-gigabyte media set should not share a job, a window, or a retry policy just because they share a destination. Split flows by urgency: the small urgent file gets a tight window, aggressive retries, and loud alerts. The big patient set gets long windows, gentle retries, and lateness tolerance measured in hours. Most "our distribution is unreliable" complaints turn out to be one flow's requirements strangling another's.
One more structural option belongs in the kit, though it should be reached for last: the regional relay. Consider a cluster of fragile sites sharing a region, especially where the long-haul leg to that region is itself the bottleneck. In that case, a small relay server in the region can receive each release once from the hub and serve the local sites from nearby. The long thin hop is paid once instead of a dozen times, and local pulls run fast and retry cheaply. The honest cost is a second tier of everything. There is another server to patch, and another set of logs to fold into verification. A new question ("did the relay get it?") stands in front of the old one. At forty branches a relay is usually more design than the problem deserves. At a few hundred, or with an overseas cluster on one congested route, it starts paying rent.
Monitor Each Site as Its Own Promise
Fleet averages lie. "Thirty-nine of forty delivered" is a healthy average and a broken promise if the fortieth is the flagship store. On thin links, lateness is routine and failure is occasional. Monitoring must treat every site as its own contract, with three distinct states worth distinguishing:
- Late but moving — the transfer is running slower than usual, margin shrinking. On a known-thin site this is a dashboard color, not a page. Nobody should be woken because a thin link is being thin within its budgeted window.
- Failing — attempts are erroring, retry budget draining. Alert when the budget exhausts or when remaining time versus remaining data says the deadline is mathematically lost. An end-of-window freshness check catches this even when the failure produced no error at all.
- Absent — the site never connected. In pull designs this is silence, detectable only by comparing who arrived against who should have. In push designs it shows as connection failures. Either way, absence after the window closes is an incident, not a curiosity.
The raw data for all three states already exists if the hub logs properly. A hub running Sysax Multi Server writes per-account activity to file or database — which branch connected, when, what it took, how long it took. Per-site transfer durations trended week over week are your earliest warning of a link quietly degrading. The site that took forty minutes in spring and takes ninety now will be next month's incident unless someone renegotiates its class. Turning those logs into the per-site morning report, and deciding who gets alerted for which state, is the subject of proving every site got the right version.
Gotcha: tune alerts per link class, or your fragile sites will train the team to ignore alarms. A page that fires every time a thin link is slow is a page nobody reads by week three. Then it is missed on the night it finally matters. Alert on eroded margin and exhausted budgets, never on ordinary slowness.
The Thin-Link Survival Checklist
Run any distribution flow against this list before trusting it to the fleet's worst link:
THIN-LINK READINESS — per flow, per link class
[ ] Usable night-time throughput measured, not assumed
[ ] Window arithmetic done: payload / throughput + half margin fits
[ ] Mirror/delta enabled: unchanged files are never re-sent
[ ] Compression applied where the payload shrinks meaningfully
[ ] Small-file crowds bundled into archives
[ ] Resume verified by killing a real transfer partway
[ ] Retry policy with backoff and a budget tied to the window
[ ] Site-side temp-name download and atomic swap in place
[ ] Slots staggered: fragile first, jitter within waves,
windows derived from LOCAL quiet hours
[ ] Big patient payloads split from small urgent ones
[ ] Per-site states defined: late / failing / absent
[ ] Alerts tuned per link class; absence check after window
Twelve lines, and every one of them is cheaper than the morning call from a branch that opened without its files.
Where to Go Next
Thin links bend a distribution design without breaking it, provided the design spends its three currencies deliberately. Those are bytes (send less), attempts (resume and retry within a budget), and hours (stagger against local windows, deliver ahead of activation). The evidence layer that proves the bent design still kept its promises is next in proving every site got the right version. The worked forty-branch design shows every tactic from this article deployed in one realistic estate, thin links, timezones, and all.
Frequently Asked Questions
How do I find out what a branch link can really do?
What is transfer resume and when do I need it?
What should I do when the file set is too big for the overnight window?
How do timezones help with distribution scheduling?
Should a slow site trigger an alert?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
