Egress, Cost, and the Cloud Transfer Bill
The cloud transfer bill rarely surprises anyone because the rates were secret. It surprises because nobody counted the crossings. A flow designed on a LAN carries habits — re-fetch to verify, let every consumer grab its own copy, sync the whole folder nightly just in case. Those habits cost nothing when bytes moved across a switch. They cost real money when the same bytes cross a cloud boundary a dozen times a month.
This article, part of our Cloud and Hybrid Transfer Architecture series, treats cost the way an architect should. Cost is a structural property of the design, not a line item to grumble about later. You will not find a single price in this article, deliberately — exact rates change and differ between providers, while the shape of the metering is stable everywhere. Learn the shape, and you can estimate any flow's bill before you build it, on any platform, this year or ten years from now.
The Asymmetry in One Sentence
Here is the whole subject compressed: data flowing into a cloud platform is cheap and usually free; data flowing out is metered per unit moved. Storage itself is billed gently by the month; requests are billed in tiny amounts that matter only in bulk. But movement outward is where transfer architecture meets the finance department. That means movement to the internet, to your own premises, often even between the provider's own regions.
Notice what kind of change this is for a budget. On premises you paid for capacity: a pipe of a certain size, the same price whether it ran full or idle. Peak demand was the number that mattered. The cloud bills consumption: every byte out is counted, idle costs nothing, and the total moved per month is the number that matters. Neither model is cheaper in general — but designs tuned for one look wasteful under the other. The LAN-era habit of moving data freely because the pipe is already paid for is precisely the habit consumption billing punishes.
The design stance that follows is simple to state: treat the cloud boundary as a toll bridge. Crossing it is fine — that is what it is for — but every crossing should be deliberate, counted, and made once rather than habitually. Almost everything else in this article is that one stance applied to the places transfer flows actually cross.
Where the Meters Sit
Picture your cloud footprint as territory with toll booths at specific gates. The structural map, valid across major platforms:
- Out to the internet — the classic egress meter. Every byte served to a partner, a customer, or your own office over the public internet is counted. This is usually the biggest number on a transfer-heavy bill.
- Out to your premises — normally the same meter as the internet. Private connectivity options (a dedicated circuit or tunnel into the provider's network) are typically metered too, at different rates — private does not mean free.
- Between regions — replication or transfer from one of the provider's regions to another is metered in most configurations. A "backup to a second region" is a recurring crossing, not a one-time one.
- Between zones — some platforms meter traffic between the isolated zones inside one region; traffic within a single zone is typically free. This one varies most, so check yours.
- Requests — each put, get, and especially each listing call carries a per-operation charge. Individually microscopic; multiplied by a polling loop against a million-key bucket, real. The listing mechanics behind this are covered in object storage in your file flows.
- Cold-tier retrieval — archive storage tiers charge to bring data back: a retrieval fee, a wait, and on some tiers a minimum storage period billed even if you delete early.
And the directions that are typically free or nearly so: inbound from anywhere, and movement within a single zone. Good hybrid designs lean their heavy traffic onto exactly those lanes.
The diagram below marks the booths on one realistic flow — the same flow the next section walks in detail.
A Worked Walkthrough: What Crosses the Meter
Take that flow apart. Halvorsen Retail's forty branches each push one nightly file — call it a few hundred megabytes apiece — to a cloud landing zone. Head office pulls the files back for batch processing. Simple, sensible, and as first built, it crossed the meter far more than anyone intended. Three multipliers were hiding in the scripts:
- The sync-everything pull. The head-office job was written as "download everything under the branch prefix." Files accumulated in the zone for a month before cleanup, so every night's run re-fetched the whole month. By the last week of the month, each night's pull was nearly thirty times the size of the night's actual new data.
- The department copy. Merchandising liked the data too, so they ran their own copy of the pull script against the same prefix. Every byte now crossed twice.
- The Friday rebuild. A weekly reporting job re-downloaded the trailing month of history to rebuild its warehouse tables — a third recurring crossing for data the hub had already fetched once.
Count the crossings per byte and the shape of the bill explains itself. A byte uploaded on the first of the month could cross the meter dozens of times before it aged out. The redesign changed no protocols and no business behavior — it changed crossings. The pull job became incremental, tracking what it had already fetched and downloading each file exactly once. A scheduled tool such as Sysax FTP Automation keeps that state naturally when a job moves files through a "fetched" folder rather than re-listing the world. The fetched files landed on an internal distribution point — an on-prem Sysax Multi Server share in Halvorsen's case. There, merchandising and the reporting job now help themselves for free. And the zone's lifecycle rule deletes files a few days after confirmed pull, which shrank both storage and the temptation to re-sync history. After the change: each byte crosses the meter exactly once, and the request bill fell too, because one narrow nightly listing replaced three wide ones.
It is worth writing the before-and-after as plain arithmetic, because this is the calculation to repeat on your own flows. Before: one byte uploaded early in the month was pulled nightly by the sync-everything job for the rest of the month. It was pulled nightly again by merchandising's copy, and pulled four more times by Friday rebuilds. That is comfortably over fifty crossings for one byte, invisible in any single night's run. After: one crossing, ever. The files, the branches, the business value — identical. The bill — a fraction of what it was. Nothing about the fix required negotiating rates; it required counting.
Remember: the meter charges for movement, not for existence. Files sitting in the zone cost gently; files commuting daily cost like commuters. When a cloud transfer bill looks wrong, audit journeys, not storage.
The Patterns That Quietly Multiply the Bill
The Halvorsen story generalizes into a short rogues' gallery — worth scanning any design for before it ships:
- Verify-by-re-download. Pulling a file back to compare it doubles every transfer. Verify instead with the checksum the storage service recorded at upload. Or carry your own hash in object metadata and compare hashes, not bytes. That is the end-to-end thinking in verifying transfers end to end without the round trip.
- Per-consumer fan-out. Ten consumers each fetching a copy is ten crossings. Pull once to the side where the consumers live, then distribute locally. If the consumers are external and global, that is a distribution problem — see our multi-site distribution series. In that case, the crossing count per recipient is a real cost of the design.
- The anxious sync. Tools that re-copy whenever metadata looks unfamiliar — timestamps shifted, sizes recorded differently between filesystem and object store — can silently re-transfer entire datasets on a schedule. Test what your sync tool does on an unchanged file before trusting it across a meter.
- Wide polling. A job that lists an entire bucket every few minutes to find one new file pays in requests forever. List a narrow prefix, or better, let the platform's arrival events do the watching.
- The round-trip backup. Backing cloud data up by pulling it on-prem, then later restoring it back up, crosses the meter both ways. Sometimes that is exactly right — an off-platform copy has real independence — but it should be a decision, not an accident of reusing the old backup script.
- Serving downloads straight from the bucket. Handing recipients presigned-style links is a clean pattern, and every download meters. Fine for modest volumes; for a large audience fetching large files, count recipients times size times months before falling in love.
Placement: The Cheapest Byte Is the One That Doesn't Cross
Zoom out and the multipliers share one root: data resting far from the things that read it. The strategic fixes are placement fixes, the same ones argued shape by shape in hybrid topologies:
- Process where the data is; move the results. If the raw feed lives in the cloud and the report is a hundredth its size, run the crunch in the cloud and pull the report. If processing must stay on-prem, pull the raw data once and keep it.
- Keep the authoritative copy on the reader-heavy side. A dataset read hourly on-prem and archived in the cloud should live on-prem and push copies up — the cheap direction — not live in the cloud and be fetched hourly.
- Choose regions to shorten the expensive legs. Locality helps latency and throughput as well as cost; the physics and the finance point the same way more often than not.
Placement is also where cost arguments must know their place. Data-residency obligations can force an authoritative copy to a particular side of a border regardless of what the meter prefers — data localization requirements covers when that override applies. Cost is an input to architecture, never the trump card.
The Two Meters: Time and Money Run Together
Cost is not the only thing that accumulates per byte — time does too, and the two meters usually point at the same design. The time-to-transfer arithmetic is unchanged from any other planning exercise: volume divided by usable throughput. Usable throughput on long paths is capped by latency as described in bandwidth-delay product. A design that ships the same bytes repeatedly is slow for the same reason it is expensive. A placement that shortens the frequent legs saves hours for the same reason it saves money.
The place this bites hardest is the big one-time move: repatriating an archive, switching providers, pulling a decade of accumulated files home. Both meters run in full — every byte pays, and every byte takes wall-clock time your cutover window must contain. Estimate both before promising dates. Estimate the money from the crossing volume, and the time from the volume divided by what the path really sustains. Measure with a test transfer; never trust the label on the pipe. When the numbers come out in weeks rather than days, you are in the territory our moving massive datasets series covers — parallelism, seeding strategies, and the honest crossover point where shipping drives beats sending bits.
An Egress-Aware Design Checklist
Run any cloud-involved flow through this before committing. It fits on one screen and catches the expensive habits while they are still free to fix:
EGRESS-AWARE DESIGN CHECKLIST Crossings [ ] Every metered crossing in this flow is drawn on the diagram [ ] Each byte crosses outward at most once per consumer side [ ] No verify-by-re-download: hashes compared, not bytes [ ] No consumer fetches its own copy of shared data [ ] History rebuilds read a local copy, not the cloud Chatter [ ] Listings are narrow (prefix-scoped), not bucket-wide [ ] Polling replaced by arrival events where volume justifies it [ ] Sync tools tested against unchanged files (no phantom re-copies) Placement [ ] Heavy traffic runs in free lanes (inbound, within-zone) [ ] Processing sits beside the data it reads most [ ] Authoritative copy lives on the reader-heavy side [ ] Region chosen to shorten the expensive, frequent legs Lifecycle [ ] Zone/relay storage expires automatically after confirmed delivery [ ] Archive tiers hold only what live flows never touch [ ] Restore-from-archive path known, priced, and tested Governance [ ] Monthly crossing volume estimated and written down [ ] First real invoice compared against the estimate [ ] Someone owns noticing when the two diverge
Estimating a Flow's Bill Before You Commit
You can price a design in an hour with nothing but arithmetic, and the discipline matters more than the precision. The method:
- Draw the flow and mark every crossing from the meter map above — internet-out, premises-out, region-to-region, zone-to-zone, retrievals.
- Put a monthly volume on each crossing: file size, times frequency, times number of consumers who cross there. Your transfer logs are the honest source for sizes and frequencies. A server that logs every session, as Multi Server does to file and database, already holds the per-account byte counts you need.
- Count the chatty operations: listings per day times pages per listing, puts and gets per file. Small numbers; multiply them anyway.
- Apply your provider's current rate sheet. This is the only step where real numbers enter, and the only step to redo when providers change rates or you change providers.
- Stress-test with a bad month: a full re-run after a failure, one partner resending everything, the history rebuild someone will inevitably request. If the bad-month number frightens nobody, the design is robust.
- Reconcile against the first live invoice. The gap between estimate and actual is your list of uncounted crossings — close it while the flow is young.
When the Meter Should Not Decide
Cost-awareness has a failure mode: letting pennies veto sound engineering. Never drop integrity verification, encryption, or logging to shave a crossing — the checklist above removes redundant journeys, not safeguards. Pay the meter without complaint for the off-platform backup copy whose whole value is being somewhere else. Pay it for the second-region copy when continuity genuinely requires it. Pay it for the compliance export a regulator is entitled to. And remember that egress is the price of freedom as well as of operation. The fact that you can pull every byte back out — at a known, finite cost — is what keeps a cloud commitment reversible. An estate that avoids ever crossing the meter by leaving everything in one platform forever has not saved money; it has bought lock-in on installments.
The Habit to Keep
Data in is cheap; data out is metered; requests and retrievals nibble at the edges. Design so each byte crosses the boundary deliberately and once. Pull once and share locally, verify with hashes, and list narrowly. Expire what has been delivered, and put processing next to its data. Estimate the crossings before you build, and reconcile against the first invoice after. The companion pieces to read next include hybrid topologies, where placement options get argued properly. Also read hybrid reference architectures, where these cost arguments appear inside three complete worked designs.
Frequently Asked Questions
Why doesn't this article give actual prices?
Is uploading to the cloud really free?
Does a private connection to the provider avoid egress charges?
How do I verify a transfer without downloading the file again?
Our cloud bill jumped and nobody knows why. Where do I look first?
Why did restoring files from the archive tier cost so much?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
