Compression in Transfer Pipelines: When and How
Ask why a transfer job zips its files and you usually get one of two answers: "we've always zipped it" or "we never got around to it." Both are the same mistake in opposite directions. Compression added out of habit burns CPU and adds a failure-prone stage to flows where it saves nothing. Compression never considered leaves four-hour transfer windows that a ten-minute zip step would cut to twenty minutes. Neither flow ever measured anything.
Compression — encoding a file so the same information takes fewer bytes — is a tradeoff. It is a very lopsided one depending on your data and your link. It can shrink a CSV feed to a twentieth of its size, or it can chew through CPU to make a video file slightly larger. This article gives you the reasoning to know which outcome you will get before you build anything.
We will cover what compression actually does (one honest paragraph of theory that predicts everything else). We will cover a decision table for when it pays, and the common formats — zip, gzip, tar — in plain words. Then we will cover the compress-then-encrypt ordering rule and why it is not optional. We will cover how to verify archives before and after they travel, and the decompress-on-arrival stage on the receiving side. This article is part of our pre- and post-processing series. In the pipeline skeleton from the anatomy of a transfer pipeline, compression lives in the transform stage, as packaging.
What Compression Actually Does
All general-purpose compression exploits one thing: redundancy. A compressor scans for patterns — repeated strings, frequently occurring symbols — and replaces them with shorter references. For example: "the next 40 characters are the same as the 40 you saw 200 bytes ago." The more repetitive the data, the more there is to replace, and the smaller the output.
That single sentence predicts everything you need to know about what compresses:
- Text formats compress dramatically. CSV files repeat the same delimiters, the same store IDs, the same date prefixes on every row. Log files repeat timestamps and message templates. XML repeats every tag name twice per element. Shrinking such files to 10–20 percent of their original size is routine, and highly repetitive feeds do even better.
- Already-compressed data does not compress. JPEG images, video and audio files, and existing zip or gzip archives have already had their redundancy removed — that is what their formats do. Running a compressor over them wastes CPU to achieve roughly nothing. The output can even be slightly larger, because the archive format adds its own bookkeeping around data it could not shrink. Modern office document formats are in this category too: they are zip containers internally, so zipping them again achieves very little.
- Encrypted data does not compress at all. Good encryption produces output that is statistically indistinguishable from random noise, and random noise has no patterns to exploit. Hold that thought — it becomes the ordering rule later.
What It Buys, What It Costs
Compression buys three things. First, wire time: fewer bytes over a slow or busy link means a shorter transfer. It also means faster retries when a transfer must restart, and less exposure to mid-transfer failures. Second, storage: archives of sent files, retention copies, and the partner's inbox all shrink. Third — often the biggest win in practice — bundling: an archive turns ten thousand small files into one transfer. Per-file overhead (a listing entry, an open, a close, sometimes a round trip each) dominates transfers of many small files on any protocol. One archive replaces all of it with a single sustained stream, a benefit you get even if the contents barely shrink. A folder of fifty thousand tiny report files that takes hours to transfer file-by-file routinely moves in minutes as a single archive. The bytes were never the problem; the per-file ceremony was.
It costs three things. CPU and time: compressing and decompressing are real work. On a fast LAN the compression step can take longer than the transfer it was meant to shorten. A new failure mode: the archive itself can be truncated or corrupt, so the pipeline must test it — a whole section below. And opacity: nothing can look inside an archive without unpacking it. Your validation checks must run before packaging (as laid out in validating files before they leave). Boundary tools like malware scanners must either unpack archives or wave them through unexamined — a limitation covered honestly in scanning limits and encrypted files.
When Compression Pays: The Decision Table
Two questions decide almost every case: how compressible is the data, and which resource is scarce — bandwidth or CPU? The table gives the honest defaults:
| Situation | Compress? | Why |
|---|---|---|
| Large text, CSV, XML, or log files over a WAN or the internet | Yes | Five- to twenty-fold shrink is typical; wire time dominates the job, so the CPU spent repays itself many times over |
| Thousands of small files, on any link | Yes — archive them | Per-file overhead dominates; one archive replaces ten thousand round trips. Bundling is the win even before any shrink |
| Media files, existing archives, modern office documents | No | Already compressed internally; you spend CPU to gain roughly nothing, sometimes a slight growth. Archive without compression if you only need bundling |
| Modest files on a fast internal network | Usually no | The transfer already takes seconds; compression adds a stage that can fail and saves time nobody was losing |
| Slow, congested, or metered link, compressible data | Yes | Every byte costs time or money; CPU is cheap by comparison. This is where compression rescues transfer windows |
| Files that will be PGP-encrypted before sending | Before encryption only | Encrypted output will not compress; see the ordering rule below |
| The partner's contract specifies a format | Follow the contract | Interoperability beats optimization; their intake automation expects exactly what was agreed |
And then measure, once, honestly: take a real production file, time compress + transfer + decompress against plain transfer, end to end. The table predicts the outcome; a ten-minute measurement on your data and your link proves it, and settles the debate permanently. Record the numbers next to the flow's documentation. After all, "we zip this feed and it saves 40 minutes a night" is an answer; "we've always zipped it" is not.
One dial worth knowing about while you measure: most compressors offer a compression level — a tradeoff between speed and ratio. The honest news is that the default level is almost always the right answer for pipelines. Cranking the level to maximum typically buys a few extra percent of shrink for a large multiple of CPU time. That matters in a nightly window. Dropping to the fastest level is worth testing only when the compression step itself is what blows the schedule.
The Formats in Plain Words: zip, gzip, tar
zip is an archive and a compressor in one: it bundles many files and folders into a single file and compresses each member individually. It is the lingua franca — natively handled on Windows desktops and available everywhere else. That makes it the polite default for partner exchange. Its per-member structure means one file can be listed or extracted without unpacking the rest.
gzip compresses exactly one stream: sales.csv in, sales.csv.gz out. It does no bundling at all — hand it a folder and it has nothing to say. It is the standard companion for single large files on Unix-like systems, and many tools read gzip streams directly without a separate extraction step.
tar is the mirror image: it bundles many files, folder structure and all, into one stream — and compresses nothing. The classic pairing is tar plus gzip (logs.tar.gz): bundle first, then compress the bundle as one continuous stream. That ordering gives so-called solid compression — patterns repeated across many similar small files get exploited, which per-member zip cannot do. So a folder of near-identical logs often ends up noticeably smaller as .tar.gz than as .zip. The tradeoff: extracting one member means reading through the stream. If a human at the destination must double-click the file, zip remains the kind choice.
The everyday commands, including the verification flags the next section relies on:
# zip: bundle and compress a staging folder into one outbound archive zip -r outbound/sales_YYYYMMDD.zip staged/ # test the archive's integrity without extracting anything unzip -t outbound/sales_YYYYMMDD.zip # gzip one large file (produces sales_YYYYMMDD.csv.gz) gzip sales_YYYYMMDD.csv gzip -t sales_YYYYMMDD.csv.gz # verify it # tar + gzip a folder tree, then verify by reading it end to end tar -czf outbound/logs_YYYYMMDD.tar.gz logs/ tar -tzf outbound/logs_YYYYMMDD.tar.gz > /dev/null
Compress, Then Encrypt: The Ordering Rule
Many pipelines both compress and encrypt, and the order is not a style preference — it decides whether the compression does anything at all. Encryption's entire job is to make output indistinguishable from random noise; random noise has no patterns; no patterns means nothing for a compressor to find. So:
- Compress, then encrypt: the compressor sees the repetitive plaintext and shrinks it fully; encryption then wraps the small result. You keep the entire benefit.
- Encrypt, then compress: the compressor sees noise, finds nothing, and emits output as large as its input — while both steps report success. The pipeline looks healthy and the benefit is silently zero.
One honest wrinkle: OpenPGP tools typically apply their own compression internally, just before encrypting. That is why a .pgp of a CSV file usually comes out smaller than the original even with no explicit zip step. So for a single file headed into PGP, an extra gzip pass is mostly redundant. Where an explicit compression step still earns its place is bundling: zip the day's many files into one archive, then encrypt that archive. The mechanics of the encrypt step, and the folder-staged pipeline around it, are covered in encrypt-before-send workflows and how PGP file encryption works — link, learn, and do not re-invent that stage.
Remember: compress before you encrypt. Encrypted bytes look random, and random does not compress — reverse the order and both steps still "succeed" while the entire benefit quietly evaporates.
Verifying Archives Before and After Transfer
An archive adds a new way to fail: the archive itself can be truncated or corrupt while every surrounding step reports success. A zip interrupted mid-write is a valid-looking file that dies at extraction time — on the partner's side, at the worst possible moment. So a pipeline that packages must also test, at three points:
- After creating the archive, run the format's own integrity test —
unzip -t,gzip -t,tar -tzfas shown above. These read every compressed block and verify its internal checksums. Only after the test passes may the pipeline consider deleting the loose source files. The cautious keep sources until the transfer itself is verified too. - After the transfer, verify the bytes arrived intact: compare sizes and checksums between the local archive and the remote copy, exactly as for any file. The techniques are in verifying transfers end to end. A checksum sidecar sent along with the archive lets the partner verify independently.
- On arrival, before use, the receiving side runs the integrity test again, before anything downstream touches the contents. A failed test on arrival is treated as a failed transfer: quarantine, alert, re-request.
Write the evidence down as you go: the run log should record each archive's name, member count, compressed and uncompressed sizes, and checksum. When a partner reports a bad extraction three days later, those four numbers tell you immediately whether the archive left your side healthy. That is the difference between a five-minute answer and an afternoon of guessing.
And because an archive is just a file, all the partial-file rules apply to it with extra force. Build it in a work folder and move it into the outbound folder only when complete. That way, no watcher ever picks up a half-written zip. The full set of those arrival-and-handoff defenses is in the partial-file safety series.
On the sending side, this packaging sequence is standard enough that automation tools generate it whole. A task built with the wizard in Sysax FTP Automation can zip the outbound files and apply OpenPGP encryption in the correct compress-then-encrypt order. It can transfer the result, archive what was sent using its file operations, and email the outcome. The script editor is available when your flow needs an extra step the wizard did not anticipate.
Decompress-on-Arrival: The Receiving Half
Whoever receives the archive runs the mirror stage. The safe sequence: verify the bytes (size, checksum), test the archive, and extract into a work area. Validate the extracted contents, and only then move them where the consuming system looks. Extracting directly into the consumer's folder is the receiving-side version of the half-written-file bug. A half-extracted tree looks exactly like a complete one to whatever reads it.
Four cautions belong in every decompress step, all boring until the day they are not:
- Disk space. Compression ratios cut both ways: a 200 MB archive of CSV files may need several gigabytes to extract. Check free space against the archive's declared uncompressed size before extracting, not during.
- Path escapes. Archive entries can carry absolute paths or
..components. A naive extractor will happily write outside the target folder — a classic attack on automated intake. Extract with tools that sanitize paths, into a dedicated folder, under an account that has no business writing anywhere else. - Expansion abuse. A tiny hostile archive can be built to expand into something enormous. For intake that accepts archives from outside, cap the extracted size and fail loudly when the cap is hit.
- Collision policy. Decide in advance what happens when an extracted name already exists — overwrite, version, or reject. "Whatever the tool does by default" is not a policy.
If arrivals land on your own transfer server, the decompress step can start itself. For example, in Sysax Multi Server, event triggers (Pro and Enterprise editions) can run your unpack-and-validate script the moment an upload completes. The server's activity log — to file or database — gives you the independent record of exactly what arrived and when. After extraction, the contents usually flow onward to the routing stage, which is the next article's territory.
The Version to Tell a Colleague
Compression is a measured tradeoff, not a habit. Text compresses enormously; media, archives, and encrypted data do not compress at all. It pays on slow links, big text, and piles of small files (where bundling, not shrinking, is the real win). It wastes CPU on fast LANs and pre-compressed formats. If you also encrypt, compress first — always — because encrypted bytes will not shrink. And an archive is a new thing that can break. So test it after creation, verify it after transfer, and test it again before anything downstream trusts its contents.
In this series, the packaging step you just built sits inside the transform stage of the pipeline anatomy. The checks that must run before the box is sealed are in validating files before they leave. The content-level reshaping that often shares the stage with compression is covered in transforming files between systems.
Frequently Asked Questions
Why did my file get bigger after I zipped it?
Does the compress-then-encrypt order really matter that much?
Should I still zip files if I send them over SFTP?
What is the practical difference between zip and tar.gz?
How do I check that an archive is not corrupt without extracting it?
unzip -t for zip, gzip -t for gzip, and listing a tar.gz end to end with tar -tzf. Run the test after creating an archive and again on arrival, and treat any failure as a failed transfer.From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
