Home › Topics › Large Files › Compression

Compression for Large Transfers: When It Pays

"Just zip it first" is the oldest advice in file transfer. For a large file it is sometimes brilliant and sometimes a waste of an hour. A 40 GB database export might shrink to 8 GB and cross a slow link in a fifth of the time. A 40 GB video file will shrink by nothing at all. The sender will have spent twenty minutes of CPU and 40 GB of scratch disk proving it. The difference is not luck; it is arithmetic you can do in advance, plus a thirty-second test.

This article gives you both. It explains in plain words why some files compress and others cannot. It shows how to test a sample before committing to the whole file. It works through the CPU-versus-bandwidth calculation on a realistic example and compares compressing on the fly with building an archive first. By the end you will be able to look at a large file and a link and make the call. You will be able to say, with numbers, whether compression will save time or burn it. It is part of our Large File Strategies series.

What Compression Can and Cannot Do

Compression is the process of rewriting data in a shorter form that can be expanded back to the exact original. It works by finding redundancy — patterns that repeat, values that are predictable from what came before — and describing them once instead of many times. A log file that contains the string ERROR: connection refused forty thousand times does not need forty thousand copies. It needs one copy and a very short way of saying "that again."

The limit on compression is entropy, which in plain words is the amount of genuine surprise in the data. Text, database dumps, XML, source code, and uncompressed images are low-entropy: full of structure and repetition. A good compressor shrinks them to a fifth or a tenth of their size. Random-looking data is high-entropy: every byte is a surprise, nothing repeats, and no compressor on earth can shrink it. You cannot summarize a list of dice rolls.

The compression ratio is the original size divided by the compressed size. A ratio of 4 means the compressed file is a quarter of the original. A ratio of 1.02 means it shrank by two percent, which for transfer purposes is the same as not shrinking at all. Every decision in this article comes down to one question: what ratio will this file achieve, and is that enough to pay for the time spent achieving it?

The Files That Will Not Shrink

Here is the fact that saves the most wasted hours: a great many large files are already compressed, and compressing them again does nothing. Their formats include compression internally, so the bytes on disk already look random. The compressor reads all 40 GB, finds no repetition, and writes out 40 GB plus a few bytes of headers — sometimes slightly larger than the original.

Kind of file Typical examples Typical ratio Compress before sending?
Text and structured data logs, CSV, SQL dumps, XML, JSON 4 to 10 Almost always
Disk and VM images, raw backups raw disk images, uncompressed backup sets 1.5 to 3 Usually, but test
Media JPEG, PNG, MP3, MP4, MKV, most audio and video 1.0 to 1.02 No
Archives and packages ZIP, 7z, gz, installers, office documents (which are ZIP inside) 1.0 to 1.05 No
Encrypted files PGP files, encrypted containers and backups 1.0 exactly Never — compress before encrypting
Compressed database backups backups made with the database's own compression option 1.0 to 1.1 No

Two rows deserve emphasis. Encrypted data is, by design, indistinguishable from random noise. If it compressed, it would still contain visible structure, which is exactly what encryption exists to hide. So a PGP-encrypted file will never shrink, and the correct order is always compress first, then encrypt. Most PGP tools compress automatically before encrypting for this reason; the encrypt-before-send workflows article covers the pipeline. And office documents, despite looking like "text," are ZIP archives internally, so a folder of spreadsheets compresses far less than its contents would suggest.

Remember: the extension tells you a lot but not everything. A file called backup.bak might be raw (compresses well) or already compressed by the backup tool (does not). When the table says "test," test — the sample check below takes thirty seconds and settles it.

Test a Sample Before You Commit

You do not need to compress 40 GB to learn whether 40 GB will compress. Compress the first few hundred megabytes and look at the ratio. Most large files are uniform enough that a sample from the start predicts the whole. If a file might change character partway (a dump with a text schema followed by binary blobs), take a second sample from the middle.

On Linux and other Unix-like systems, read the first 256 MB with head -c, push it through the compressor, and count the output bytes with wc -c:

$ head -c 268435456 export.sql | gzip -6 | wc -c
41283910

$ head -c 268435456 vm-image.raw | gzip -6 | wc -c
171559366

$ head -c 268435456 training-video.mp4 | gzip -6 | wc -c
268401187

Reading the output: 268,435,456 bytes went in each time. The SQL export came out at 41 MB — a ratio of about 6.5, an excellent candidate. The disk image came out at 172 MB, a ratio of about 1.6, worthwhile on a slow link and marginal on a fast one. The video came out at 268.4 MB: a ratio of 1.0001, meaning the compressor found nothing, and sending it compressed would be pure waste. If you have zstd installed, substitute zstd -3 for gzip -6; it runs several times faster and the ratios are similar.

On Windows, PowerShell can carve out the same sample with a .NET file stream and compress it with the built-in Compress-Archive:

$src = 'D:\export\export.sql'
$in  = [System.IO.File]::OpenRead($src)
$buf = New-Object byte[] (256MB)
$n   = $in.Read($buf, 0, $buf.Length)
$in.Close()
$out = [System.IO.File]::Create('D:\export\sample.bin')
$out.Write($buf, 0, $n)
$out.Close()

Compress-Archive -Path 'D:\export\sample.bin' -DestinationPath 'D:\export\sample.zip' -Force
$ratio = (Get-Item 'D:\export\sample.bin').Length / (Get-Item 'D:\export\sample.zip').Length
'{0:N2}' -f $ratio

6.41

The script reads up to 256 MB into a buffer (PowerShell understands 256MB as a number of bytes). It writes the buffer out as sample.bin, zips it, and prints the ratio. Under about 1.1 means "do not bother"; over 2 is worth the trouble on most links; in between, the arithmetic in the next section decides. Delete the sample files afterward.

The CPU-Versus-Bandwidth Arithmetic

Compression trades processor time for bytes on the wire. Whether that trade wins depends on three numbers. The first is the link speed in bytes per second. The second is the compressor speed — how many bytes of input it can eat per second on your hardware. The third is the ratio from the sample test. For our running example — a 40 GB file on a 200 megabit link, about 22 MB per second in practice — the uncompressed transfer takes about thirty minutes. There are two ways to use a compressor, with different arithmetic.

Streaming: compress and send at the same time

In streaming compression the compressor's output is fed straight into the network connection, so compressing and sending overlap. The transfer runs at the speed of whichever is slower: the compressor consuming input, or the link carrying the compressed output. In plain words, streaming compression pays whenever the compressor can eat input faster than the link could have sent it raw, provided the ratio is meaningfully above 1. If the compressor is slower than the link, you have swapped a network bottleneck for a CPU bottleneck.

Archive first, then send

Building an archive means making a compressed file on disk before the transfer starts. The two steps happen one after the other, so their times add. The rule: with a ratio of r, the compressor must be faster than the link by a factor of r divided by (r minus 1). At a ratio of 3 it must be one and a half times the link speed. At 1.5 it must be three times. At 6 it only needs to be 1.2 times.

The table works the example through with realistic single-core compressor speeds. Exact figures vary with hardware; the relationships do not.

Compressor (typical speed) Ratio on the SQL export Streaming, 200 Mbit link Archive first, 200 Mbit link Streaming, 10 Gbit LAN
None 1 30 min 30 min about 32 s
zstd level 3 (about 400 MB/s) 6 5 min about 7 min about 100 s (slower)
gzip level 6 (about 40 MB/s) 6.5 about 17 min about 21 min about 17 min (much slower)
xz level 6 (about 8 MB/s) 8 about 83 min (slower) about 87 min (slower) about 83 min (far slower)

Read across the rows. The fast compressor wins on the slow link: forty gigabytes of input at 400 MB per second takes 100 seconds. But the compressed 6.7 GB still needs five minutes on the wire. So the link remains the bottleneck and the transfer runs six times faster than raw. The slower compressor still wins — seventeen minutes beats thirty — but now the CPU is the bottleneck and the link sits idle half the time. The very slow, very thorough compressor loses badly: an eight-to-one ratio is worthless if producing it takes three times longer than sending the file raw.

Now read the last column. On a fast local network the raw transfer takes half a minute, and every compressor loses. None can process input at anything like 1.25 GB per second on one core. This is the most common compression mistake: a script written for a slow WAN link, copied to a fast LAN. There it now triples the transfer time while pegging a CPU. Compression is a remedy for a slow link, and a slow link only.

Rule of thumb: compression pays when the compressor is faster than the link and the sample ratio is above about 1.5. On links below a few hundred megabits, a fast compressor at a low level nearly always wins on compressible data. On gigabit and faster links, compress only when the ratio is very high, or not at all. Higher compression levels rarely pay for transfer: they cost several times the CPU for a few percent better ratio.

Two refinements. Most compressors have a multi-threaded mode — zstd -T0 uses every core, and 7z multi-threads by default. That mode multiplies the compressor speed by the core count and moves the break-even point toward faster links. And the far end has to decompress, but decompression is typically several times faster than compression, so it rarely changes the answer.

Streaming Compression Versus Archives

The arithmetic shows streaming is faster when both are possible. But the two approaches differ in more than speed, and the differences decide which one fits a given transfer.

Streaming in practice

The classic streaming pipeline compresses on the sender, pipes the output through an SSH connection, and decompresses on the receiver. No compressed file is ever written to disk on either side:

# Send one large file, compressing in flight; nothing extra touches disk
$ zstd -3 -T0 -c export.sql | ssh backup@10.20.30.40 'zstd -d -c > /srv/inbound/export.sql.part'

# Send a whole directory tree the same way, bundled with tar
$ tar -cf - /srv/export | zstd -3 -T0 | ssh backup@10.20.30.40 'zstd -d | tar -xf - -C /srv/inbound'

The -c flag tells zstd to write to standard output instead of a file; -d means decompress. The receiver writes to a .part name so that anything watching the inbound folder ignores the file until it is complete and renamed. This is the pattern from temp names and atomic renames. Streaming needs no scratch space and starts immediately, but it has one large weakness: a stream cannot be resumed. If the connection drops at 90 percent, there is no compressed file on either side to pick up from; you start over. For a file that takes minutes, that is acceptable. For one that takes hours on a flaky path, it is not.

Several protocols offer streaming compression as a built-in option, and they inherit the same arithmetic. SSH's -C flag (and SFTP over it) compresses the whole session with a moderate, single-threaded algorithm. That is helpful on slow links with compressible data, harmful on fast ones. rsync -z does the same and sensibly skips files whose extensions mark them as already compressed. Some FTP servers support a compressed transfer mode called MODE Z; support is uneven, so test before relying on it. In every case: if the link is fast or the data is already compressed, turn the option off.

Archives in practice

An archive is a compressed file on disk, made before the transfer with a tool such as 7-Zip, gzip, or zstd. It is sent as an ordinary file and unpacked after arrival. It costs scratch space on both ends and the serial time from the table. In exchange you get everything an ordinary file gets: the transfer can resume after an interruption. You can hash the archive and verify it on the far side, and you can split it into chunks. The archive format carries its own checksum so a corrupted archive announces itself on unpacking. For anything measured in hours, those properties outweigh the extra minutes:

# Linux: archive with zstd (fast, multi-threaded), keeping the original
$ zstd -3 -T0 export.sql -o export.sql.zst
export.sql : 15.38%   (40.0 GiB => 6.15 GiB, export.sql.zst)

# Windows: archive with 7-Zip at a fast level; -mx1 is fastest, -mx5 is default
C:\> 7z a -mx1 D:\export\export.7z D:\export\export.sql

The zstd output line reports the result directly: the archive is 15.38 percent of the original, a ratio of about 6.5. With 7-Zip, compare the two file sizes afterward. Low levels (-mx1, zstd -1 to -3, gzip -1) are the right choice for transfer: nearly the ratio of the slow levels at several times the speed.

The choice, then: stream when the transfer is short enough that a restart is cheap and scratch space is tight. Archive when the transfer is long, the path is unreliable, or you need the file to be resumable, hashable, or splittable.

Compression Inside an Automated Flow

In a scheduled job, compression is a pre-processing step on the sender and a post-processing step on the receiver. It needs the same care as any other step. Three things go wrong repeatedly.

  • Scratch space. The archive step doubles the sender's disk use for the duration, and the unpack step doubles the receiver's. Jobs that ran for months fail the first time the export grows past half the free space. Check space before compressing, and delete the archive after a verified transfer.
  • Naming and detection. The receiver must recognize the archive, unpack it, verify the result, and only then hand the file to whatever consumes it. A watcher that passes export.7z to a loader expecting export.sql produces a confusing failure. Put the unpack step in the flow explicitly.
  • Compressing what is already compressed. A job that zips everything in the outbound folder spends its longest minutes on the media files that gain nothing. Branch on the sample ratio or the extension, and skip those.

Automation tools generally support this directly. Sysax FTP Automation, for instance, can zip files as a pre-processing step before a scheduled upload and unzip after a download. So the compress-transfer-unpack sequence lives in one job definition. The broader design of these pipelines is covered in compression in transfer pipelines.

A Decision Procedure

Everything above reduces to five steps:

  1. Check the format. Media, archives, encrypted files, compressed backups: send uncompressed. Text, dumps, raw images: continue.
  2. Test a sample. First 256 MB through a fast compressor. Ratio under 1.1: send uncompressed. Over 1.5: continue. In between: continue only on a slow link.
  3. Compare speeds. Compressor input speed versus link speed in bytes per second. Compressor slower: send uncompressed, or use a faster compressor or more threads. Compressor faster: continue.
  4. Pick streaming or archive. A few minutes of transfer with tight scratch space: stream. Anything long or over an unreliable path: archive, so you can resume and verify.
  5. Use a low level. The fast settings give most of the ratio for a fraction of the CPU. Reserve high levels for long-term storage, not files transferred once.

The Version to Keep in Your Head

Compression is a trade of CPU for bandwidth, and it only wins when the CPU has something to find and bandwidth is the scarcer resource. Test a sample, because the extension is a hint, not a promise. Compare the compressor's speed with the link's, because a compressor slower than the link makes the transfer slower. That is why compression helps on a WAN and hurts on a LAN. Stream when the transfer is short and disk is tight; archive when the transfer is long and you need resume and verification. And never compress after encrypting.

Compression changes what you send; the other articles in this series change how you send it. What changes when files get large explains why long transfers need resume in the first place. Our article on choosing protocols and settings for multi-gigabyte moves covers where protocol-level compression options live and when to switch them off. For the slow-link side of the equation — measuring what your link actually delivers — start with the diagnosing slow transfers series.

Frequently Asked Questions

Why did my zipped video end up slightly larger than the original?
Video formats are already compressed internally, so the data looks random to a compressor. It finds nothing to shrink and adds its own headers on top, producing a file a few bytes to a few kilobytes larger. Send media files uncompressed.
Should I use the maximum compression level for a transfer?
Almost never. High levels cost several times the CPU for a few percent better ratio. For a one-time transfer that extra CPU time usually exceeds the seconds saved on the wire. Use a low or default level for transfers; reserve high levels for long-term storage.
Does SFTP or FTPS compress data automatically?
Not by default. SSH (and therefore SFTP) can compress the session if both sides enable it. rsync has a compression flag, and some FTP servers support a compressed mode. But none are on unless you turn them on. On fast links or already-compressed data, leave them off.
Should I compress before or after PGP encryption?
Before, always. Encrypted data is indistinguishable from random noise and cannot be compressed. Most PGP tools compress automatically as part of encryption, so check whether yours does before adding a separate step.
How can I tell whether a file will compress without compressing all of it?
Compress the first 256 MB and compute the ratio, as shown in this article. Large files are usually uniform enough that a sample predicts the whole. If the file might change character partway through, take a second sample from the middle.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.