Home › Topics › Benchmarking › Tuning Knobs

The Tuning Knobs Worth Benchmarking

"I turned on compression, set it to eight connections, bumped the buffer, and switched the cipher. It's slower now. Which one do I put back?" All of them, and then we start again. Every tool exposes a dozen settings. Every forum thread recommends a different three. An evening of changing them all at once ends with a job that is faster for reasons nobody can name. Or it is slower for reasons nobody can undo. Settings applied without a method are not reverted. They accumulate, and within a year the server has a configuration no one dares touch.

This article is the method. It starts with the control run and the one-change rule. Then it walks through the knobs genuinely worth testing: parallelism, windows and buffers, compression, cipher choice, and client-side versus server-side. For each, it covers the effect to expect, how to test it, and when to skip it. It closes with a knob-by-knob table and a short list of settings that are not worth your evening. It is part of our Benchmarking Transfer Performance series. It assumes the harness and the five-runs rule from designing a transfer benchmark.

The physics of why some knobs work lives in our acceleration series, linked where it applies. Here the subject is the experiment.

The Control Run and the One-Change Rule

A control run is the benchmark of the job exactly as it is today. It uses the current client, the current settings, the real file mix, and the real link. It has five repeats, median and spread. It is the number every experiment is compared against. Without it a tuning session cannot tell an improvement from noise. If the control's spread is wide (runs varying by twenty percent or more), the environment is not stable enough to detect a tuning effect. In that case, the first job is to find out why, not to start turning knobs.

The one-change rule is exactly what it says. Change one setting and run five times. Compare the median with the control's median against both spreads. Write the verdict down, and only then change the next thing. Two changes at once produce a result that explains neither. Worse, two changes can cancel. Parallelism helps, a buffer change hurts, and the median lands where it started. Both are wrongly recorded as useless. A change that helped is kept and becomes the new control. A change that did not is reverted, and the fact is written down so nobody tries it again.

Acme learned the rule the slow way. Their nightly export went from forty minutes to fifty-five after an evening in which four settings changed together. Nobody could say which to put back. Reverting all four restored the forty minutes. They tested one at a time the following week. Four parallel connections helped, two settings did nothing, and compression was the culprit. It had been switched on for a folder of already-zipped archives. The evening had cost a week.

Each experiment gets a short record; the "Next" line is what keeps a session from wandering:

EXPERIMENT   parallel-uploads-mixC
Question     Do 4 parallel connections shorten the partner feed (mix C, WAN)?
Control      1 connection; median 14 min 10 s; spread 13:40 to 15:05 (5 runs)
Change       client profile: max concurrent transfers 1 -> 4 (nothing else changed)
Result       median 4 min 05 s; spread 3:50 to 4:30 (5 runs)
Server       CPU 12% -> 30%; disk queue steady; no errors in the activity log
Verdict      KEEP: gain far larger than either spread; no saturation signs
Next         try 8 only if the feed is still outside its window; else stop

Remember: a control run, one change, five repeats, a written verdict, then the next change. A knob whose effect is smaller than the control's spread has no measurable effect, whatever the forum thread promised.

Knob 1: Parallelism

Parallelism means running more than one transfer at the same time. That means several files on separate connections, or in some tools several streams carrying pieces of one large file. It is the knob with the largest effect and the clearest limit.

Expected effect. Large for many small files, because each connection spends most of its time waiting on per-file round trips and a second connection fills that idle time. Large on a WAN for the same reason, and because a single connection often cannot keep a long link full. Nothing for one large file on a LAN that is already near wire speed, because there is no idle time to fill. Past a certain point, the effect is negative. Once the disk queue lengthens or a core pins, more connections compete for the saturated resource and the total slows.

How to test. Run the control at one connection, then two, four, and eight, five repeats each. Watch and record CPU and disk queue on both ends during every run. The curve almost always rises steeply, flattens, and then dips. Keep the last setting before the flattening, not the peak of a lucky run. Test with the real mix. The parallelism that suits five thousand small files rarely suits ten large ones.

When to skip. A single large file on a LAN already running near the network's ceiling. A server whose disk or CPU is already busy in the control run. Skip it for a partner whose server limits concurrent connections per account. Check before you test, because the eighth connection may be refused rather than slow (connection limits and per-user caps explains why). Our article on parallel streams explains why streams and connections behave differently.

Knob 2: Windows, Buffers, and Request Sizes

Three related settings live here, in different places. The TCP receive window is an operating-system setting that governs how much data a connection can have in flight before it must wait for an acknowledgment. Modern operating systems tune it automatically (receive window autotuning). The only tuning most administrators ever need is to confirm nobody has turned it off. Application buffers are the chunk sizes a client reads and writes with. Request sizes and outstanding requests are SFTP-specific. They govern how large each read or write request is, and how many the client keeps in flight before waiting for replies.

Expected effect. On a LAN with sub-millisecond round trips, close to nothing for all three. On a WAN, potentially large for the window and moderate for the request settings, because a long link multiplies every pause. The arithmetic (how much data a link of a given speed and delay needs in flight to stay full) is the bandwidth-delay product. A window smaller than that product caps a single connection no matter what else you tune.

How to test. Measure the round-trip time first; if it is under a millisecond, skip this knob entirely. Otherwise, confirm autotuning is enabled on both ends, then run the control for one large file across the WAN. If a single connection cannot get near the link's capacity while CPU and disk are idle, the window is the first suspect. Our guide to TCP tuning first has the fix. For SFTP, then test the client's request-size and outstanding-request options, one at a time, on the large-file mix. For other clients, test the buffer option if there is one. For big single files, protocol and settings for big moves collects what matters.

When to skip. Any LAN. Any WAN where the control already fills the link. Any client that does not expose the setting. Registry-level TCP tweaks are on the not-worth-it list below, where they earned their place.

Knob 3: Compression On or Off

Compression shrinks data before it is sent and expands it on arrival, trading CPU on both ends for fewer bytes on the wire. Some protocols offer it inside the connection; some pipelines compress files into archives before transfer. The benchmark question is the same for both.

Expected effect. Large, sometimes several times faster, for compressible data such as CSV exports, logs, XML, and plain text, on a link that is the bottleneck. Zero or negative for data that is already compressed or encrypted: archives, media, PGP-encrypted files, most office documents. On those, the CPU works hard to save nothing, and the transfer gets slower. On a LAN where the link is not the bottleneck, compression usually costs more CPU time than it saves in wire time even for text.

How to test. First check compressibility cheaply: put a sample of the real files into an archive and compare sizes. A file that shrinks by less than a tenth is not worth compressing. Then run the control and the compressed cell on the real files and record CPU on both ends. Never use random test data for this. It does not compress and makes every compression setting look useless. Where the pipeline can compress once before transfer instead of on the fly, test that as a separate cell. It is often the better shape (see compression in transfer pipelines).

When to skip. Already-compressed or encrypted data. A LAN with a fast link and a busy CPU. A partner feed where the receiving side cannot handle archives.

Knob 4: Cipher Choice

The cipher is the algorithm that encrypts the connection, negotiated from lists on both ends. It is the knob people most enjoy testing and the one that most rarely changes anything.

Expected effect. Honestly small on modern hardware. Processors accelerate the common ciphers. The difference between two policy-approved choices is typically a few percent of one core, invisible in wall-clock time on a gigabit link. The exceptions are real but narrow. They include a cipher the hardware does not accelerate or an old or very small CPU. They also include a ten-gigabit link outrunning one core, or a server encrypting for many clients at once.

How to test. Look at CPU during the control run first. If no core is near full use, the cipher is not your bottleneck and the test will confirm it. If one is, choose two ciphers your policy already allows. Force each in the client and run the large-file mix on the LAN, five repeats each, recording CPU. Most clients print the negotiated cipher with a verbose flag. Log it so the experiment can be repeated.

When to skip. Whenever the CPU is idle during transfers, which is most of the time. And never test a cipher your policy forbids. A weak cipher that wins a benchmark has traded a security property for a rounding error. The policy exists for reasons our cipher policy basics guide explains.

Knob 5: Client-Side or Server-Side?

The same improvement can often be made on either end, and where you make it matters as much as what you change. A client-side change affects one job and is reverted by editing one profile. A server-side change affects every user and partner at once. A benchmark run at midnight cannot see the customer who connects at nine. Prefer client-side changes when the effect is equivalent. Make server-side changes in a maintenance window with a written rollback. Then rerun the control for a job you did not intend to affect. (There is always one.)

Before choosing either, find the bottleneck, because a knob that does not touch it does nothing. The basic monitors on each platform show four patterns during a control run. The Windows tools are the resource and performance monitors. The Linux tools are the top, iostat, and sar family. The patterns are:

  • A core pinned on either end: encryption, compression, or per-request processing is the limit. More connections spread work across cores; compression off or a hardware-accelerated cipher lightens each one.
  • Disk queue long or disk at full use: storage is the limit. Parallelism will make it worse. Move the landing folder to a faster or less busy disk; test with a plain local copy first.
  • Network card at line rate: the link is full. Nothing helps except fewer bytes (compression, if the data allows) or a bigger link. The article on measuring bandwidth impact shows which job is filling it.
  • Everything idle, transfer still slow: waiting, not working. Per-file round trips or a starved window. Parallelism, pipelining, keep-alive, and session reuse are the knobs that fill idle time.

Server-side, the transfer server's activity log is part of the instrument. Sysax Multi Server records every login and transfer with timestamps, to a file or a database. So you can see whether a small-file job's time went into connecting, authenticating, or moving data. That is the difference between tuning the login path and tuning the disk. Client-side, a scheduled job is the ideal vehicle. A task in Sysax FTP Automation with one profile setting changed runs the same transfer set for five nights. Its email notification confirms each run, so nobody needs to be awake for it.

The Knob Table

The table gathers the knobs above with the ones that hide in plain sight (scanning, disk placement, name lookups, logging). Those often dwarf protocol settings and are found only by benchmarking the real job on the real server.

Knob Expected effect How to test When to skip
Concurrent transfers Large for small files and WAN; harmful past disk or CPU saturation 1, 2, 4, 8 on the real mix; watch CPU and disk queue on both ends One large file on a LAN near wire speed
Parallel streams for one file Helps one large file on a long link; needs support on both ends 1 vs 4 streams, large-file mix, WAN LAN, or a link that is already full
TCP window autotuning Large on WAN if it was disabled; nothing on LAN Confirm enabled both ends; large file over WAN before and after RTT under a millisecond
SFTP request size and outstanding requests Moderate on WAN; small on LAN Client options, one at a time, large-file mix over WAN Client lacks the option; LAN
Compression Large on text over a slow link; negative on compressed data Real files, on and off; record CPU; archive-once as a separate cell Compressed, encrypted, or media data; fast LAN
Cipher choice Small on modern CPUs; visible without hardware acceleration Two policy-allowed ciphers, large file on LAN, watch CPU CPU idle during transfers
TLS session resumption, HTTP keep-alive Large for small files over FTPS or HTTPS Small-file mix with and without; count handshakes in the log Large-file jobs
Antivirus on-access scanning of the landing folder Can dominate small-file jobs Lab only, with security's agreement: on-access vs scheduled scan Never skip the test; skip the change if policy forbids it
Landing folder disk placement Moderate; large if shared with the OS or logs Local copy test on the disk; move folder; rerun the batch mix Disk already faster than the link
Reverse DNS and similar lookups on connect Seconds per connection; severe for connect-per-file patterns Time the login alone; check the server's connect settings Login already under a second
Logging verbosity Small to moderate on small-file jobs Debug vs normal level, small-file mix Already at normal level
Jumbo frames Small; LAN only; every device on the path must agree Only with ten-gigabit links and full control of the path Almost always

Two rows deserve emphasis because they are so often the real answer. On-access antivirus scanning of a landing folder can add tens of milliseconds to every file, minutes on a five-thousand-file job. The fix is a conversation with the security team about scheduled scanning or scanning elsewhere in the flow (never a quiet exclusion). The article on scanning integration points describes the options. And a server that performs a reverse name lookup on every connection can add seconds per login when the lookup times out. That is invisible in a throughput benchmark and glaring in a job that connects once per file.

The Knobs Not Worth Your Evening

Some settings are better left alone. That is not because they never matter. Their expected effect is smaller than any realistic benchmark's spread, or their fragility outweighs any gain.

  • Registry-level or kernel-level TCP tweaks on a modern operating system. Autotuning already handles the window; hand-set values are usually worse and are forgotten within a year. Confirm autotuning is on and stop.
  • Cipher shopping when the CPU is idle. The knob does nothing, and the temptation to reach for a weaker cipher is the only outcome.
  • Jumbo frames on a mixed LAN. One device that does not agree turns a small gain into fragmentation or silent packet loss.
  • Changing three things at once because each seemed harmless. Three harmless changes are one unexplainable result.
  • Tuning for the benchmark rather than the job. A setting that wins on a single 2 gigabyte test file and loses on the nightly mix of two hundred medium files has optimized the wrong thing. The control run is always the real job.
  • Tuning a link that is already full. When the network card sits at line rate, every knob on this page is a placebo. Our article on whether you need acceleration is the honest guide to what the remaining options are.

I have turned most of the knobs on that list at one time or another. I turned the registry one twice, because the first time I had not written it down. A written verdict is the only thing that stops a knob being rediscovered every year by someone new.

Rule of thumb: find the bottleneck, then turn only the knobs that touch it. Use parallelism for waiting and compression for compressible bytes on a slow link. Use window checks for long links and disk placement for busy storage. Leave the cipher alone unless a core is pinned.

One Knob at a Time, Forever

Tuning is a habit, not a phase. Every meaningful change (a new client build, a moved folder, a security baseline applied by another team) is a knob turned. That is true whether or not anyone called it that. The control-run discipline is how you notice. The experiment records from this article are the raw material for the baseline and regression alarm in reporting results and catching regressions. And the next time someone asks which of four settings to put back, the answer is in a file, not a memory.

If the knob you are reaching for is the protocol itself, comparing protocols fairly applies the same parity rules to that larger change. And if the numbers improved but the users did not notice, measuring what users feel explains which knobs they can perceive. Those are usually the login and listing times no throughput benchmark measures.

Frequently Asked Questions

Which knob should I try first?
The one that touches your bottleneck. If the transfer is slow while CPU, disk, and network all sit idle, it is waiting on round trips. In that case, parallelism comes first. If the disk is busy, move the landing folder. If the link is full, only compression of compressible data or a bigger link will help.
How many parallel connections is too many?
The number after which the disk queue grows or a CPU core pins on either end. Test 1, 2, 4, and 8 with five runs each and keep the last setting before the gains flatten. Check too whether the partner's server limits connections per account.
Does turning compression on always help?
No. It helps compressible data such as text and CSV on a link that is the bottleneck. It hurts already-compressed or encrypted data, where the CPU works to save nothing. Archive a sample of real files first. If it shrinks by less than a tenth, leave compression off.
Why must I change only one thing at a time?
Two changes produce one result that explains neither. Two changes can cancel so that both are wrongly recorded as useless. A control run, one change, five repeats, and a written verdict is slower per experiment and far faster per answer.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.