Home › Topics › Capacity & Concurrency › Tuning Checklist

A Capacity Tuning Checklist for Transfer Servers

Capacity problems on transfer servers are rarely caused by one big mistake. They are caused by a dozen small defaults nobody changed. Those include a file handle limit of 1024 on a server that holds four hundred sessions, and a listen backlog sized for a workstation. They include logs on the same disk as the data, an idle timeout of an hour, and no alert on the disk queue. Each is a five-minute fix. The trouble is remembering all of them, on every server, and checking again when the load has doubled since the build.

This article is that memory. It collects the settings and checks from the rest of our Capacity & Concurrency series into one checklist. The checklist walks the server from the bottom up: operating-system limits, the network stack, disk and storage, and the transfer server's own settings. It also covers the monitoring hooks that tell you when any of it drifts. It closes with the quarterly review that keeps capacity ahead of demand instead of one month-end behind it. Every item names the command that checks it and the value that is usually right. So a junior administrator can work through it without a senior one in the room.

How to Use the Checklist

Run it three times in a server's life: once at build, once after the first real peak, and then every quarter. Three rules make it useful rather than dangerous. Record the current value before changing anything — a text file of "was / now / why / date" per server is the whole change log. Change one setting at a time when the server is already in trouble, so you know which change helped. And retest after material changes using the ramp from load testing a transfer server safely, because a tuning change that was never measured is a hope.

The values suggested below assume the example server used throughout this series. It has 8 cores, 16 GB of memory, a gigabit interface, and around 300 partners with a month-end peak of about 310 concurrent sessions. Scale them to your own sizing worksheet from sizing a transfer server.

Section 1: Operating-System Limits

The operating system has ceilings of its own beneath the transfer server, and they fail with messages that never mention capacity.

  • File handles per process. Every session, open file, and log costs at least one. On Linux the default soft limit is often 1024. Check with ulimit -n as the service account (and ulimit -Hn for the hard limit). For a service started by systemd, the value that counts is LimitNOFILE= in the unit file. A sensible setting is 65536. The symptom when it is wrong is "too many open files" in the server log while disk space is fine. Windows processes have a far higher ceiling, but track \Process(name)\Handle Count for a steady climb that indicates a leak.
  • System-wide file handles. sysctl fs.file-max is the machine's total; cat /proc/sys/fs/file-nr shows allocated and maximum. Rarely the limit on a dedicated server, but check once.
  • Processes and threads. A process-per-connection server needs a process limit above its session limit: ulimit -u, or TasksMax= in the systemd unit. Set it to at least twice the global session limit.
  • Swap. A transfer server should have swap configured as insurance and never use it. On Linux set vm.swappiness=10 so the kernel prefers dropping cache to swapping the server. On Windows leave the page file system-managed and alert on \Memory\Available MBytes. Any sustained swap activity means the memory row of the worksheet was wrong.
  • Power management. Virtual and physical servers alike can be throttled by a balanced power plan. Windows: powercfg /getactivescheme should report High performance. Linux: check the CPU frequency governor is not set to power saving on a server that encrypts.
$ ulimit -n; ulimit -Hn
1024
524288
$ sysctl fs.file-max vm.swappiness
fs.file-max = 9223372036854775807
vm.swappiness = 60
$ cat /proc/sys/fs/file-nr
5344	0	9223372036854775807

That output is a server that needs two changes: the soft handle limit is the 1024 default, and swappiness is at the desktop value of 60. The OS-level hardening article covers the security side of these same files, and the two reviews are worth doing together.

Section 2: The Network Stack

These settings live below the transfer server and above the network card. Most have safe defaults on modern systems, but each one has a failure mode that looks like something else.

  • Listen backlog. The queue of completed connections waiting for the server to accept them. On Linux, sysctl net.core.somaxconn should be at least 4096 (older kernels default to 128); the server's own listen setting caps it too. Check the live value with ss -ltn: Send-Q is the backlog size for each listening port. A Recv-Q that approaches it during a burst means connections are timing out at the door.
  • Ephemeral port range. Only matters when the server makes outbound connections — active-mode FTP data connections, or a gateway forwarding to a back-end. Linux: sysctl net.ipv4.ip_local_port_range, normally 32768 60999. Windows: netsh int ipv4 show dynamicport tcp, normally starting at 49152 with 16384 ports. Widen only if TIME_WAIT counts approach the range size.
  • TIME_WAIT count. ss -s reports it; netstat -an | find /c "TIME_WAIT" on Windows. On an FTP server, a TIME_WAIT count that dwarfs established connections means data connections are being recycled faster than ports free up. In that case, widen the passive range rather than shortening kernel timers.
  • Socket buffer limits. Distant partners need large TCP windows to fill the pipe. Linux autotunes within net.ipv4.tcp_rmem and net.ipv4.tcp_wmem; raise the third (maximum) value to 16 MB if partners are more than about 50 ms away. Windows: netsh int tcp show global should show receive window autotuning at normal. Our article on TCP tuning first covers the arithmetic.
  • Interface speed and duplex. A gigabit card negotiated down to 100 Mbit/s half duplex is a surprisingly common finding. Linux: ethtool eth0 | grep -E 'Speed|Duplex'. Windows: Get-NetAdapter | Select-Object Name, LinkSpeed.
  • Connection tracking. A host firewall tracks every connection in a table with a maximum. Linux: sysctl net.netfilter.nf_conntrack_max and net.netfilter.nf_conntrack_count; when count reaches max, new connections are silently dropped. Set the maximum to several times the peak connection count, remembering that each FTP transfer is a connection of its own. The same applies to the perimeter firewall, which our firewalls and NAT series covers.

Section 3: Disk and Storage

Storage is where transfer servers actually run out of capacity, so this section deserves the most time.

  • Separate volumes. The operating system, the server's logs, and the transfer data must be on different disks or virtual disks. Add a fourth for a logging database if you use one. A log volume that fills stops the server; a log stream sharing spindles with three hundred uploads steals IOPS from all of them.
  • Storage type matches the IOPS estimate. Take the disk row of the sizing worksheet — concurrent transfers times operations per stream — and compare it with what the volume can do. Spinning disks deliver one to two hundred random IOPS each; if the estimate is in the hundreds or thousands, the data volume belongs on solid-state storage. Confirm during a real peak with iostat -x 5 (aqu-sz under 2 per spindle, w_await under 20 ms spinning or 5 ms solid-state). Or use the PhysicalDisk counters Avg. Disk Queue Length and Avg. Disk sec/Write.
  • Virtual disk quotas. On virtual infrastructure, find the IOPS and throughput quota attached to the data disk. It is frequently the real ceiling, far below the hardware beneath it.
  • Last-access timestamps. By default some filesystems write a timestamp every time a file is read, turning every download into a write. Linux: mount the data volume with noatime (or confirm relatime). Windows: fsutil behavior query disablelastaccess; disable updates on the data volume if they are enabled.
  • Directory sizes. Listing a directory of 50,000 files costs CPU and IOPS every time a partner's client asks, and scripted clients ask often. Keep active directories under a few thousand entries with archive subfolders; the many-small-files problem article explains why.
  • Scanning and indexing. Real-time antivirus scanning of the data volume doubles the I/O of every upload. Scan at a defined point in the flow instead, as our article on scanning integration points describes, and exclude the data volume from any search indexer.
  • Free space thresholds. Alert at 20 percent free on the data volume and 30 percent on the log volume. A burst needs room for a full night's arrivals plus temporary files. Growth and cleanup are the subject of our storage growth and quotas and cleanup series.

Remember: when the checklist finds the data volume on spinning disks shared with the logs, stop and fix that before tuning anything else. No kernel setting compensates for a saturated disk queue.

Section 4: Transfer Server Settings

The server's own configuration is where the earlier articles in this series land. Every value here should be written down next to the reason for it.

  • Global session limit. Set, not unlimited: about 70 percent of the resource ceiling and at least 25 percent above the measured peak. Example: ceiling 600, peak 311, limit 400.
  • Per-user cap. Twice a normal client's parallelism, minimum 2; higher for named batch accounts. Example: 4.
  • Per-IP cap. Per-user cap times accounts expected behind one address, plus 2; a short exception list for NAT-heavy partners. Example: 6.
  • Login rate limit. Enough unauthenticated connections to absorb a burst of logins without pinning the CPU; on OpenSSH, MaxStartups 20:30:60 for an eight-core server. Pair it with brute-force auto-blocking, which our lockout and throttling design article covers. A server such as Sysax Multi Server provides IP allow and block lists with automatic blocking. That keeps scanners from spending your login budget.
  • Passive port range (FTP). Twice the expected simultaneous transfers, opened on the firewall, with the correct external address announced. Example: 200 ports. The passive range guide has the procedure.
  • Idle timeout. Long enough for a legitimate pause between commands, short enough to reap ghosts: 10 minutes for partners, shorter during bursts. Our timeouts and keepalives series covers the interaction with slow clients.
  • Transfer timeout. A stalled transfer should be dropped after a few minutes of no data, not left holding a slot and a passive port for an hour.
  • Ciphers and compression. Prefer a hardware-accelerated AES mode on CPUs that support it; disable SSH-level compression, which costs a core per session for almost nothing on already-compressed files. Our cipher policy basics article has the policy side.
  • Logging level. Debug-level logging under load is a write per command per session and can halve throughput. Run at the normal level and enable debug per session or per account when diagnosing.
  • Service recovery. On Windows, the service's recovery options should restart it on failure, with a short delay. That way, a crash at three in the morning does not wait for a human.

The details of choosing each number, and what the server says when a limit is hit, are in connection limits and per-user caps.

Section 5: Monitoring Hooks

Every item above can drift: a partner grows, a patch resets a default, a volume fills. The monitoring hooks are how you find out before the partners do. The server health monitoring series covers the tooling. This section is the list of what to collect and the thresholds that turn a number into a page.

Metric Source Warn Critical
Concurrent sessions ss -s / \TCPv4\Connections Established 70% of global limit 90% of global limit
Refusals per hour Server log (421, MaxStartups, refusal events) Any outside a known burst More than 5% of logins
Disk queue length iostat -x aqu-sz / Avg. Disk Queue Length 2 per spindle for 5 min 4 per spindle for 5 min
Write latency w_await / Avg. Disk sec/Write 20 ms 100 ms
CPU busy vmstat / \Processor(_Total)\% Processor Time 70% for 10 min 90% for 5 min
Available memory free -m / \Memory\Available MBytes 2 GB 1 GB or any swap-in
Network throughput sar -n DEV / \Network Interface(*)\Bytes Total/sec 70% of link for 10 min Flat at line rate for 15 min
Free space (data / log) df -h / \LogicalDisk(*)\% Free Space 20% / 30% 10% / 15%
Handle count of the server process ls /proc/PID/fd | wc -l / \Process(name)\Handle Count Rising week over week at constant load 80% of the handle limit

Two more hooks are not alerts but records. Keep the once-a-minute session sampling log from the limits article running permanently; it is the raw material for every review below. And keep the sizing worksheet, the last load-test results, and this checklist's change log in one folder per server. That way, the next person can see what was decided and why. Our alerting that gets read article covers making the pages above useful rather than ignored.

The Quarterly Capacity Review

The review is a one-hour meeting with yourself, once a quarter, that answers a single question: will this server still be fine at the next peak? Copy the template and fill it in each time.

Capacity review: sftp-01            Reviewer: ____   Quarter: ____

1. Peaks this quarter (from the session log and server log)
   max concurrent sessions: ____   p95: ____   logins/min peak: ____
   refusals: ____ (of which outside known bursts: ____)
2. Resources at the peak
   CPU: ____%   disk queue: ____   write latency: ____ ms
   available memory: ____ GB   network: ____ MB/s   free space: ____%
3. Limits vs reality
   global limit ____ vs peak ____  (margin ____%)
   ceiling from worksheet ____     (peak / ceiling = ____%)
4. Growth
   partners now ____ (last review ____)   volume now ____ GB/night (last ____)
   quarters until peak reaches 70% of ceiling at this rate: ____
5. Changes since last review (patches, hardware, partners, schedules): ____
6. Actions: retest? ____  resize? ____  spread partners? ____  adjust limits? ____
   next review date: ____

The number that matters most is line four: how many quarters until the peak reaches 70 percent of the ceiling. If the answer is fewer than the time it takes your organization to buy and build a server, the procurement starts now. It does not wait for the review where the answer is zero. If a change in line five is material — new storage, a server upgrade, fifty new partners — rerun the load test before trusting last year's knee.

Gotcha: the review is only as good as the session log behind it. If the sampling loop stopped three months ago because someone rebooted the server, line one is blank and the review is guesswork. Make the sampler a service, and check it is running as the first item of every review.

What to Take Away

Capacity tuning is a checklist, not a heroic diagnosis. Raise the operating system's file handle and process limits above the session limit. Size the listen backlog, socket buffers, and connection-tracking table for the peak. Put the data, logs, and system on separate volumes with storage that matches the IOPS estimate. Set every server limit and timeout deliberately and write down why. Wire the counters that reveal saturation into alerts with thresholds you chose in advance. Then review it every quarter against the measured peak, and the month-end stampede becomes a line in a log rather than a night on call.

Each section here is a summary of a longer article. The reasoning behind the resources is in how transfer servers handle load. The numbers come from sizing a transfer server. The night the checklist exists for is the month-end burst described earlier in this series.

Frequently Asked Questions

Which item on the checklist matters most?
Storage layout and type. A data volume on spinning disks shared with the logs is the most common cause of a transfer server collapsing under concurrent load. No other setting compensates for it. Fix that first, then work through the rest.
What does "too many open files" mean on a server with plenty of disk space?
It means the server process has hit its file handle limit, which counts network sockets and log files as well as data files. On Linux check ulimit -n for the service account, or LimitNOFILE in the systemd unit, and raise it well above the session limit.
How often should I rerun the load test?
After any material change: new storage, a hardware or VM resize, a server software upgrade, or a large batch of new partners. Otherwise once a year is enough, provided the quarterly review shows the peak is still comfortably below the last measured knee.
Do I need to tune the network stack on a modern operating system?
Usually only two things: check the listen backlog is not at an old small default, and raise the maximum socket buffer if partners are far away. The rest of the network items are checks rather than changes, looking for surprises such as a card negotiated to a low speed or a full connection-tracking table.
Should the checklist values be the same on every server?
The checks should be the same; the values should not. A server with 300 partners and a server with 20 need different limits. A server on solid-state storage tolerates concurrency that a spinning-disk server cannot. Derive each value from that server's sizing worksheet and record the reasoning.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.