A Capacity Tuning Checklist for Transfer Servers
Capacity problems on transfer servers are rarely caused by one big mistake. They are caused by a dozen small defaults nobody changed. Those include a file handle limit of 1024 on a server that holds four hundred sessions, and a listen backlog sized for a workstation. They include logs on the same disk as the data, an idle timeout of an hour, and no alert on the disk queue. Each is a five-minute fix. The trouble is remembering all of them, on every server, and checking again when the load has doubled since the build.
This article is that memory. It collects the settings and checks from the rest of our Capacity & Concurrency series into one checklist. The checklist walks the server from the bottom up: operating-system limits, the network stack, disk and storage, and the transfer server's own settings. It also covers the monitoring hooks that tell you when any of it drifts. It closes with the quarterly review that keeps capacity ahead of demand instead of one month-end behind it. Every item names the command that checks it and the value that is usually right. So a junior administrator can work through it without a senior one in the room.
How to Use the Checklist
Run it three times in a server's life: once at build, once after the first real peak, and then every quarter. Three rules make it useful rather than dangerous. Record the current value before changing anything — a text file of "was / now / why / date" per server is the whole change log. Change one setting at a time when the server is already in trouble, so you know which change helped. And retest after material changes using the ramp from load testing a transfer server safely, because a tuning change that was never measured is a hope.
The values suggested below assume the example server used throughout this series. It has 8 cores, 16 GB of memory, a gigabit interface, and around 300 partners with a month-end peak of about 310 concurrent sessions. Scale them to your own sizing worksheet from sizing a transfer server.
Section 1: Operating-System Limits
The operating system has ceilings of its own beneath the transfer server, and they fail with messages that never mention capacity.
- File handles per process. Every session, open file, and log costs at least one. On Linux the default soft limit is often 1024. Check with
ulimit -nas the service account (andulimit -Hnfor the hard limit). For a service started by systemd, the value that counts isLimitNOFILE=in the unit file. A sensible setting is 65536. The symptom when it is wrong is "too many open files" in the server log while disk space is fine. Windows processes have a far higher ceiling, but track\Process(name)\Handle Countfor a steady climb that indicates a leak. - System-wide file handles.
sysctl fs.file-maxis the machine's total;cat /proc/sys/fs/file-nrshows allocated and maximum. Rarely the limit on a dedicated server, but check once. - Processes and threads. A process-per-connection server needs a process limit above its session limit:
ulimit -u, orTasksMax=in the systemd unit. Set it to at least twice the global session limit. - Swap. A transfer server should have swap configured as insurance and never use it. On Linux set
vm.swappiness=10so the kernel prefers dropping cache to swapping the server. On Windows leave the page file system-managed and alert on\Memory\Available MBytes. Any sustained swap activity means the memory row of the worksheet was wrong. - Power management. Virtual and physical servers alike can be throttled by a balanced power plan. Windows:
powercfg /getactiveschemeshould report High performance. Linux: check the CPU frequency governor is not set to power saving on a server that encrypts.
$ ulimit -n; ulimit -Hn 1024 524288 $ sysctl fs.file-max vm.swappiness fs.file-max = 9223372036854775807 vm.swappiness = 60 $ cat /proc/sys/fs/file-nr 5344 0 9223372036854775807
That output is a server that needs two changes: the soft handle limit is the 1024 default, and swappiness is at the desktop value of 60. The OS-level hardening article covers the security side of these same files, and the two reviews are worth doing together.
Section 2: The Network Stack
These settings live below the transfer server and above the network card. Most have safe defaults on modern systems, but each one has a failure mode that looks like something else.
- Listen backlog. The queue of completed connections waiting for the server to accept them. On Linux,
sysctl net.core.somaxconnshould be at least 4096 (older kernels default to 128); the server's own listen setting caps it too. Check the live value withss -ltn:Send-Qis the backlog size for each listening port. ARecv-Qthat approaches it during a burst means connections are timing out at the door. - Ephemeral port range. Only matters when the server makes outbound connections — active-mode FTP data connections, or a gateway forwarding to a back-end. Linux:
sysctl net.ipv4.ip_local_port_range, normally32768 60999. Windows:netsh int ipv4 show dynamicport tcp, normally starting at 49152 with 16384 ports. Widen only ifTIME_WAITcounts approach the range size. TIME_WAITcount.ss -sreports it;netstat -an | find /c "TIME_WAIT"on Windows. On an FTP server, aTIME_WAITcount that dwarfs established connections means data connections are being recycled faster than ports free up. In that case, widen the passive range rather than shortening kernel timers.- Socket buffer limits. Distant partners need large TCP windows to fill the pipe. Linux autotunes within
net.ipv4.tcp_rmemandnet.ipv4.tcp_wmem; raise the third (maximum) value to 16 MB if partners are more than about 50 ms away. Windows:netsh int tcp show globalshould show receive window autotuning atnormal. Our article on TCP tuning first covers the arithmetic. - Interface speed and duplex. A gigabit card negotiated down to 100 Mbit/s half duplex is a surprisingly common finding. Linux:
ethtool eth0 | grep -E 'Speed|Duplex'. Windows:Get-NetAdapter | Select-Object Name, LinkSpeed. - Connection tracking. A host firewall tracks every connection in a table with a maximum. Linux:
sysctl net.netfilter.nf_conntrack_maxandnet.netfilter.nf_conntrack_count; when count reaches max, new connections are silently dropped. Set the maximum to several times the peak connection count, remembering that each FTP transfer is a connection of its own. The same applies to the perimeter firewall, which our firewalls and NAT series covers.
Section 3: Disk and Storage
Storage is where transfer servers actually run out of capacity, so this section deserves the most time.
- Separate volumes. The operating system, the server's logs, and the transfer data must be on different disks or virtual disks. Add a fourth for a logging database if you use one. A log volume that fills stops the server; a log stream sharing spindles with three hundred uploads steals IOPS from all of them.
- Storage type matches the IOPS estimate. Take the disk row of the sizing worksheet — concurrent transfers times operations per stream — and compare it with what the volume can do. Spinning disks deliver one to two hundred random IOPS each; if the estimate is in the hundreds or thousands, the data volume belongs on solid-state storage. Confirm during a real peak with
iostat -x 5(aqu-szunder 2 per spindle,w_awaitunder 20 ms spinning or 5 ms solid-state). Or use thePhysicalDiskcountersAvg. Disk Queue LengthandAvg. Disk sec/Write. - Virtual disk quotas. On virtual infrastructure, find the IOPS and throughput quota attached to the data disk. It is frequently the real ceiling, far below the hardware beneath it.
- Last-access timestamps. By default some filesystems write a timestamp every time a file is read, turning every download into a write. Linux: mount the data volume with
noatime(or confirmrelatime). Windows:fsutil behavior query disablelastaccess; disable updates on the data volume if they are enabled. - Directory sizes. Listing a directory of 50,000 files costs CPU and IOPS every time a partner's client asks, and scripted clients ask often. Keep active directories under a few thousand entries with archive subfolders; the many-small-files problem article explains why.
- Scanning and indexing. Real-time antivirus scanning of the data volume doubles the I/O of every upload. Scan at a defined point in the flow instead, as our article on scanning integration points describes, and exclude the data volume from any search indexer.
- Free space thresholds. Alert at 20 percent free on the data volume and 30 percent on the log volume. A burst needs room for a full night's arrivals plus temporary files. Growth and cleanup are the subject of our storage growth and quotas and cleanup series.
Remember: when the checklist finds the data volume on spinning disks shared with the logs, stop and fix that before tuning anything else. No kernel setting compensates for a saturated disk queue.
Section 4: Transfer Server Settings
The server's own configuration is where the earlier articles in this series land. Every value here should be written down next to the reason for it.
- Global session limit. Set, not unlimited: about 70 percent of the resource ceiling and at least 25 percent above the measured peak. Example: ceiling 600, peak 311, limit 400.
- Per-user cap. Twice a normal client's parallelism, minimum 2; higher for named batch accounts. Example: 4.
- Per-IP cap. Per-user cap times accounts expected behind one address, plus 2; a short exception list for NAT-heavy partners. Example: 6.
- Login rate limit. Enough unauthenticated connections to absorb a burst of logins without pinning the CPU; on OpenSSH,
MaxStartups 20:30:60for an eight-core server. Pair it with brute-force auto-blocking, which our lockout and throttling design article covers. A server such as Sysax Multi Server provides IP allow and block lists with automatic blocking. That keeps scanners from spending your login budget. - Passive port range (FTP). Twice the expected simultaneous transfers, opened on the firewall, with the correct external address announced. Example: 200 ports. The passive range guide has the procedure.
- Idle timeout. Long enough for a legitimate pause between commands, short enough to reap ghosts: 10 minutes for partners, shorter during bursts. Our timeouts and keepalives series covers the interaction with slow clients.
- Transfer timeout. A stalled transfer should be dropped after a few minutes of no data, not left holding a slot and a passive port for an hour.
- Ciphers and compression. Prefer a hardware-accelerated AES mode on CPUs that support it; disable SSH-level compression, which costs a core per session for almost nothing on already-compressed files. Our cipher policy basics article has the policy side.
- Logging level. Debug-level logging under load is a write per command per session and can halve throughput. Run at the normal level and enable debug per session or per account when diagnosing.
- Service recovery. On Windows, the service's recovery options should restart it on failure, with a short delay. That way, a crash at three in the morning does not wait for a human.
The details of choosing each number, and what the server says when a limit is hit, are in connection limits and per-user caps.
Section 5: Monitoring Hooks
Every item above can drift: a partner grows, a patch resets a default, a volume fills. The monitoring hooks are how you find out before the partners do. The server health monitoring series covers the tooling. This section is the list of what to collect and the thresholds that turn a number into a page.
| Metric | Source | Warn | Critical |
|---|---|---|---|
| Concurrent sessions | ss -s / \TCPv4\Connections Established |
70% of global limit | 90% of global limit |
| Refusals per hour | Server log (421, MaxStartups, refusal events) |
Any outside a known burst | More than 5% of logins |
| Disk queue length | iostat -x aqu-sz / Avg. Disk Queue Length |
2 per spindle for 5 min | 4 per spindle for 5 min |
| Write latency | w_await / Avg. Disk sec/Write |
20 ms | 100 ms |
| CPU busy | vmstat / \Processor(_Total)\% Processor Time |
70% for 10 min | 90% for 5 min |
| Available memory | free -m / \Memory\Available MBytes |
2 GB | 1 GB or any swap-in |
| Network throughput | sar -n DEV / \Network Interface(*)\Bytes Total/sec |
70% of link for 10 min | Flat at line rate for 15 min |
| Free space (data / log) | df -h / \LogicalDisk(*)\% Free Space |
20% / 30% | 10% / 15% |
| Handle count of the server process | ls /proc/PID/fd | wc -l / \Process(name)\Handle Count |
Rising week over week at constant load | 80% of the handle limit |
Two more hooks are not alerts but records. Keep the once-a-minute session sampling log from the limits article running permanently; it is the raw material for every review below. And keep the sizing worksheet, the last load-test results, and this checklist's change log in one folder per server. That way, the next person can see what was decided and why. Our alerting that gets read article covers making the pages above useful rather than ignored.
The Quarterly Capacity Review
The review is a one-hour meeting with yourself, once a quarter, that answers a single question: will this server still be fine at the next peak? Copy the template and fill it in each time.
Capacity review: sftp-01 Reviewer: ____ Quarter: ____ 1. Peaks this quarter (from the session log and server log) max concurrent sessions: ____ p95: ____ logins/min peak: ____ refusals: ____ (of which outside known bursts: ____) 2. Resources at the peak CPU: ____% disk queue: ____ write latency: ____ ms available memory: ____ GB network: ____ MB/s free space: ____% 3. Limits vs reality global limit ____ vs peak ____ (margin ____%) ceiling from worksheet ____ (peak / ceiling = ____%) 4. Growth partners now ____ (last review ____) volume now ____ GB/night (last ____) quarters until peak reaches 70% of ceiling at this rate: ____ 5. Changes since last review (patches, hardware, partners, schedules): ____ 6. Actions: retest? ____ resize? ____ spread partners? ____ adjust limits? ____ next review date: ____
The number that matters most is line four: how many quarters until the peak reaches 70 percent of the ceiling. If the answer is fewer than the time it takes your organization to buy and build a server, the procurement starts now. It does not wait for the review where the answer is zero. If a change in line five is material — new storage, a server upgrade, fifty new partners — rerun the load test before trusting last year's knee.
Gotcha: the review is only as good as the session log behind it. If the sampling loop stopped three months ago because someone rebooted the server, line one is blank and the review is guesswork. Make the sampler a service, and check it is running as the first item of every review.
What to Take Away
Capacity tuning is a checklist, not a heroic diagnosis. Raise the operating system's file handle and process limits above the session limit. Size the listen backlog, socket buffers, and connection-tracking table for the peak. Put the data, logs, and system on separate volumes with storage that matches the IOPS estimate. Set every server limit and timeout deliberately and write down why. Wire the counters that reveal saturation into alerts with thresholds you chose in advance. Then review it every quarter against the measured peak, and the month-end stampede becomes a line in a log rather than a night on call.
Each section here is a summary of a longer article. The reasoning behind the resources is in how transfer servers handle load. The numbers come from sizing a transfer server. The night the checklist exists for is the month-end burst described earlier in this series.
Frequently Asked Questions
Which item on the checklist matters most?
What does "too many open files" mean on a server with plenty of disk space?
How often should I rerun the load test?
Do I need to tune the network stack on a modern operating system?
Should the checklist values be the same on every server?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
