Home › Topics › Bandwidth Management › Measuring

Measuring Transfer Impact on the Network

Every fix in this series — a throttle, a scavenger class, a window, a connection cap — is a claim that something got better. Without measurement, that claim is an opinion, and opinions lose arguments to whoever complains loudest. The administrator who can say "the link peaked at 96 percent for eighty minutes on Tuesday. Latency for the branch went from 18 to 340 milliseconds. And 70 percent of the bytes were the backup account" gets the window approved. The one who says "I think it's the backup" gets a ticket reassigned.

This article is about the measurement that makes bandwidth management stick. It covers what three numbers to watch and how to read them from interface counters on the transfer host. It explains what the network team's SNMP graphs show and hide, and how to spot saturation from a rising ping. It shows how to work out from server logs which account or job took what share. It covers how to compare before and after a change honestly. It shows how to turn it all into a weekly report short enough that someone reads it. It is the last article in our Bandwidth Management series. Measuring the speed of a transfer as a benchmark is a different discipline, covered in our Benchmarking Transfer Performance series. This article is about the effect a transfer has on everyone else.

The Three Numbers

Bandwidth impact comes down to three measurements, and everything else in this article is a way of obtaining one of them.

Utilization is how full the link is: bytes sent per second, divided by the link's capacity, expressed as a percentage. It is measured at an interface — on the edge router by the network team, or on the transfer host by you. It is always measured in one direction at a time, because a link that is 96 percent full outbound may be 20 percent full inbound.

Latency under load is the round-trip time other users experience while the transfer runs, compared with the round-trip time when it does not. Utilization says the link is busy; latency says whether that busyness is hurting anyone. A link at 90 percent with stable latency is a link with good queue management. A link at 90 percent with latency ten times its idle value is a bloated buffer, as described in why bulk transfers crush the WAN.

Per-flow share is how the bytes divide between jobs, accounts, and partners. This is the number that turns "the link was full" into "the backup account moved 31 GB between 16:00 and 17:20." It comes from server and job logs rather than from the network. With these three you can set a bandwidth budget for each flow. That is an agreed amount of bytes per day and an agreed ceiling during business hours. You can check it weekly.

Reading Interface Counters

Every network interface keeps a running count of bytes sent and received since boot. Read the count twice, subtract, divide by the seconds between readings, and multiply by eight. You have the average rate in bits per second over that interval. That is all a utilization graph is: this calculation, repeated and plotted. The first article in this series used it once by hand. Here it becomes a logger that writes a line every minute so you have a record to look back at.

#!/bin/bash
# ratelog.sh: append "time,tx_mbit,rx_mbit" for eth0 every 60 seconds
IF=eth0; OUT=/var/log/ratelog.csv
t1=$(cat /sys/class/net/$IF/statistics/tx_bytes)
r1=$(cat /sys/class/net/$IF/statistics/rx_bytes)
while sleep 60; do
  t2=$(cat /sys/class/net/$IF/statistics/tx_bytes)
  r2=$(cat /sys/class/net/$IF/statistics/rx_bytes)
  printf "%s,%d,%d\n" "$(date +%H:%M)" $(( (t2-t1)*8/60/1000000 )) $(( (r2-r1)*8/60/1000000 )) >> $OUT
  t1=$t2; r1=$r2
done

# One-off views without a script:
ip -s link show eth0          # totals since boot
sar -n DEV 5 12               # 12 samples, 5 seconds apart, per interface

On Windows, the same counters are available through PowerShell or the built-in performance counters. typeperf is the quickest way to log them to a file:

# Bytes sent per second, sampled every 60 seconds, 1440 samples (one day), to CSV
typeperf "\Network Interface(*)\Bytes Sent/sec" -si 60 -sc 1440 -f CSV -o C:\logs\ratelog.csv

# Quick check of the totals since boot
Get-NetAdapterStatistics -Name "Ethernet" | Select-Object SentBytes, ReceivedBytes

Remember that typeperf reports bytes; multiply by eight for bits. Keep the resulting file for at least a month. A spreadsheet chart of it is a perfectly good utilization graph, and it has one advantage over the network team's. It shows what this host sent, so a spike on it is unambiguously a transfer from this machine.

SNMP in Plain Words

The network team's graphs come from SNMP (simple network management protocol). It is nothing more mysterious than a way for a monitoring server to ask a router "what is the value of your outbound byte counter?" every few minutes. The router answers with the same kind of running count you just read from /sys. The monitoring server subtracts the previous answer and plots the rate. The counters have names — the 64-bit ones are called ifHCInOctets and ifHCOutOctets — and you may hear them, but you never need to touch them. What you do need is read access to the graph of the WAN interface. That is a reasonable thing to ask for. Most network teams grant it to a colleague who asks with a specific reason.

There is one thing SNMP graphs routinely hide, and it matters for transfers: the polling interval. Most monitoring polls every five minutes, so each point on the graph is the average over five minutes. A transfer that pins the link at 100 percent for ninety seconds and then stops shows on a five-minute graph as a bump to 30 percent. The phones froze for ninety seconds; the graph says nothing happened. The diagram below shows the difference.

Two charts of the same ten minutes. The upper chart, sampled every ten seconds, shows the link at one hundred percent for ninety seconds. The lower chart, averaged over five-minute polling intervals, shows the same event as a single bar at thirty percent.

Two consequences. When you are hunting a specific problem, ask for one-minute polling on the WAN interface for a week, or rely on your own per-minute host logger, which sees the peak. And when you read a long-term graph, look at the 95th percentile line if the tool draws one. It is the level the link stayed below for 95 percent of the intervals. This ignores the very worst spikes but is far more honest than the average. A link whose 95th percentile sits above 80 percent of capacity during business hours is a link that is saturated regularly, whatever the average says.

Spotting Saturation from Latency

Utilization needs access to counters; latency needs only a ping. A steady stream of pings from inside the site to the first router on the far side of the WAN link is the cheapest saturation detector there is. Log it with timestamps. A full queue shows up as a jump in round-trip time before anyone has opened a ticket. The idle baseline for a branch circuit is typically tens of milliseconds and stable; under a bulk transfer it rises into the hundreds and wanders.

# Linux: one ping every 5 seconds to the far side of the link, timestamped
ping -i 5 10.20.0.1 | while read line; do
  echo "$(date +%H:%M:%S) $line"
done >> /var/log/pinglog.txt

# Sample lines: idle, then the backup starts at 16:00
15:58:10 64 bytes from 10.20.0.1: icmp_seq=21 ttl=62 time=18.4 ms
16:00:35 64 bytes from 10.20.0.1: icmp_seq=50 ttl=62 time=287 ms
16:00:40 64 bytes from 10.20.0.1: icmp_seq=51 ttl=62 time=341 ms

# PowerShell equivalent (older Windows PowerShell names the field ResponseTime)
while ($true) { $r = Test-Connection 10.20.0.1 -Count 1
  "{0:HH:mm:ss} {1} ms" -f (Get-Date), $r.Latency >> C:\logs\pinglog.txt
  Start-Sleep 5 }

A working rule of thumb: suppose the loaded round trip is more than three times the idle round trip. If it stays that way for longer than a minute, the link is saturated and the queue is bloated. Line the ping log up against the rate log by time and the picture is complete. The rate log says the host was sending 48 Mbit/s. The ping log says latency went to 300 ms at the same minute. The two together are the evidence pack from the first article, now collected automatically. Note that the ping target must be beyond the congested link; pinging the local router tests only the LAN. The subtler question of whether a transfer is slow because of latency or because of bandwidth is a diagnosis in its own right, covered in our Diagnosing Slow Transfers series.

Per-Flow Accounting from Server Logs

Counters tell you the link was full; logs tell you who filled it. Every transfer server records each session and, in most cases, the bytes moved per file. So a few lines of text processing produce a table of bytes per account per hour. That table is the basis for fairness discussions, partner conversations, and the per-flow budget.

OpenSSH's SFTP server logs at the INFO level when its subsystem line is configured with -l INFO. Each file close then produces a line with the bytes read and written, and each session start names the account. The script below joins the two by process id and sums bytes per account per hour. The log format is the usual syslog one, with lines like Mar 14 02:10:33 sftp01 sftp-server[4821]: close "/inbound/backup.tar" bytes read 0 written 31457280000.

awk '
  /sftp-server\[[0-9]+\]: session opened for local user/ {
    match($0, /sftp-server\[[0-9]+\]/); pid = substr($0, RSTART, RLENGTH)
    for (i = 1; i <= NF; i++) if ($i == "user") acct[pid] = $(i+1)
  }
  /sftp-server\[[0-9]+\]: close / {
    match($0, /sftp-server\[[0-9]+\]/); pid = substr($0, RSTART, RLENGTH)
    hour = substr($3, 1, 2); r = 0; w = 0
    for (i = 1; i <= NF; i++) { if ($i == "read") r = $(i+1); if ($i == "written") w = $(i+1) }
    bytes[acct[pid] " " hour] += r + w
  }
  END { for (k in bytes) { split(k, p, " "); printf "%-12s hour %s  %8.2f GB\n", p[1], p[2], bytes[k] / 1e9 } }
' /var/log/auth.log | sort

# Sample output (bytes count in the hour the file closed)
backup       hour 17      31.46 GB
partner-a    hour 09       5.80 GB
partner-b    hour 09       0.41 GB

Read as a share, that is the whole story of the afternoon: one account moved 31 GB, finishing in the hour the phones failed. Traditional FTP servers keep a transfer log with one line per file including the byte count and account, which sums the same way with a simpler script. Windows transfer servers keep activity logs in their own format. A server such as Sysax Multi Server records per-account activity that can be exported and summed by account and hour in a spreadsheet. That is entirely adequate for a weekly report. Whatever the source, the techniques for reading these logs, and for gathering them from several servers into one place, are in reading transfer logs and centralizing logs.

Job logs fill in what server logs cannot. A scheduled job that records its own start time, end time, and bytes moved gives you the per-job view directly. That includes jobs that push to a partner's server whose logs you will never see. Both rsync and lftp print totals at the end of a run. Capturing those lines into the job's log is a one-line change that pays for itself the first time a partner asks why their file was late.

Remember: the network sees bytes; only the logs see names. Keep both. A utilization spike with no matching log entry means an unmanaged transfer exists somewhere, and finding it is more valuable than any throttle.

Before-and-After Comparisons That Hold Up

Every change in this series should be followed by a comparison, and a comparison is only fair if the two sides differ in one thing. The rules are the same ones that make a benchmark honest, spelled out in why transfer benchmarks lie, but the impact version is short.

  • Same job, same weekday, same hour. Tuesday's backup against Tuesday's backup a week later. Monday and Friday traffic differ; month-end differs from everything.
  • Same three numbers each side. Peak utilization, loaded latency, and the flow's bytes and duration. Record all three even when only one was expected to change, because the throttle that fixed latency may have pushed the job past its window.
  • Normalize for size. If the job moved 30 GB one week and 36 the next, compare megabytes per second and minutes per gigabyte, not wall-clock time.
  • Note what else changed. A partner's server upgrade, a new application at the site, a carrier fault. One line in the record saves an hour of confusion later.
Measure Before (backup at 16:00, unthrottled) After (moved to 01:40, scavenger mark)
Peak outbound utilization, business hours 96% 44%
Loaded latency to far side, business hours 340 ms 21 ms
Backup bytes and duration 31.5 GB in 82 min 31.8 GB in 79 min
Helpdesk tickets mentioning phones or RDP 4 0

The last row is not a network measurement, and it is the one management reads. Keep it.

The Weekly Report

Measurement that nobody looks at decays into measurement that nobody collects. A weekly report, one page, the same layout every week, is what keeps windows from silently shrinking and partners from silently widening. It should take fifteen minutes to produce from the logs above and thirty seconds to read.

Section Contents Where it comes from
Link health, per WAN link Business-hours peak and 95th percentile; hours above 80%; worst loaded latency SNMP graph or rate log; ping log
Top flows Five largest accounts or jobs by bytes, with share of total and daytime bytes Server log accounting
Budget check Each budgeted flow: agreed bytes and daytime ceiling versus actual; flag breaches Accounting plus the inventory
Window health Window fill percentage; jobs that overran, landed softly, or were deferred by admission control Job logs
Unexplained Utilization spikes with no matching job or account Rate log versus accounting
Actions At most three, each with an owner You

The budget line deserves a definition, because it is the thing this whole series has been building toward. A bandwidth budget per flow is a written agreement, per job or partner, of three numbers. These are bytes per day you expect, the ceiling during business hours (often zero, meaning "off-peak only"), and the window it runs in. The backup's budget might read "30 GB per night, zero daytime bytes, window 01:40 to 06:00." Partner A's might read "6 GB per day, 8 Mbit/s daytime ceiling, four connections." Once the numbers are written down, the weekly report is simply a comparison of actual against agreed. A breach is then a conversation with a name attached rather than a mystery. Where the report should live, and how to present it alongside job success and failure, is covered in transfer dashboards and reporting. The wider health of the server itself is the subject of our Transfer Server Health Monitoring series.

The Version to Tell a Colleague

Three numbers describe a transfer's impact: how full the link is, how much latency other users see while it runs, and which account or job moved the bytes. Interface counters on the host give the first. A timestamped ping to the far side of the link gives the second. Server logs summed per account per hour give the third. The network team's SNMP graphs are useful but average over five minutes, so ask for the 95th percentile and keep your own per-minute log when hunting a spike. Compare before and after on the same job and weekday, normalized for size, and record what else changed. Then write a one-page weekly report with a budget per flow, so that every limit in this series is checked against reality instead of remembered.

If you have read the series in order, the loop is now closed: the problem, the fixes, and the measurement that proves them. With a week of data in hand, one article most worth revisiting is off-peak scheduling, to check the window arithmetic against real fill. The other is fairness between flows, to see whether the top-flows table shows anyone taking more than their share.

Frequently Asked Questions

The network team's graph shows the link at 30 percent, but users say it was unusable. Who is right?
Probably both. A five-minute SNMP average turns ninety seconds at 100 percent into a bump at 30 percent. Ask for one-minute polling, or log the interface counters on the transfer host every minute yourself, and compare against a timestamped ping log.
What is a 95th percentile, in plain words?
Sort every measurement interval in a period from lowest to highest and take the value 95 percent of the way up. It is the level the link stayed below almost all the time, ignoring the rarest spikes. A business-hours 95th percentile above 80 percent of capacity means the link is regularly saturated.
How do I find out which user or partner used the most bandwidth?
From the transfer server's logs, not from the network. SFTP and FTP servers log bytes per file with the account name; sum them per account per hour. Windows transfer servers with activity logs can be exported and summed in a spreadsheet.
Where should I ping to detect saturation on the WAN link?
A stable address on the far side of the congested link, such as the first router beyond your edge or the head-office gateway. Pinging your own local router only tests the LAN. Compare loaded against idle; more than three times the idle round trip, sustained, means the queue is full.
What goes in a bandwidth budget for a flow?
Three agreed numbers per job or partner: expected bytes per day, the ceiling allowed during business hours (often zero), and the window it runs in. The weekly report then compares actual against agreed, and a breach becomes a conversation with a named owner.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.