Home › Topics › Storage Growth › The Disk Eater

Finding the Directory That Quietly Ate the Disk

"It'll be the archive. It's always the archive." The data volume is at 97%, uploads are failing, a partner is on the phone, and the channel is filling with confident guesses. Sometimes the guess is right; you still have to prove it. A server with a few hundred thousand files spread over dozens of folders can hide a terabyte almost anywhere. The folder people blame first is rarely the one that grew. Guessing is slow. What you need is a method that starts at the top of the tree and narrows to the culprit in a few steps. It uses tools that are already installed.

This article is that method. Confirm which volume is actually full. Size directory trees with the everyday tools on Linux and Windows. Sort by age rather than just by size. Check the usual suspects first, and know where space hides from directory tools. It ends with a worked hunt from the first alert to the culprit in about twenty minutes. That is less time than the argument about the archive usually takes. It is part of our Storage Growth series, and it assumes you have read why transfer servers fill up. That article names the patterns you will be looking for.

Step One: Confirm Which Volume Is Full

A volume is one formatted chunk of disk space that the operating system presents as a unit. That is a drive letter on Windows, or a mount point on Linux. A mount point is the folder where the volume is attached to the directory tree. A server usually has several volumes, and "the disk is full" always means one particular volume. Before you size anything, find out which one. On Linux, df -h lists every mounted volume with its size, used space, free space, and mount point:

$ df -h
Filesystem      Size  Used Avail Use% Mounted on
/dev/sda2        40G   19G   19G  50% /
/dev/sdb1       1.8T  1.7T   26G  99% /srv/transfer
/dev/sdc1       200G   61G  130G  32% /var/log

Read the Mounted on column first. Here the operating system volume (/) is fine, the log volume is fine, and the data volume mounted at /srv/transfer is the problem. Notice that Used plus Avail is less than Size. Linux filesystems typically reserve a small percentage for the system, so "full" for ordinary users arrives a little before the numbers reach 100%. On Windows, PowerShell's Get-PSDrive gives the same picture, and fsutil gives exact byte counts for one volume:

PS> Get-PSDrive -PSProvider FileSystem

Name  Used (GB)  Free (GB) Provider    Root
----  ---------  --------- --------    ----
C        41.20      78.80 FileSystem  C:\
D      1791.50      26.10 FileSystem  D:\
L        61.30     138.70 FileSystem  L:\

PS> fsutil volume diskfree D:
Total free bytes             : 28025176064
Total bytes                  : 1979120943104
Total quota free bytes       : 28025176064

Newer Windows builds also print a human-readable size in parentheses after each number. Either way, you now know which volume to hunt on, and you know the size of the hole you need to fill. There is about 26 GB free on a volume that wants at least 10% headroom. So you are looking for well over 150 GB to reclaim, not a few stray files.

Remember: if the full volume is the operating system volume, stop and treat it as an emergency. Free a few gigabytes immediately (temp folders, old update caches, the recycle bin) before you start any careful analysis. A full OS volume can take the whole server down. The volume layout article explains why that should never be possible in the first place.

Step Two: Size the Tree on Linux

The workhorse is du (disk usage), and the one incantation to memorize is this:

$ du -xh --max-depth=1 /srv/transfer | sort -h
2.1G    /srv/transfer/tmp
9.8G    /srv/transfer/staging
230G    /srv/transfer/archive
1.5T    /srv/transfer/partners
1.7T    /srv/transfer

Each flag earns its place. -x keeps du on the one volume you are hunting. That way, a mounted archive volume or a network share under the tree does not get counted. -h prints human-readable sizes. --max-depth=1 shows only the immediate children plus the total on the last line, which is exactly the "which branch?" view you want. sort -h understands those human-readable suffixes, so 230G sorts after 9.8G. The biggest entry lands at the bottom, just above the total. It takes a while on a big tree, because du genuinely walks every file; expect a minute or two per terabyte. It is not hung; it is counting.

Now descend. The biggest branch is partners at 1.5 TB, so run the same command against it, and then against whichever child wins:

$ du -xh --max-depth=1 /srv/transfer/partners | sort -h
1.2G    /srv/transfer/partners/contoso
44G     /srv/transfer/partners/acme
1.4T    /srv/transfer/partners/northwind
1.5T    /srv/transfer/partners

$ du -xh --max-depth=1 /srv/transfer/partners/northwind | sort -h
12M     /srv/transfer/partners/northwind/inbox
3.1G    /srv/transfer/partners/northwind/archive
1.4T    /srv/transfer/partners/northwind/outbox
1.4T    /srv/transfer/partners/northwind

Three levels down and the culprit has a name. If you prefer to explore interactively, use ncdu. This text-mode disk-usage browser, available in most package repositories, walks the tree once. It then lets you move up and down with the arrow keys, sorted by size, without re-scanning. Start it with ncdu -x /srv/transfer; the -x means the same as it does for du. It can also delete from inside the browser, which is convenient and slightly dangerous. So read the section on what to delete before you touch that key.

Step Three: Size the Tree on Windows

Windows has no du built in, but PowerShell can sum a tree. To size one folder, pipe every file under it into Measure-Object and add up the Length property:

PS> (Get-ChildItem -Path D:\Transfer\partners\northwind -Recurse -File -Force -ErrorAction SilentlyContinue |
       Measure-Object -Property Length -Sum).Sum / 1GB
1433.7

-Recurse descends into subfolders. The -File flag skips folder entries themselves. The -Force flag includes hidden files. The -ErrorAction SilentlyContinue flag keeps one unreadable folder from aborting the count. Dividing by 1GB (PowerShell understands the suffix) gives gigabytes. To get the "which branch?" view that du --max-depth=1 gives, wrap that in a loop over the top-level folders:

PS> Get-ChildItem -Path D:\Transfer -Directory | ForEach-Object {
      [PSCustomObject]@{
        Folder = $_.FullName
        GB     = [math]::Round((Get-ChildItem $_.FullName -Recurse -File -Force -ErrorAction SilentlyContinue |
                 Measure-Object -Property Length -Sum).Sum / 1GB, 1)
      }
    } | Sort-Object GB -Descending | Select-Object -First 10

Folder                          GB
------                          --
D:\Transfer\partners        1478.3
D:\Transfer\archive          230.4
D:\Transfer\staging            9.8
D:\Transfer\tmp                2.1

Change the path and run it again to descend, exactly as with du. It is slow on large trees, but it needs nothing installed. For repeated hunts, a graphical disk-usage viewer shows the whole volume at a glance. Any that draws the tree as nested rectangles will do. It is worth having on every file server. The command-line version, however, works over a remote session at three in the morning, which is when you will need it. The disk did not check your calendar first.

Step Four: Sort by Age, Not Just by Size

Size tells you where the bytes are. Age tells you whether they are still needed. Every file carries a modification time (the last time its contents were written). On a transfer server the age of a file is a strong hint about its purpose. A file in an outbox that is three months old has almost certainly been collected or abandoned. On Linux, find filters by both size and age in one pass. This lists every file over 1 GB that has not been modified in ninety days, with its size:

$ find /srv/transfer -xdev -type f -size +1G -mtime +90 -print0 | xargs -0 du -h | sort -h | tail -5
2.9G    /srv/transfer/partners/northwind/outbox/inventory_full_YYYYMMDD.csv
2.9G    /srv/transfer/partners/northwind/outbox/inventory_delta_YYYYMMDD.csv
3.0G    /srv/transfer/archive/dw_extract_YYYYMMDD.bak
3.0G    /srv/transfer/archive/dw_extract_YYYYMMDD.bak.gz
3.1G    /srv/transfer/staging/vendor_catalog.zip.part

-xdev is find's version of "stay on this volume". -size +1G means larger than one gigabyte. The -mtime +90 flag means modified more than ninety days ago. The plus sign matters: -mtime 90 without it means exactly ninety days. -print0 with xargs -0 handles filenames with spaces safely. To measure how much space all files older than a threshold occupy, regardless of individual size, use the file list. Hand it to du and ask for a grand total:

$ find /srv/transfer/partners/northwind/outbox -xdev -type f -mtime +30 -print0 | du -ch --files0-from=- | tail -n 1
1.3T    total

That single line is often the whole diagnosis. There is 1.3 TB of files older than a month in a folder whose contents are supposed to be collected nightly. The PowerShell equivalent filters on LastWriteTime. This shows the twenty largest files older than ninety days, with their age in days rather than a date. The age in days is the number you actually reason with:

PS> Get-ChildItem D:\Transfer -Recurse -File -Force -ErrorAction SilentlyContinue |
      Where-Object { $_.LastWriteTime -lt (Get-Date).AddDays(-90) -and $_.Length -gt 1GB } |
      Sort-Object Length -Descending |
      Select-Object -First 20 FullName,
        @{n='GB';e={[math]::Round($_.Length/1GB,1)}},
        @{n='AgeDays';e={((Get-Date) - $_.LastWriteTime).Days}}

One caution about age. Some tools preserve the original timestamp when they copy, so a file that arrived yesterday can carry a modification time from two years back. If the ages look implausible, check the creation time as well (CreationTime in PowerShell, stat on Linux), and treat age as a hint rather than proof. Timestamps are witnesses, not confessions.

The Usual Suspects

Before you walk the whole tree, spend ten seconds on the places that fill transfer servers most often. They are the same five patterns from the previous article, with the folders where each one hides:

Suspect Where it hides How to confirm
Uncollected outbox A partner's outbox or out folder Hundreds of files, oldest is months old, no downloads in the activity log
Archive without expiry archive, processed, done, backup Largest branch; -mtime +365 finds thousands of files
Logs The server's log folder, job log folders, /var/log Many rolled files; one enormous current file means debug logging is on
Temp and partials tmp, staging, and inside inboxes find -name '*.part' -o -name '*.tmp' -o -name '*.filepart'
The stray project A folder with a person's name or "old", "misc", "migration" Nobody can say who owns it; nothing has changed in months

Check the server's own log folder even if the data volume is the one that is full. Some servers keep logs under the data root by default. A server such as Sysax Multi Server writes its activity logs with rollover, so its log folder holds a series of files rather than one growing one. A folder like that is healthy as long as something deletes the oldest files. The server health monitoring series treats log growth as a signal in its own right, starting with log growth and service-log health.

Acme's log volume taught me the debug-logging row of that table. A partner's uploads had been failing at odd hours. Someone turned on verbose logging to catch the next failure. The fault turned out to be on the partner's side, and the verbose setting stayed on. Nobody had set rollover for the verbose file, because nobody had planned to keep it. So it grew as one file for eleven weeks until the log volume filled and the server stopped writing any log at all. The top-level du found it in one pass: a single file occupying most of the volume. Turning the setting off took a minute. The rollover rule took ten. The on-call note now says which setting to check first when a log volume fills.

Space That Directory Tools Cannot See

Sometimes du adds up to far less than df says is used. That gap has a short list of causes, and knowing them saves an hour of confusion:

  • Deleted files still held open. On Linux, deleting a file that a process still has open removes the name but not the data. The space returns only when the process closes it. A rotated log that the server is still writing to is the classic case. lsof -nP +L1 lists open files whose names have been deleted, with the process holding them. Restart that process (or make it reopen its logs) and the space comes back.
  • Snapshots and shadow copies. Windows volume shadow copies keep old versions of changed blocks, and storage snapshots do the same on the array or hypervisor. Neither shows up as files. vssadmin list shadowstorage shows how much a Windows volume has given to shadow copies. A busy transfer volume with many changing files can hand over a surprising share.
  • Reserved space. Linux filesystems typically reserve a few percent for the system, which is why Used plus Avail falls short of Size in df. It is not lost, but ordinary users cannot write into it.
  • Running out of inodes. An inode is the small on-disk record that holds a file's metadata. A filesystem has a fixed number of them. A volume can report free bytes yet refuse new files because every inode is used. df -i shows inode usage, and a transfer folder with millions of tiny files is exactly how it happens. The many small files problem explains why those folders arise.
  • Other volumes mounted under the tree. Without -x or -xdev, a mounted archive volume gets counted as if it lived on the data volume. So the numbers add up to more than the volume holds rather than less. Use the flags.

A Worked Hunt: Twenty Minutes from Alert to Culprit

Here is the whole method as it plays out on a real server, with the times from the on-call log. The alert arrives at Mar 14 02:10: data volume at 99%, three partner uploads failed in the last ten minutes.

  1. 02:11, confirm the volume. df -h shows /srv/transfer at 99% with 26 GB free, while the OS and log volumes are healthy. That rules out the "everything is about to die" scenario, so the hunt can be careful rather than frantic.
  2. 02:13, stop the bleeding. The uploads that are failing are retrying every five minutes and leaving a fresh partial each time. Check the staging folder: find /srv/transfer/staging -name '*.part' -mmin -60 lists nine partials from the last hour, 20 GB in total. Those are safe to delete (their transfers will start again from scratch anyway), and doing so buys breathing room. It is the only deletion in the first fifteen minutes.
  3. 02:16, size the top level. du -xh --max-depth=1 /srv/transfer | sort -h takes about two minutes and puts partners at the bottom of the list with 1.5 TB. The archive is 230 GB, which is notable but not tonight's problem.
  4. 02:19, descend. The same command against partners shows northwind at 1.4 TB, and against northwind shows outbox at 1.4 TB. Two commands, four minutes.
  5. 02:24, check age. find .../northwind/outbox -type f -mtime +30 | wc -l reports several thousand files; the grand-total du puts 1.3 TB of them older than thirty days. The outbox is supposed to be collected nightly.
  6. 02:27, check the activity log. A search for northwind's account in the server log shows logins and downloads stopping abruptly some months ago, right after a password rotation. Their pull job has been failing to authenticate ever since, and nobody on either side noticed because your export job kept succeeding.
  7. 02:30, decide. The culprit is named and the cause is known. The files are exports the partner may still want, so nothing further gets deleted at 02:30. The on-call note records the folder, the size, the cause, and the recommendation. Contact the partner in the morning, agree what they need, then apply the outbox cleanup rule that should have existed all along.

Twenty minutes, one safe deletion, and a clear answer. Notice what did not happen: nobody deleted the archive because it "looked old", nobody rebooted anything, and nobody added disk in the dark. I have done all three at one time or another, and none of them found the folder. If the exchange runs through a scheduled tool such as Sysax FTP Automation, its job logs are one more place the failure would have been visible. That is a reason to keep job logs somewhere you actually read. The article why jobs fail silently covers that habit.

What to Do Once You Have Found It

Finding the folder is the diagnosis; what you do next depends on what the folder holds. The safe order is:

  1. Delete only what is unambiguously debris. Abandoned partials, temporary files older than any plausible transfer, extracted copies whose originals still exist. These are yours to remove, and removing them is usually enough to get the volume out of the danger zone.
  2. Move, do not delete, anything that might be a record. If the culprit is an archive or an outbox full of business data, move the old files rather than deleting them. Use another volume or an archive location. Whether they may be deleted at all is a retention question; the where transferred files accumulate article covers the governance side.
  3. Fix the cause, not just the symptom. An outbox that filled because a partner stopped collecting will fill again unless the partner's job is fixed and an age-based rule is added. Cleanup jobs and quotas that stop this from recurring are covered in the quotas and automated cleanup series.
  4. Write it down. Record the folder, the size, the growth pattern, and the fix. The next full-disk alert on this server should start from that note, not from zero.

Then measure the growth rate, because the hunt only tells you where the last terabyte went. The capacity planning article turns today's numbers into a runway, and growth monitoring makes sure the next culprit is found while it is still small. Culprits are much easier to interview at 40 GB than at 1.4 TB.

The Method in One Paragraph

Confirm the full volume with df -h or Get-PSDrive. Size the top level of that volume with du -xh --max-depth=1 | sort -h or the PowerShell folder loop. Then descend into the biggest branch until one folder holds most of the space. Sort that folder by age with find -mtime or a LastWriteTime filter to learn whether the files are still wanted. If the numbers do not add up, check for deleted-but-open files, shadow copies, and inodes. Delete only debris tonight, move the rest, and fix the rule that let the folder grow. That is the whole method, and it works on any transfer server you will ever be handed. They tend to be handed over at night.

Frequently Asked Questions

Why does du report less space than df says is used?
Usually because a deleted file is still held open by a process, so its space has not been released yet. The command lsof +L1 finds those. Snapshots, shadow copies, and the filesystem's reserved percentage also consume space without appearing as files. Check those three before assuming the tools are wrong.
Is it safe to delete files while the server is running?
Deleting abandoned temporary files and old partials is generally safe, since any transfer that needs them will start over anyway. Do not delete a file that a transfer might be writing right now, and never delete business data during a hunt. Move it aside and let the retention policy decide.
The du command is taking forever. Is something wrong?
Probably not. Both du and the PowerShell loop have to touch every file, so a tree with millions of small files can take many minutes. Run it against one branch at a time, or use ncdu, which scans once and then lets you browse without rescanning.
What is an inode and why would I run out of them?
An inode is the small record that stores a file's metadata on Linux filesystems. Each volume has a fixed number of them. A folder holding millions of tiny files can use them all up while plenty of bytes remain free. Then new files fail with a "no space" error. df -i shows how many are left.
Which folder should I check first?
The archive or processed folder and any partner outbox, because those two patterns account for most full transfer servers. Then the server's own log folder and any staging or temp area. The top-level du will confirm or correct your guess in a couple of minutes.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.