War Story: The Deleted Inbox
"We seem to be missing about five weeks of paperwork," the customs broker said at a quarter past nine. It was the voice of someone who assumes the fault is on their side and is being polite about it. It was not on their side. A housekeeping script had run without incident for two years. At five past midnight, it had deleted every file older than two weeks from every partner folder on a freight forwarder's transfer server. For nine hours the only system that knew was the one that had done it.
This is the postmortem of that night: estate, timeline, discovery, and the response including its mistakes. Then come the contributing factors, plural, because there is never one villain. Acme, Kestrel Customs, and the people are composites with invented names. The mechanism is exact, down to the line of shell that did the damage. This article is part of our War Stories series. It ends with a checklist that finds the same failure in your own estate in under an hour.
The Estate, and the One Partner Who Fetched Late
Acme is a mid-sized freight forwarder. Its document system publishes shipping paperwork (bills of lading, commercial invoices, packing lists, arrival notices) for nine partners. Each partner has an account on Acme's Windows SFTP server, xfer.example.com, jailed to its own tree:
D:\xfer\partners\<partner-code>\inbox files published FOR the partner to download D:\xfer\partners\<partner-code>\outbox files the partner uploads TO Acme D:\xfer\archive\<partner-code>\ Acme's own copy of everything exchanged, kept 14 days
The document system pushes files into partner inboxes over SFTP as a service account called docpub. A poller sweeps every outbox within minutes. The archive tree is the one that grows; a nightly purge trims it.
That purge does not run on the Windows server. A small Linux utility host, util01.example.com, mounts the transfer share at /srv/xfer. It runs the housekeeping from cron, because the administrator who wrote it preferred bash. That administrator left a year before the incident. The script stayed, as scripts do.
One partner matters most here. Kestrel Customs, a customs broker, does not fetch daily. Its clerks pull a shipment's paperwork when the vessel is a day or two from port, three to six weeks after Acme publishes it. So the Kestrel inbox routinely held four hundred or more documents. Most were older than two weeks, and all were still needed. Every other partner fetched daily. That asymmetry is why an estate-wide bug looked like a single-partner disaster.
The Timeline
- Mar 13 16:20 — Priya, the transfer administrator, closes a routine ticket. Kestrel wants its folder renamed to match the partner code in its new internal system. She renames
D:\xfer\partners\kestreltoD:\xfer\partners\kestrel-customsand updates the SFTP account's home directory. She has the partner test a login and closes the ticket. Eleven minutes. - Mar 14 00:05 — The nightly purge starts on
util01. It iterates the partner directories, looks each up in its configuration file, and prunes the matching archive folder. For the renamed directory the lookup finds nothing. The path variable is empty. The delete runs anyway, against the wrong root. - Mar 14 00:05–00:09 —
findremoves 431 files older than fourteen days from beneath/srv/xfer/partners: 412 in Kestrel's inbox, the rest stale test files elsewhere. The run takes four minutes instead of a few seconds. The log records exactly what happened. Nobody reads it. - Mar 14 01:00 — The nightly backup of the share runs, faithfully capturing the already-emptied inbox.
- Mar 14 08:50 — At Kestrel, a clerk pulling documents for four containers arriving that afternoon finds only the last two weeks. Tomas, Kestrel's operations lead, assumes their own fetch job misfiled something and spends twenty minutes checking.
- Mar 14 09:15 — Tomas calls Acme. Dana, the on-call administrator, picks up.
- Mar 14 09:50 — Dana finds the cause. The purge is disabled at 09:55.
- Mar 14 10:20 — Restore from the previous night's backup begins. It restores more than it should.
- Mar 14 11:40 — Kestrel's inbox verified complete against the server's own upload log.
- Mar 14 14:30 — Operations finishes reversing seventeen duplicate shipment updates created by the over-broad restore.
- Mar 17 10:00 — Postmortem meeting.
Discovery: Four Lines and a Trailing Space
The discovery came from the partner, not from monitoring, and the postmortem wrote that down first.
The first wrong turn
Dana's first theory was the natural one: Kestrel's own automation had deleted or moved the files. She searched the SFTP server's activity log for deletes by the kestrel account. None. By docpub. None. Forty minutes, and in hindsight the absence was the clue. Nothing had passed through the server's front door. So something had removed the files directly across the share. Only util01 did that, and its purge log told the story in four lines:
Mar 14 00:05 purge amberline: files older than 14 days from /srv/xfer/archive/amberline Mar 14 00:05 purge brightwater: files older than 14 days from /srv/xfer/archive/brightwater Mar 14 00:05 purge kestrel-customs: files older than 14 days from Mar 14 00:09 purge lindqvist: files older than 14 days from /srv/xfer/archive/lindqvist
A blank where a path should be, and a four-minute gap after it. Dana said later she understood the whole incident the moment she saw the trailing space. (A small thing to lose a morning to. Most of them are.)
The script
Here it is, condensed, every load-bearing line intact. The configuration file /etc/xfer/purge.conf held one row per partner: code, archive directory, retention in days.
#!/bin/bash
# purge-archive.sh -- nightly prune of each partner's archive folder
CONF=/etc/xfer/purge.conf
cd /srv/xfer/partners || exit 1 # glob the partner directories from here
for p in */ ; do
p=${p%/}
read -r ARCHIVE_DIR RETAIN_DAYS < <(awk -v p="$p" '$1==p {print $2, $3}' "$CONF")
echo "$(date '+%b %d %H:%M') purge $p: files older than ${RETAIN_DAYS:-14} days from $ARCHIVE_DIR"
find $ARCHIVE_DIR -type f -mtime +${RETAIN_DAYS:-14} -delete
done
Three details combine into the failure, each harmless on its own:
- The lookup is keyed on the directory name. Rename a directory without updating the configuration and the lookup returns nothing;
ARCHIVE_DIRis empty. The day count had a default; the path did not. A default for the dangerous variable would have been wrong too. What was missing was a refusal. - The variable is unquoted. Written
find "$ARCHIVE_DIR" ..., an empty value producesfind '', which fails with "No such file or directory" and deletes nothing. Writtenfind $ARCHIVE_DIR ..., an empty value simply vanishes from the command line. What actually ran wasfind -type f -mtime +14 -delete, with no starting point at all. - GNU find with no starting point searches the current directory. And the current directory, because of the
cdthat made the glob convenient, was/srv/xfer/partners. That was the root of every partner's inbox and outbox.
The Windows version of the same bug: in PowerShell, a path built as "$Root\archive" with $Root empty becomes \archive. That resolves against the root of the current drive. A cleanup aimed at D:\xfer\archive quietly becomes a cleanup of D:\archive. Different syntax, identical shape: an empty base path that changes the root instead of stopping the run.
What the Responders Did, Including the Mistakes
Disabling the job was the right first move and took thirty seconds. Recovery was where the mistakes happened, and each was reasonable under pressure. I have made two of the three myself.
The backup that ran after the purge
The share was backed up nightly at 01:00, fifty-five minutes after the purge. So the most recent backup contained the damage. The one before it, from Mar 13 01:00, contained everything. Every deleted file was at least fifteen days old, so all existed then. Good news with a sting in it. Backup retention was seven days. A week's delay in noticing instead of nine hours, and the last good copy would have rolled off the day someone went looking. The postmortem recorded this as a near miss, not a success.
Restoring more than was lost
The backup tool offered a one-click restore of the whole partners tree onto the live share. Dana took it, because a partner was waiting and a selective restore meant hand-picking paths. It put back Kestrel's 412 documents. It also put back every file sitting in any partner's outbox at 01:00 the previous day. Those were uploads long since swept and processed. The poller swept them again at 10:30 and re-submitted seventeen shipment updates to the transport system. Operations spent the afternoon reversing them. A restore is a write, and writes into a live flow get processed. The poller has no idea it is looking at history. Why duplicates happen lists restores as a classic source.
Verifying the restore
A backup answers "what existed at 01:00 yesterday," not "what should exist now." For that Dana used the activity log on Sysax Multi Server, written to file and, in Acme's setup, to a database. It recorded every upload docpub had made into the Kestrel inbox and every download kestrel had made from it. A query for files published in the previous forty-five days returned 927 names. After the restore the inbox held 927. The same log listed what Kestrel had already downloaded. Tomas used that to tell his clerks which shipments needed no re-fetch. No backup catalog could answer either question. Only the log of the flow could.
Seven Things That All Had to Be True
Had the postmortem stopped at "unquoted variable," it would have fixed one line and left the estate exactly as fragile. The meeting listed seven factors. Every one had to be true.
- The empty variable was allowed to proceed. No
set -u, no parameter check, no quoting. A script that deletes must refuse to run on a path it cannot name. - The script ran from the worst possible working directory, as an account with delete rights over the entire share. It needed only
/srv/xfer/archive. Least privilege would have turned a catastrophe into a permissions error in the log. - The configuration lived where the change process could not see it. The rename checklist covered the directory and the SFTP account, not
/etc/xfer/purge.confon another host. Nobody on the current team knew the file existed. The script had no owner. - No dry run, no ceiling, no allowlist. Nothing compared the path to an expected root or counted candidates before deleting them. Nothing asked whether 431 deletions was normal for a job that usually removed a few dozen.
- The log was written and never read. The blank path and the four-minute gap were on disk within seconds. No rule looked for either.
- Backup timing and retention matched the purge badly. The backup ran after the purge, and its retention was shorter than a plausible time-to-notice.
- One partner's legitimate pattern concentrated the damage. Kestrel's long dwell time was agreed in writing. But "this inbox holds five weeks of unfetched documents" was recorded nowhere an administrator would see it. So nobody treated the folder as the fragile thing it was.
Priya's rename appears nowhere on that list as a cause. She followed the checklist. The checklist was incomplete, and the script punished an incomplete checklist with deletion. The third factor is really an offboarding failure a year in the making. The author left, and nobody asked what jobs he had left behind or who owned them now. The blameless postmortem method exists to keep the conversation on the system rather than on whoever was holding the ticket.
What Was Actually Changed
The script was rewritten; the version below is the copyable part of this article. It refuses an empty, unlisted, nonexistent, or out-of-root path. It counts before it deletes and stops above a ceiling. It supports a dry run and never changes directory.
#!/bin/bash
# purge-archive.sh <partner-code> -- DRY_RUN=1 lists without deleting
set -euo pipefail
PARTNER=${1:?usage: purge-archive.sh <partner-code>}
CONF=/etc/xfer/purge.conf
ALLOWED_ROOT=/srv/xfer/archive
MAX_DELETE=${MAX_DELETE:-500}
ARCHIVE_DIR=""; RETAIN_DAYS=""
read -r ARCHIVE_DIR RETAIN_DAYS < <(awk -v p="$PARTNER" '$1==p {print $2, $3}' "$CONF") || true
: "${ARCHIVE_DIR:?no purge.conf row for $PARTNER - refusing to run}"
: "${RETAIN_DAYS:?no retention value for $PARTNER - refusing to run}"
# Only beneath the allowlisted root, and only if it really exists
case "$ARCHIVE_DIR" in
"$ALLOWED_ROOT"/?*) ;;
*) echo "refusing: $ARCHIVE_DIR is outside $ALLOWED_ROOT" >&2; exit 2 ;;
esac
[ -d "$ARCHIVE_DIR" ] || { echo "refusing: $ARCHIVE_DIR does not exist" >&2; exit 2; }
# Count first; above the ceiling, a human decides
COUNT=$(find "$ARCHIVE_DIR" -type f -mtime +"$RETAIN_DAYS" | wc -l)
[ "$COUNT" -le "$MAX_DELETE" ] || { echo "refusing: $COUNT candidates exceed $MAX_DELETE" >&2; exit 3; }
if [ "${DRY_RUN:-0}" = 1 ]; then
find "$ARCHIVE_DIR" -type f -mtime +"$RETAIN_DAYS" -print; exit 0
fi
find "$ARCHIVE_DIR" -type f -mtime +"$RETAIN_DAYS" -print -delete
echo "$(date '+%b %d %H:%M') purge $PARTNER: removed $COUNT files older than $RETAIN_DAYS days"
Around the script, the estate changed in five ways. The purge now iterates the rows of purge.conf, never the filesystem. A partner with no row is simply not purged. A weekly report lists directories without one. The cron account can now modify only archive. The configuration moved into version control beside the partner register. The rename checklist gained one line: search every job host and configuration directory for the old partner code before closing the ticket. Two alert rules watch the purge log: any refusal, and any count above three times the trailing average. And the backup moved to 23:30, ahead of the purge, with thirty-day retention. The restore procedure says: side location first, then copy only what was lost. One rewrite and five estate changes, for one trailing space. That is about the usual exchange rate. Age-based cleanup jobs covers the general design.
The Lessons, and Where to Learn Each Fix
Each lesson links to the article that teaches the fix properly. The story is the reason to read them.
- A delete must fail closed. Unset variables, empty strings, missing directories, and unexpected roots stop the script; they never redirect it.
set -euo pipefail, quoting, and${var:?message}open every script that removes anything. See hardening bash transfer jobs and bash error handling. - Dry-run before you trust, and count before you act. A list of what would be deleted, read once by a human, catches a wrong root instantly. A ceiling catches it every night after. Rsync dry runs and verification teaches the same habit for mirroring.
- Purge from the archive, never from the flow. Partner-facing folders need gentler retention rules that reflect how each partner actually collects: automated purge policies.
- Housekeeping accounts get the narrowest rights that work. Least privilege in practice and the blast radius of one account.
- A log nobody reads is a diary, not monitoring. The evidence was complete at 00:05. What to log and reading transfer logs cover what belongs in a job log and how to make anomalies visible.
- Restore narrowly, and expect reprocessing. Anything restored into a watched folder will be processed again. The way to make that harmless is in safe reprocessing patterns.
- Names have invisible consumers. A partner code appears in more places than the folder. Build the search into the change checklist, as the partner onboarding runbook describes.
Remember: the dangerous scripts in your estate are not the clever ones. They are the small, boring housekeeping jobs written years ago by someone who has left. They run with more rights than they need, from a working directory nobody has thought about.
Check Your Estate
An hour with this list finds the same failure waiting in most estates. I say so with some confidence. I have written the unguarded version of that purge script myself and been lucky rather than careful. Run it for every script that deletes or purges.
DELETE-SCRIPT CHECKLIST (one copy per script that removes files) [ ] Every path variable is quoted; the script starts with set -euo pipefail [ ] An empty, unset, or unlisted path stops the run; it never defaults [ ] The delete path is checked against an allowlisted root first [ ] Nothing depends on the current working directory [ ] A dry-run mode exists and a person has read its output at least once [ ] Candidates are counted first; above a ceiling the run refuses [ ] The account cannot modify anything outside the purge root [ ] Refusals and blank paths in the log raise an alert somebody receives [ ] Backups run before the purge; retention exceeds the time-to-notice [ ] The script has a named owner and a row in the job inventory [ ] Folders with long-dwell partner files are known and excluded
The last line deserves emphasis. Ask each partner how long files sit before they collect them. Write the answer in the partner register. Treat any long-dwell folder as fragile: excluded from every purge, backed up longer, named in the runbook. The line above it is the hit-by-a-bus test applied to one script. It would have caught this a year early.
The Version to Tell a Colleague
A purge script looked up each partner's archive path by directory name. A folder was renamed, and the lookup came back empty. The empty path was unquoted. GNU find with no starting point searched the current directory: the root of every partner's inbox. Everything older than two weeks went. The bug was estate-wide; the damage landed on the one partner whose inbox legitimately held five weeks of files. The fix was not one line. Refuse an empty path, allowlist the root, and count before deleting. Narrow the account, read the log, and move the backup ahead of the purge. And give the script an owner, because the last one left a year ago. The script, as scripts do, stayed.
The next story is the opposite failure, a job that did far too much: the looping job. For a flow that stopped without anyone noticing, read the file that never arrived. The method behind all of them is in running a blameless postmortem.
Frequently Asked Questions
Why did an empty variable delete files instead of just failing?
find $DIR -delete became find -delete. GNU find with no starting point searches the current directory. Quoted, it would have produced an error and deleted nothing.Would set -u have been enough on its own?
set -u stops a script that uses a variable never set. But this one was set, to an empty string, by a lookup that found no match. ${DIR:?message} stops on both unset and empty. Use both, and quote everything.Is a dry-run mode really worth adding to a ten-line script?
What should I do first if a script has deleted the wrong files?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
