Home › Topics › War Stories › The Silent Miss

War Story: The File That Never Arrived

Grace in finance had the bank statement on one screen and the ledger on the other. The card settlements stopped on Jun 09 like a line drawn with a ruler. Fourteen stores, three weeks, nothing. Her ticket to IT read settlement import broken?. The question mark turned out to be the most accurate punctuation in the incident. For three weeks every system involved had reported that all was well. The SFTP server saw no failed logins. The import job exited zero every day. The partner's monitoring showed no errors. The daily summary email went out with the subject line it always had.

This is the postmortem of a silent miss: the failure that consists of nothing happening. It is the most common serious incident in file transfer and the least alarmed. Almost all monitoring is built to notice events, and an absence is not an event. Northgate Retail, Pennywhistle Payments, and the people are composites; the mechanism is exact. It is part of our War Stories series. It ends with a freshness check you can copy. It also has a checklist for the feeds in your own estate that could vanish without a sound.

The Estate and the Flow

Northgate Retail runs fourteen stores. Its card processor, Pennywhistle Payments, delivers a settlement file every morning. One CSV covers all stores. It lists the previous day's card transactions and the amounts to be paid into the retailer's bank. The flow is a push. The partner connects to Northgate's SFTP server, sftp.example.com, as the account pennywhistle. It uploads the file into its inbox around 04:00. At 05:00, a scheduled import job on jobs01.example.com picks up whatever is in that inbox. It loads it into the ledger and moves the file to an archive folder.

04:00  partner uploads   /partners/pennywhistle/inbox/SETTLE_daily_<day>.csv     (push, partner-initiated)
05:00  import job        loads every file in the inbox into the ledger, moves it to D:\feeds\archive\pennywhistle\
05:01  summary email     "Settlement import complete" to a shared finance mailbox

Three kinds of monitoring existed, each reasonable. The SFTP server alerted on failed logins and lockouts. The import job alerted when it exited with an error. The partner's platform alerted its operators when a sending job failed. Nothing anywhere alerted when a file that should have arrived did not.

The Timeline

  • Jun 09 04:02 — The last normal delivery. The partner logs in, uploads the file, disconnects. The 05:00 import loads it. Nobody will see another settlement file for twenty-one days.
  • Jun 09 22:00 — Pennywhistle cuts over to a new outbound transfer platform. Their old sending job is disabled. The replacement job exists on the new platform but is left disabled pending a checklist step. That step is confirm receiving endpoint with each customer. The step is assigned to an engineer who starts two weeks of leave the next morning.
  • Jun 10 04:00 — Nothing connects to sftp.example.com as pennywhistle.
  • Jun 10 05:00 — The import job runs, finds an empty inbox, logs 0 files found, nothing to do, and exits 0. The summary email goes out: Settlement import complete.
  • Jun 11 – Jun 29 — The same, every day. Twenty green runs. Twenty identical emails. On the partner's side, a disabled job produces no failures, so their monitoring is equally content.
  • Jun 30 15:20 — Grace, in finance, starts the month-end reconciliation: bank statement against ledger. Card settlements are present through Jun 09 and absent after it, for every store.
  • Jun 30 15:45 — Ticket to IT: settlement import broken?
  • Jun 30 16:00 — Felix, the administrator on the ticket, checks the import job's history. Green every day. Inbox empty. He runs the import by hand, twice, expecting it to "fetch" something.
  • Jun 30 16:40 — He checks the SFTP server's activity log. The last login by pennywhistle was Jun 09 04:02. The partner has not connected since.
  • Jun 30 16:50 — Call to Ines, the partner's operations contact. Their dashboard shows no failures. At 17:25 she finds the disabled job and the unticked checklist step.
  • Jun 30 17:40 — The partner agrees to backfill all twenty-one files overnight.
  • Jul 01 03:00 – 03:20 — The backfill arrives. Felix, watching, starts the import at 03:05, while files are still being written.
  • Jul 01 03:30 — The import rejects nineteen files: settlement date older than three days. The first error alert of the entire incident fires, during the recovery.
  • Jul 01 09:00 – Jul 02 17:00 — Validation relaxed under a documented exception, one partially imported day reversed and reloaded, twenty-one days posted. Finance closes the month two days late.
  • Jul 04 14:00 — Postmortem, with the partner joining for the first half.

Why Every Alert Stayed Quiet

This is the heart of the story, and it deserves more care than "we should have had an alert." Each alert was correctly designed for its failure. None was designed for absence.

The server alert watched for bad logins. Failed passwords, lockouts, unknown accounts. A partner that never connects produces none of those. To the server, a quiet account and a healthy account are indistinguishable.

The import alert watched the exit code. The job did exactly what it was written to do: import every file present. When no files are present, importing all of them takes a fraction of a second and succeeds. The log said so:

Jun 09 05:00:02  import start    inbox=D:\feeds\in\pennywhistle
Jun 09 05:00:04  imported SETTLE_daily_jun08.csv   14 stores   2,318 rows   OK
Jun 10 05:00:02  import start    inbox=D:\feeds\in\pennywhistle
Jun 10 05:00:02  0 files found. Nothing to do.   exit 0
Jun 11 05:00:02  0 files found. Nothing to do.   exit 0
Jun 12 05:00:02  0 files found. Nothing to do.   exit 0

Read as a story, that log is alarming by the second day. Read by a rule that asks only "was the exit code zero?", it is twenty-one days of success. The difference is the whole incident. Why jobs fail silently catalogs this shape: the job that succeeds at doing nothing. It is the most common silent failure there is.

The partner's alert watched for failing jobs. A disabled job does not fail. Their cutover checklist was the only thing that knew the feed was not running. It was waiting for someone on leave.

The summary email never changed. Its subject line was Settlement import complete whether it had loaded two thousand rows or none. The body said 0 files, but nobody opens the four-hundredth identical message. An email that says the same thing on a good day and a bad day is not a notification. It is wallpaper. Alerting that gets read is about exactly this: the subject line must carry the number that would make someone look.

The business had the signal and was not looking. The finance dashboard showed zero card settlements for every day after Jun 09. It was a month-end tool, opened at month-end.

Remember: a failure alert can only fire when something runs and breaks. The thing that should run may never start because a job was disabled, a partner changed, or a schedule lost its trigger. Then the only monitor that can notice is one that expected something and counted the minutes since it last happened.

The Reconciliation Scramble, Including the Mistakes

Forty minutes on the wrong side

Felix's first hour went to re-running the import, because the ticket said "import broken." The import is the thing IT owns. (I have done this too. The thing you own is the thing you poke.) But the flow is a push: the partner sends, Northgate receives. Re-running the receiver cannot produce a file the sender never sent. The evidence that settled it was the server's log showing the account's last login three weeks earlier. It was one query away the whole time. Sysax Multi Server logs every login and upload with account and time, to file and optionally to a database. A query for last successful upload per partner account is the fastest answer to "did they stop, or did we?" The first question in any missing-file incident is which side owns the next move. The article on push vs pull makes that reflex automatic.

Importing while the backfill was landing

The partner's backfill wrote twenty-one files over twenty minutes, with no temporary names and no marker to say "done." Felix started the import five minutes in. It loaded four complete files and one still being written, about sixty percent of a day's rows. It archived the partial file as if finished. That day had to be reversed and reloaded. A receiver never processes a file the sender may still be writing. The sender's half of the fix is temporary names and atomic renames. The receiver's half, when the sender cannot change, is size-stability checks.

The recovery tripped the only alert

The import validated that a settlement date was no more than three days old, to catch a partner re-sending stale data by mistake. Twenty-one backfilled files failed that check, and the job finally exited non-zero. The first alert of the incident came during the fix. The validation was right to exist. What was missing was a documented catch-up. That is a mode or an exception process for the ordinary situation where a feed resumes after a gap. Catch-up after failures designs one before you need it.

Root Cause and Contributing Factors

  • No freshness check. Nothing on Northgate's side recorded that a settlement file was expected daily by 04:30. So nothing could notice its absence. This is the root; everything else made the gap longer.
  • Zero files treated as success. The import's authors wrote "nothing to do" as a normal outcome. For a daily feed, an empty inbox on a business day is a warning at least, and after two days an alarm.
  • Failure-only alerting on both sides. Server, import, and partner platform all alerted on errors. A disabled job and an absent file generate none.
  • A partner change with no notice. The partner migrated its platform without telling customers. Their own checklist step for confirming endpoints existed precisely because they knew this mattered. It stalled on one person's leave. The technical side of a swap like theirs is our platform migration mechanics series. The coordination side, telling every counterparty before the switch, is what failed here. The article they needed is partner coordination in migrations.
  • No written expectation. No SLA row said "daily by 04:30, contact Ines if absent by 06:00." Both sides assumed the other would notice.
  • A summary email with a constant subject line, and a dashboard nobody opened between month-ends.
  • No catch-up procedure, so the recovery produced the incident's only errors and an extra day of delay.

Neither Felix nor the engineer on leave appears on that list. One followed the ticket; the other followed a checklist that was correct and unfinished. The system both worked in had no way to represent "this should have happened by now."

What Was Actually Changed

The central change was an expected-arrivals register, a small CSV with one row per feed. Each row states what should arrive, where, and by when. A check runs every fifteen minutes against it. The check reads the archive folder the import moves files into. So "arrived today" means "arrived and imported today." Both are copyable.

# D:\ops\expected-feeds.csv  -- one row per feed that must arrive
feed,folder,pattern,due_by,business_days_only,contact
pennywhistle-settlement,D:\feeds\archive\pennywhistle,SETTLE_daily_*.csv,04:30,no,ines@pennywhistle.example.com
warehouse-stock,D:\feeds\archive\warehouse,STOCK_*.csv,06:15,yes,ops@warehouse.example.com
# check-expected-feeds.ps1 -- run every 15 minutes from Task Scheduler; alerts when a feed is late
$Now   = Get-Date
$Feeds = Import-Csv 'D:\ops\expected-feeds.csv'

foreach ($f in $Feeds) {
    if ($f.business_days_only -eq 'yes' -and $Now.DayOfWeek -in 'Saturday','Sunday') { continue }
    $due = Get-Date $f.due_by                       # today, at the feed's due time
    if ($Now -lt $due) { continue }                 # not late yet

    $latest = Get-ChildItem -Path $f.folder -Filter $f.pattern -File -ErrorAction SilentlyContinue |
              Sort-Object LastWriteTime -Descending | Select-Object -First 1
    $arrivedToday = $latest -and ($latest.LastWriteTime -ge $Now.Date)

    if (-not $arrivedToday) {
        $lastSeen = if ($latest) { $latest.LastWriteTime.ToString('MMM dd HH:mm') } else { 'never' }
        Send-Alert -Severity High `
            -Subject "MISSING FEED: $($f.feed) was due $($f.due_by); last seen $lastSeen" `
            -Body    "Folder $($f.folder), pattern $($f.pattern). Partner contact: $($f.contact)"
    }
}

Send-Alert is whatever reaches a person in your estate: mail to an on-call address, a ticket, a pager. Two refinements are worth adding once the basic check works. Alert once per feed per day rather than every fifteen minutes. Escalate to a second contact if the feed is still missing two hours later. If your import deletes files instead of archiving them, point the check at the server's log. The query for last successful upload by this account answers the same question.

Around the check, four smaller things changed. The import now exits with a warning when it finds zero files on a day a feed is expected. The summary email's subject line carries the counts. Settlement import: 1 file, 2,290 rows reads differently from Settlement import: 0 files, 0 rows. A working SLA was written with the partner. It sets delivery by 04:30 daily, seven days' notice of platform changes, and named contacts on both sides. It also includes an agreement that on a day with genuinely no settlements the partner sends an empty file rather than nothing. So silence always means a problem. A catch-up mode, switched on by a flag, accepts older settlement dates and logs the exception. And the backfill procedure now says: the partner uploads to a separate backfill folder with temporary names. The partner confirms completion before anyone starts the import. The register's contact column is flow ownership and contacts in miniature. An alert that names Ines beats one that names a folder.

The Lessons, and Where to Learn Each Fix

  1. Every recurring feed needs an expectation and a clock. Write down what should arrive and by when, and check the clock against it. Freshness checks and expected files is the full design, including weekends, holidays, and feeds that legitimately skip days.
  2. "Nothing to do" is a result, not a success. A job that expects input should say so loudly when it finds none. Why jobs fail silently lists the other shapes of green failure.
  3. Subject lines carry numbers or they carry nothing. Alerting that gets read makes the daily message impossible to ignore on the one day it matters.
  4. Write the expectation down with the partner. Delivery windows, change notice, contacts, and the empty-file convention belong in a working SLA. The template is partner SLAs and expectations.
  5. Silence should never be ambiguous. An empty file, a marker, or an acknowledgment turns "no data today" into a message. The article on acknowledgment patterns covers the options.
  6. Know which side moves next. Decide whether the flow is push or pull. Then check the receiving server's log for the sender's last visit. The article on reading transfer logs shows that query.
  7. Design the catch-up before the gap. Backfills are ordinary; validation that rejects them and imports that start mid-upload turn a gap into a second incident. See catch-up after failures and marker and control files.
  8. Monitor the monitor. A freshness check that stops running is a silent miss of its own. The heartbeat that closes the loop is monitoring the monitoring.

Check Your Estate

List every feed that arrives on a schedule, from partners, from internal systems, from anywhere, and answer these for each. Most estates find one or two feeds that stopped weeks ago and that nobody has missed yet. "Yet" is doing a lot of work in that sentence.

SILENT-MISS CHECKLIST  (one row per recurring inbound feed)

[ ] The feed has a row in an expected-arrivals register: pattern, folder, due time, days, contact
[ ] Something checks the clock against that row and alerts a person when the feed is late
[ ] The consuming job treats "zero files on an expected day" as a warning, and two days as an alarm
[ ] The daily summary's subject line contains the count that would change on a bad day
[ ] The partner has agreed to send an empty file or marker on genuinely empty days
[ ] A working SLA names the delivery window, change-notice period, and contacts on both sides
[ ] You can answer "when did this partner last connect?" from the server log in under a minute
[ ] A catch-up procedure exists: backfill folder, temp names or completion marker, validation exception
[ ] The freshness check itself has a heartbeat; its silence is also alarmed
[ ] Someone outside IT would notice the absence within a day — and you know who

The last line tests the business, not the technology. If the honest answer is "finance, at month-end," you have found a feed whose absence can hide for weeks.

The Version to Tell a Colleague

A partner migrated its sending platform and left the new job disabled behind a checklist step. The import found an empty inbox every morning, did nothing, and exited zero. The server saw no failed logins because nothing logged in. The summary email said "complete" as always. Every alert was a failure alert, and nothing failed. Finance found the gap at month-end, with a question mark. The fix was a register of what should arrive and by when, and a check that watches the clock against it. The job treats an empty inbox on a delivery day as a warning. The subject line has a number in it. And a written agreement says that silence always means a problem.

The mirror image of this story, a job that did far too much while every alert stayed quiet, is the looping job. For another incident where the server's log answered the key question in one query, see the leaked credential. The method for reviewing your own incidents is in running a blameless postmortem.

Frequently Asked Questions

Why didn't the import job's success email count as an alert?
Because it looked identical on good days and bad days. The subject line never changed, and people stop opening a message that has said the same thing hundreds of times. An alert must differ visibly when something is wrong. That means a count in the subject line, or no email on good days and a loud one on bad.
What is a freshness check, in one sentence?
A scheduled check that knows when a file should have arrived and alerts when the clock passes that deadline without a matching file appearing. It watches the clock rather than waiting for an error, which is why it can see an absence.
Should the partner have caught this before we did?
Both sides could have, and neither had a check for absence. The partner's monitoring watched for failing jobs; a disabled job never fails. The useful lesson is to write the expectation down on both sides, with contacts and a notice period for changes. Then either side noticing is enough.
How do I tell whether the sender stopped or our receiver broke?
Check the receiving server's activity log for the sending account's last successful login and upload. If the last visit was weeks ago, the sender stopped. If they connected today and no file was processed, the problem is yours. That one query would have saved forty minutes here.
Why send an empty file on days with no data?
So that silence is never ambiguous. If the partner always sends something, a file with headers only or a small marker, then the absence of anything is always a fault. The freshness check can then alert without a list of exceptions. Sending the file or marker also gives your import a positive record that the day was checked.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.