Home › Topics › Watch Folders › Error Design

Designing Watch Folders That Handle Failure

Anyone can build a watch folder that works when the files are good. The happy path is twenty lines: see file, claim file, process file, archive file. What separates a workflow you can trust from a demo that happens to be in production is everything that twenty-line version leaves out. There is the malformed file, the destination that stopped answering, the upload that never finishes. There is the file that fails, and fails, and fails again at two-minute intervals. Eventually someone notices the log is forty thousand lines of the same error.

Failure handling in a watch folder is not an exception path bolted on at the end. It is half the design, and the half that determines whether the workflow runs unattended for years or pages you monthly. This article — part of our watch folders and event-driven transfers series — covers that half. It explains why the error folder is a first-class part of the anatomy, and how to decide between retrying and parking. It covers how reprocessing loops start and how to make them impossible. It covers what to alert on, and how to keep the whole thing legible to the operator who inherits it.

Failure Is a Normal Input

Start with the mindset shift. A watch folder is an unattended intake open to the outside world — partners, applications, devices, colleagues. Over a long enough run, everything that can arrive will arrive. That includes a file with the wrong name, a CSV missing its header, a zip that does not open. It includes a file three hundred times the usual size, an empty file, a file that is still growing an hour later. And independent of the files, the world around the workflow fails too. The destination server goes down for patching, credentials expire, a disk fills, and a network path flaps.

Those two families — the file is bad versus the world is bad — behave completely differently. Telling them apart is the root decision of all error design. A bad file is a permanent failure. No number of retries will give the CSV its missing columns, and every retry just burns cycles and spams logs. A bad world is a transient failure: the file is fine, the moment is wrong. A retry after a sensible delay will very likely succeed. The full taxonomy, with real error messages classified, is the opening article of our retry and error handling series. For watch-folder purposes the two-way split carries almost all the weight.

One more principle before the machinery: one bad file must never stop the flow for the good files behind it. The whole point of per-file, arrival-driven processing — the case made in the hot folder pattern — is that each file has an independent fate. An error design that halts the watcher on first failure converts one bad file into an outage for every sender. Isolate the casualty, keep the line moving.

The Error Folder Is a Feature, Not an Apology

The instrument that makes isolation real is the error folder — error/ in the standard layout, sometimes called a quarantine or dead-letter area. Its job: hold files that cannot proceed, out of the pipeline's way and in front of a human's eyes. Four properties make an error folder useful rather than decorative:

  • Same volume as the inbox. The quarantine move — the act of parking a failed file — must be a rename: instant, atomic, unable to fail halfway. An error folder on another disk or share turns your failure path into a slow copy that can itself fail. That is exactly the wrong moment for new problems.
  • Invisible to producers. Senders write to in/ and nowhere else. A partner who can see — or worse, write into — your error area will "helpfully" delete, resend, or edit evidence you need.
  • Carries context, not just corpses. A bare file in an error folder is a puzzle. Park a small companion note beside every casualty saying what happened, so diagnosis starts from facts instead of archaeology.
  • Watched by a person, alerted by the system. An unmonitored error folder is a landfill with better intentions. Every arrival triggers a notification to someone whose job includes acting on it.

The context note deserves a concrete shape. A plain-text sidecar, named after the file it describes, answers the diagnostic questions in order:

error\acme_orders_YYYYMMDD_002.csv
error\acme_orders_YYYYMMDD_002.csv.why.txt
--------------------------------------------------
parked   : Mar 14 02:11
flow     : acme-orders intake
stage    : validate (after claim, before transfer)
error    : header mismatch - expected 9 columns, found 7
attempts : 1 (permanent failure - no retry)
file     : 48,213 bytes, arrived Mar 14 02:10 via sftp drop
next     : confirm layout with sender; resend under new name
owner    : transfers on-call

If this pattern looks familiar, it should. It is the same quarantine discipline used when malware scanning intercepts a file. Our guide to quarantine workflow design explores the shared ideas — isolation, evidence, controlled release — in the security context. A transfer error folder is the operational cousin: same moves, lower stakes, identical need for discipline.

Retry or Park: Making the Decision Explicit

Every failure the watcher meets ends in one of two actions: retry (try again later, bounded) or park (quarantine move now, human follows up). The classification from earlier drives the choice. Writing the mapping down — rather than leaving it implicit in code — makes the workflow's behavior predictable during an incident. The table below is a starting matrix you can adapt per flow:

Failure Class Response Park when
Name violates the intake contract Permanent Park immediately, notify sender's contact First occurrence
Validation fails (structure, counts, archive integrity) Permanent Park immediately with the validator's message First occurrence
Destination unreachable / timeout Transient Retry with increasing delays Retry budget spent (e.g. 5 tries / 2 hours)
Authentication rejected Permanent until fixed One confirming retry, then park the flow's files and page — credentials do not heal themselves Second failure
Destination disk full Transient, slow Long-interval retries, alert early — healing needs a human elsewhere Budget spent
File never settles / never stops growing Suspect Leave in inbox, alert on age (see stuck-file rule) Age threshold, park with note
Unexplained processing crash Unknown Treat as transient, small budget (2–3 tries) Budget spent — likely a poison file

Three rules travel with the table. First, retries are always bounded. A retry budget (attempts, or a time window, or both) is decided in advance. Exhausting it converts the failure to a park; retry-forever is how flows wedge silently. Second, retry delays grow — immediate, then minutes, then longer, so a struggling destination is not hammered. The timing craft, including backoff and jitter, is covered in the retry and error handling series. Third, the attempt count travels with the file, in a sidecar or a ledger, not in the watcher's memory. A watcher restart must not reset every file's count to zero. If you build on a folder-monitoring tool rather than raw script, much of this is configuration. Sysax FTP Automation, for example, pairs its folder monitoring with built-in retry and error handling. So bounded retries and failure actions are settings on the task instead of logic you maintain.

The Reprocessing Loop, and How to Make It Impossible

The signature failure of amateur watch folders is the infinite reprocessing loop. A file fails, remains where the watcher looks, is found again on the next scan, fails identically, and repeats — every thirty seconds, all weekend. The log grows monstrous, and real alerts drown in the noise. If each attempt does partial work (a half-upload, a notification), the damage multiplies. Loops start in a handful of well-known ways:

  • Failure leaves the file in the inbox. The watcher processes in place, hits the error, and moves nothing. So the next sweep finds the same file with no memory of the disaster. This is the default behavior of the twenty-line demo.
  • The error folder lives inside the watched path. A recursive watcher pointed at the intake root sees error/ as more arrivals. The file is parked and immediately rediscovered — the loop now includes the quarantine itself.
  • Retries without a counter. Each attempt is genuinely the first as far as the code knows. Bounded retry needs remembered attempts; amnesia makes every budget infinite.
  • A human short-circuits the quarantine. Someone drags the file from error/ back to in/ without fixing the cause. The loop is manual, but it is still a loop.

The countermeasures are structural, which is what makes them reliable. First, claim before processing. The file leaves in/ before any work starts, so failure strands it in work/, not back in the scan path. Next, park on failure, always. The quarantine move is the error handler's last act, and it is a move, so the file cannot be in two rooms. Also, watch the inbox only, never subfolders. And make reprocessing a deliberate act — a documented re-entry step a human performs after fixing the cause. Reprocessing should not be a side effect of where a file happens to sit. Files that defeat even bounded retries — failing on every attempt for reasons nobody has diagnosed yet — are poison files. The dead-letter thinking for handling them systematically is part of the retry and error handling pillar.

Remember: a failed file must always end the attempt in a different folder than it started. If failure can ever leave a file where the watcher will find it again unchanged, you have built a loop. You are waiting for a bad file to start it.

Alerting: Parked, Stuck, and Silent

An error design that quietly files casualties away is only half done — the other half is telling someone. Three conditions cover what a watch folder needs to say:

Parked. A file entering error/ is, by definition, something automation could not fix. It should generate one alert, at the moment of parking. The alert should carry the same facts as the sidecar note: flow, filename, stage, error, and the first diagnostic step. Send one alert per park. One per retry attempt trains people to delete the flow's alerts unread. Using only a daily digest turns a two-minute fix into a next-day discovery.

Stuck. Some failures never raise an error. A file sits in in/ because it never passes its settle check, or in work/ because the watcher died mid-job. Nothing failed, so nothing alerted — the file is simply stuck. The countermeasure is an age sweep: a periodic check. It can share the reconciliation sweep from the detection article. The age sweep alerts on any file older than a threshold in any transit room. Ten minutes in work/ when processing normally takes forty seconds is a finding, whatever the logs claim.

Silent. The failure no folder can show you: nothing arrived at all, because the sender's side broke. Absence detection — expected-file deadlines, grace windows — is its own discipline, covered in the transfer job monitoring series. Here it is enough to note that an empty inbox means either "all is well" or "nobody can reach us." Only a freshness check can tell the two apart.

Wire the alerts to wherever your team actually looks. Email is the traditional channel for file-flow automation, and tools in this space treat it as standard equipment. Sysax FTP Automation, for instance, can send email notifications as part of its task error handling. That covers the parked case without extra plumbing. Whatever the channel, log every failure event as well. The alert is for now, the log line is for the postmortem, and the two should agree. Our guides on what to log and alerts from transfer logs set the standard for both halves.

The Operator's View: Keep It Obvious

The final test of error design is not technical; it is human. Can the person on duty — possibly hired after you built this — answer three questions in under a minute? Is anything stuck? Is anything parked? When did this flow last succeed?

The four-room layout answers the first two with a directory listing. Having in/ and work/ empty means nothing stuck. The folder error/ lists the open incidents, each with its sidecar note. The newest timestamp in done/ answers the third. This glanceability is worth protecting with a few habits:

  • Casualties keep their names. Park acme_orders_YYYYMMDD_002.csv as itself, with the story in the sidecar. Do not rename it into failed_0047.dat. That destroys the one identifier the sender and the logs share.
  • The error folder gets emptied by resolution, not by cleanup. Every departure from error/ is either a fixed-and-reprocessed file or a documented write-off. A quarterly "clear out the error folder" purge means the folder stopped being information months ago.
  • Reprocessing has a written procedure. Fix the cause, then re-enter the file through the front door. Move it back into in/ (or a dedicated re-entry point) and watch it all the way through. Because reruns happen, processing should be safe to repeat; the duplicate detection and idempotency series covers making that true.
  • Startup policy is written down. After a crash, anything in work/ is suspect. The runbook says whether the watcher parks it with a note (safe default) or re-queues it (fine when processing is idempotent). Deciding at 3 a.m. is design debt coming due.

Everything above compounds into a quiet operational virtue: failures become boring. A bad file arrives, gets parked with its story, raises one alert, and waits; the flow keeps serving everyone else; the fix follows the runbook. That boringness is the entire goal.

The Design in Summary

Watch-folder error design reduces to five commitments. Classify failures as file-bad or world-bad, because the first should never be retried and the second usually should. Give the workflow a real error folder — same volume, producer-invisible, context beside every casualty, a human notified on arrival. Bound every retry with a budget and growing delays, and let exhaustion convert to a park. Make loops structurally impossible: claim before processing, park on failure, watch only the inbox, reprocess only deliberately. And keep the operator's view legible enough that the directory tree itself reports the system's health.

With intake covered by the arrival contract and failure covered here, you have the two hard halves of the pattern. The series capstone, building a watch folder workflow end to end, assembles them into a complete working flow with folder trees, log lines, and every failure path wired. For the deeper failure craft — backoff timing, poison-file handling, recovery design — continue into the retry and error handling series.

Frequently Asked Questions

Should a failed file stay in the inbox so it gets retried automatically?
No — that is the recipe for an infinite reprocessing loop. Retries should be deliberate. The file is claimed out of the inbox first, then retried from the working area on a bounded schedule. It is parked in the error folder when the budget runs out. Failure must always leave the file somewhere the next scan will not blindly rediscover it.
How many times should a transient failure be retried?
Decide a budget in advance — commonly three to five attempts with growing delays, or a time window like two hours. Park the file when the budget is spent. The exact numbers matter less than having them written down and bounded. Permanent failures, like a file that fails validation, should not be retried at all.
What is a poison file?
A file that fails every processing attempt, usually for a reason nobody has diagnosed yet — an encoding surprise, a crash-triggering edge case. Bounded retries exist largely to stop poison files from clogging the pipeline. Once the budget is spent, they are parked with their error context for a human postmortem.
Can the error folder be a subfolder of the watched folder?
It can sit under the same flow root (like intake\acme\error), but the watcher must scan only the inbox itself, never subfolders. If a recursive watcher can see the error folder, every parked file is instantly rediscovered as a new arrival. You have built a quarantine-shaped loop.
Who should be notified when a file is parked?
Someone whose job includes acting on it — a team alias or on-call rotation, not an individual's inbox that goes quiet during vacations. For contract violations like bad names or failed validation, notify the sender's contact too, since the fix is usually on their side.
Is it safe to just move a parked file back into the inbox?
Only after the cause is fixed — otherwise it will fail identically and land back in quarantine, or loop. Follow a written reprocessing step: fix the cause, re-enter the file through the normal intake path, and watch it complete. If reruns can ever double-process, add idempotency protection first.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.