Error Messages Worth Logging (and Acting On)
Transfer failed. That is the whole message. It is 03:10, the file did not arrive, and the person squinting at the screen has never seen this job before. That is because the person who wrote it is asleep, on leave, or at another company. Everything that reader will ever know about the failure is what the job wrote down at the moment it happened. An error message is not decoration on a failure; it is the only witness statement the failure leaves behind, and this one says "something happened."
Most transfer jobs write terrible witness statements. Failed doing what? To where? Which file? What did the server say? Was it retried? This article is about writing better ones. It covers capturing exit codes and client output properly, and enriching bare errors with job, file, and host context. It covers deduplicating the alert storm a repeating failure creates. It gives you a concrete style guide, bad line versus good line, you can apply to any script or tool today. It is part of our Retry Logic and Error Handling series. Earlier articles decided what to do about failures; this one makes sure they leave evidence worth having.
Five Questions Every Error Message Must Answer
Strip away formats and tooling and a useful error record answers five questions:
- What operation failed? Uploading, downloading, connecting, renaming — and the object involved: which file, which host, which folder.
- When? A timestamp precise enough to line up with other systems' logs, with the time zone unambiguous.
- What exactly went wrong? The verbatim evidence: the server's reply, the exit code, the exception text — not a paraphrase.
- What did the job do about it? Retried three times? Quarantined the file? Gave up? The response is part of the story.
- What should a human do now? Even one clause — "check the account is not locked" — turns a report into an instruction.
Hold any log line you have ever written against those five questions and the gap is usually obvious. ERROR: upload failed answers one of five. The rest of this article is the machinery for answering all of them without heroics. Heroics at three in the morning are usually a logging failure wearing a cape.
Capture the Evidence: Exit Codes and Client Output
The raw material comes from two places, and scripts routinely lose both. Not on purpose. By default.
The first is the exit code — the number a process leaves behind, zero for success. The classic bug is checking it too late. In a shell script, $? holds the exit code of the most recent command. So an innocent echo between the transfer and the check overwrites the evidence. Capture it into a named variable on the very next line, always. PowerShell has the same trap with $LASTEXITCODE for external commands. In Python, client libraries raise exceptions instead — the equivalent discipline is a try/except that records the exception's actual text rather than swallowing it.
The second is error output. Command-line clients write their complaints to stderr — the error stream, separate from normal output. A script that redirects only stdout throws the diagnosis away. The pattern that keeps everything is to send stderr to a per-run log file and, on failure, pull the last few lines into the summary message:
RUNLOG="/var/log/flows/invoices-out/run_YYYYMMDD_HHMMSS.log" sftp -b upload.batch acct@sftp.partner.example >>"$RUNLOG" 2>&1 rc=$? # capture immediately - nothing in between if [ $rc -ne 0 ]; then detail=$(tail -n 3 "$RUNLOG" | tr '\n' ' | ') log_error "rc=$rc detail=\"$detail\"" fi
Two refinements pay for themselves. Run clients with their verbose flag always, writing to the run log. Verbosity costs nothing when the run succeeds and is priceless when it does not. You cannot retroactively add -v to last night. And record silence explicitly: a hung transfer killed by a watchdog produces no error text at all, so the wrapper must write the absence down. The message "no output; killed after 120 seconds" is a real diagnosis. Per the classification article, it is a very different one from a crisp server rejection.
Exit Codes Are a Language: Speak It Deliberately
Capturing the client's exit code is half the job. The other half is choosing what your wrapper exits with, because that number is the entire interface between the job and its scheduler. A wrapper that exits zero no matter what tells the scheduler every run succeeded. That is a depressingly common pattern, usually caused by a final logging command succeeding and setting the status. No amount of beautiful log text repairs that lie. The scheduler acts on the number, not the prose. No scheduler has ever been moved by prose.
Treat your exit codes as a tiny published vocabulary, kept consistent across every job you own. A workable set:
- 0 — full success: everything delivered and verified.
- 1 — partial: the run finished but some items failed; details in the log. (The batch-accounting side of this is covered in partial failures in multi-file jobs.)
- 2 — total failure after retries: transient trouble that outlived the budget.
- 3 — permanent failure: auth, missing path, security warning — no retry was attempted, human needed.
- 4 — configuration error: the job could not even start sensibly (missing credential file, bad parameter).
The exact numbers matter less than the consistency. Once monitoring knows that 3 means "do not just re-run this," it can route the failure differently than a 2. A re-run might genuinely cure the latter. In a shell script this is exit 3; in PowerShell, exit 3 at script level; in Python, sys.exit(3). Document the vocabulary in the script header where the next maintainer will trip over it, and resist inventing per-job dialects. Five codes, used identically everywhere, beat twenty codes nobody remembers.
Enrich with Context: Job, File, Host, Attempt
The client's error text says what the protocol saw. It does not say which flow this was, which file was in flight, which machine ran the job, or whether this was attempt one or attempt five. That context lives only in the wrapper — so the wrapper must attach it at write time, because nobody can reconstruct it later. The workhorse format is a single line of key=value pairs:
Mar 14 02:41:07 level=ERROR flow=invoices-out host=app01 remote=sftp.partner.example file=inv_YYYYMMDD_0004.pdf attempt=3/3 rc=1 detail="remote open: 553 File name not allowed" action=quarantined runlog=run_YYYYMMDD_HHMMSS.log
Everything a responder needs is on that line, and — just as important — every value is searchable. grep flow=invoices-out finds this flow's history; grep "553" finds every name rejection across all flows. attempt=3/3 distinguishes a final failure from a stumble that recovered. Long diagnostic material — full verbose output, stack traces — does not belong on the summary line. It stays in the run log, and the summary line points to it by name. The one-line-summary-plus-detail-file split is the same discipline our logging series recommends generally. The article on what to log covers the full field list. And reading transfer logs is its mirror image, written for the person searching.
The quiet hero among those fields is the run id — a value generated once at the top of the run. A timestamp plus the process id is plenty. The id is stamped onto everything the run produces: every summary line, the run log's file name, the manifest, the notification. It is the thread that ties the pieces back together. When a partner asks about one file three weeks from now, the search goes like this. Find the file's error line, read its run id, open that run's log and manifest, and you are done. Without the id, the same investigation is an exercise in matching timestamps across files and guessing which of Tuesday's four runs was the guilty one. That is twenty minutes of forensics to replace one string you could have written down for free. I have done that forensics. I do not recommend it.
The Style Guide: Bad Line vs Good Line
Rules stick better next to examples. Each pair below is the same failure written twice — as jobs usually write it, and as the five questions demand:
| Bad line | Good line |
|---|---|
Transfer failed |
upload inv_YYYYMMDD_0004.pdf to sftp.partner.example failed: 553 File name not allowed (attempt 3/3, quarantined) |
Could not connect |
connect to sftp.partner.example:22 timed out after 30s (attempt 1/6, retrying in 34s) |
Disk problem on server |
local write failed: no space left on /data (194 MB free, file needs 1.2 GB) - job stopped, no retry |
Login error, will retry |
auth rejected for acct@sftp.partner.example - NOT retrying (lockout risk); check credential rotation |
The rules those good lines follow, stated as a checklist you can apply in code review:
- Name the operation and the object. A verb and a file: "upload X to Y," never just "failed."
- Quote the verbatim evidence. The server's exact reply —
553 File name not allowed— not your paraphrase of it. Paraphrases destroy the reply code, and the reply code is the classification. - Prefer values to adjectives. "194 MB free, needs 1.2 GB" ages better than "disk almost full." Numbers let the reader judge severity; adjectives make them guess.
- State what the job did. "Retrying in 34s," "quarantined," "stopped, no retry" — the response is half the story, and it tells the reader whether anything is still owed.
- Make timestamps unambiguous. Pick one time zone for all job logging — UTC or the server's local zone — and use it everywhere. An incident investigated across systems logging in different zones wastes its first half hour reconciling clocks.
- Never log secrets. Passwords, keys, and full connection strings do not belong in logs — scrub them even from verbose client output, which sometimes echoes more than you expect.
- Keep the shape stable. The same failure should produce the same message text every time, because stable text is what searching, counting, and deduplication all key on.
Remember: write the message for a reader who has never seen this job, at the worst hour, with the original author unavailable. If that reader can pick a next step from the message alone, the message is good.
A Small Library of Message Templates
The fastest way to make every job speak this dialect is to stop composing messages ad hoc. Keep a small library of templates — one per failure class, using the classes from the classification article. The shape is fixed and only the values change:
AUTH_FAIL auth rejected for {account}@{remote} - NOT retrying (lockout risk); check credential rotation
TIMEOUT connect to {remote}:{port} timed out after {limit}s (attempt {n}/{max}, retrying in {delay}s)
NAME_REJECT upload {file} to {remote} failed: {server_reply} (attempt {n}/{max}, quarantined)
DISK_LOCAL local write failed: no space left on {path} ({free} free, needs {size}) - job stopped, no retry
UNKNOWN {operation} {file} to {remote} failed rc={rc}: {detail} (attempt {n}/{max})
The template name itself goes into the logged line, and that one habit compounds quietly. Jobs written years apart produce identical text for identical problems, so searching and counting work across the whole estate. grep NAME_REJECT finds every name rejection ever, regardless of which script hit it. New flows inherit good messages for free instead of re-learning the five questions. Deduplication gets its stable key without extra work, because the template name plus the object is the key. And code review gains a checklist question — "which template does this failure path use?" — that catches missing error handling before it ships. Five templates cover most transfer estates; add a sixth when reality invents a class you did not have, not before. Reality is reliable about this.
One Problem, One Alert: Deduplication
A repeating failure writes the same line every run — which is correct for the log and disastrous for the inbox. Forty-seven identical emails by morning do not communicate forty-seven times as much as one. They communicate less, because somewhere around the fifth the reader stops opening them. The different, genuinely new alert that arrives at 04:00 dies unread in the pile. Deduplication is what keeps notifications meaningful, and it rests on three techniques. None of them involves a bigger inbox.
Alert on state changes, not states. The first failure after a success is news; the eleventh consecutive failure is not. Track the previous outcome per flow and notify when it flips. Include the flip back, because "recovered" is the second-most useful message a pipeline can send. A tiny state file per flow is enough to implement this in a script.
Rate-limit repeats with a counter. When the same condition persists, escalate the summary instead of repeating the alert. In that case, send one message per hour at most per deduplication key, carrying "occurred 47 times since Mar 14 02:10." The natural key is flow plus error class plus object — invoices-out / 553 / inv_YYYYMMDD_0004.pdf — so that two different problems in one flow still alert separately.
Escalate on age. Deduplication must never become suppression. A condition that has been failing quietly for six hours deserves a louder message than it got at minute one. A daily sweep of everything currently red is the backstop that catches alerts lost in transit. The craft of keeping notifications honest at scale — severity tiers, digests, who gets woken — is its own series: see alerts from transfer logs and the broader transfer job monitoring pillar.
Where Messages Should Land
One failure, three audiences, three destinations. The mistake to avoid is sending the same text to all three.
- The run log gets everything: verbose client output, every attempt, every timing. Its audience is deep investigation, and its friend is the run id that summary lines point to.
- The job summary log gets one enriched line per event — the
key=valueline from earlier. Its audience is searching and counting: how often, which flows, since when. - The notification channel gets state changes only, phrased as the good lines above. Its audience is a human deciding whether to get up.
On Windows estates, some of this comes assembled. Sysax FTP Automation tasks include retry and error handling and can send an email notification when a task fails. That is the notification layer, wired to the job outcome without hand-rolled state files. On the receiving side, Sysax Multi Server logs all activity to both a file and a database. That quietly solves the "searchable summary" layer for the server's half of every conversation. Failure analysis becomes a query — every refused login this week, every session from that partner's address — rather than a grep across rotated text files. Client-side wrapper logs plus server-side activity records, joined on timestamp and file name, reconstruct almost any incident. Getting both streams into one searchable place is the subject of centralizing logs.
Close the Loop: The Message Feeds the Machine
Well-written errors are not only for humans. The stable, quoted server reply is what your classifier matches on to choose retry versus stop. The action=quarantined field is what the dead-letter sidecar copies, per the poison-file pattern. The attempt counter in the message is the same counter the batch manifest reports at the end of the run. Sloppy messages break all of it at once — a paraphrased reply defeats the classifier, an unstable shape defeats deduplication, a missing file name defeats the manifest. Message discipline is not polish on top of error handling; it is load-bearing.
Start With One Flow
Kestrel Payroll had a flow whose only failure message was Transfer failed. The help desk had learned to route every one of its tickets straight to the one engineer who could decode it. Each ticket cost that engineer half an hour. The work was to open the run log, guess which of the night's runs was the guilty one, find the server reply, and explain it. The retrofit took one afternoon, most of it spent on the immediate exit-code save and the quoted server reply. The next month's tickets for that flow averaged a few minutes, and most closed from the alert text alone. That was because 553 File name not allowed told the desk which partner to call before anyone opened a log. The engineer got the half-hours back. The flow had not become more reliable; it had become more honest.
Retrofit this in an afternoon. Pick your most troublesome flow and add the capture pattern (verbose run log, immediate exit-code save, tail-on-failure). Switch its summary line to key=value with the five questions answered, and add first-failure/recovery notifications. The next incident will pay you back the afternoon. Then read transient vs permanent failures if you have not, because classification is what turns captured errors into decisions. Read designing transfer jobs that recover themselves, where good records become the raw material of self-healing runs. The next witness statement will name the file, the host, and the hour.
Frequently Asked Questions
What should every transfer error message include?
How do I capture a command's error output in a script?
2>&1 into the same log as stdout. Save the exit code into a variable on the very next line, since the next command overwrites it. On failure, pull the log's last lines into your summary message.Why did my log miss the error even though the job failed?
How do I stop getting fifty alerts for the same failure?
What exit code should my wrapper script return?
Is it safe to log full connection details?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
