Logging Patterns for Bash Transfer Jobs
Every scheduled transfer job eventually faces its judgment morning. A partner says a file never arrived, a downstream system processed something twice, and someone asks the only question that matters — what actually happened last night? A job with good logs answers in ninety seconds. A job without them turns the question into an afternoon of guesswork, re-running things by hand and hoping the problem reproduces.
Interactive commands do not need logs because you are the log — you watch the output scroll by. An unattended job has no watcher, so whatever it writes down is the entire record of the run. That flips logging from a nicety into a design requirement, on par with the transfer itself.
This article, part of our Bash & Cron transfer automation series, builds transfer-job logging from the ground up. It covers a timestamped logging function worth copying, capturing what the transfer tools themselves say, and choosing between log files and syslog. It also covers recording outcomes a monitor can parse, and rotation so the logs never become their own incident.
What a Transfer Log Must Be Able to Answer
Design logs backward from the questions they will be asked. For a transfer job, the recurring ones are concrete: Did the job run at all last night, and when did it start and finish? Which files did it move — names, sizes, how many? Did it succeed, and if not, what failed and with what error? What did it connect to, and as which account? Was anything skipped, retried, or left behind?
Every pattern in this article exists to make one of those questions answerable by reading, not by re-running. Notice what the list implies: a log line per run is not enough (it cannot name files). Raw tool output alone is not enough either (it cannot say what the job decided — what it skipped, what it concluded). You need both the job's narration and the tools' testimony, stitched into one timeline. What to record is also an estate-level policy question with audit implications. Our guide to what to log covers that layer. But none of the policy matters until the script can write a decent record at all, so that is where we start.
A log() Function Worth Copying
The foundation is a small function that gives every line the same shape: timestamp, job name, process ID, severity, message.
readonly JOB_NAME="push-orders"
readonly LOG_FILE="/var/log/transfer/${JOB_NAME}.log"
log() { # usage: log LEVEL message...
local level="$1"; shift
printf '%s %s[%d] %s: %s\n' \
"$(date '+%b %d %H:%M:%S')" "$JOB_NAME" "$$" "$level" "$*" >> "$LOG_FILE"
}
die() {
log ERROR "$*"
exit 1
}
Used as log INFO "run started" and log ERROR "upload failed", it produces lines like these:
Mar 14 02:17:01 push-orders[4711] INFO: run started Mar 14 02:17:04 push-orders[4711] INFO: uploading orders.csv (18244 bytes) Mar 14 02:17:09 push-orders[4711] INFO: run finished ok
Each piece earns its place. The timestamp on every line is non-negotiable: durations, hangs, and "which run was this?" all come from timestamps. The format above matches classic syslog style so the lines feel native next to system logs. (Teams that grep across long histories often prefer a fully numeric, big-endian stamp instead — the kind built from %Y%m%d-style tokens that sorts alphabetically into time order. The tradeoffs live in our file naming and datestamping series.) The job name makes lines self-identifying once logs are aggregated. The process ID ($$) separates interleaved lines if two runs ever share a file — the overlap scenario from the locking article. The process ID also lets you match a log to a still-running process. The level gives you something to filter on: grep ' ERROR: ' is the fastest incident triage there is.
Two implementation details are deliberate. printf, not echo: echo's handling of options and backslashes varies between shells and modes. That means a message beginning with -n or containing \t can silently change shape. printf with an explicit format never does. And we use $* rather than "$@" in the message slot. Here we want all remaining arguments joined into one string with spaces — the one common case where $* is the right choice. The log directory itself should be created once, owned by the service account that runs the job, with group-readable permissions — not world-readable, for reasons the "what not to log" section makes plain.
One Timeline, Not Three: Capturing Tool Output
Your script narrates, but rsync, sftp, and curl also talk — progress, statistics, and above all error messages. A run's record is incomplete unless their words land in the same timeline as yours. Three patterns, in increasing order of polish:
Per-command capture appends each tool's output to the log where it runs: rsync ... >> "$LOG_FILE" 2>&1. The 2>&1 sends stderr — where error messages travel — to the same place as stdout. Simple, explicit, and the pattern used in the fundamentals article's skeleton. Its weakness: tool lines carry no timestamps or job prefix, so bracket each command with log lines to keep the timeline readable.
Whole-script capture does it once at the top: exec >>"$LOG_FILE" 2>&1 redirects the script's own stdout and stderr for the rest of the run. That catches even output you forgot existed — a stray diagnostic from a helper, bash's own error messages. Belt and braces; the same tradeoff about unstamped lines applies.
Stamped capture pipes tool output through the log function so every line gets the full prefix:
rc=0
rsync -a --stats -e "ssh -i $SSH_KEY -o IdentitiesOnly=yes" \
"$SRC_DIR"/ "$REMOTE" 2>&1 |
while IFS= read -r line; do log RSYNC "$line"; done || rc=$?
log INFO "rsync exit status $rc"
The pieces: 2>&1 merges the streams before the pipe. IFS= read -r reads each line without trimming whitespace or eating backslashes. And the exit-status capture deserves a careful look, because pipelines are where status gets lost. Normally a pipeline reports only its last command's status — the harmless read loop — so a failed rsync would vanish. With set -o pipefail active (as strict mode makes it), the pipeline instead reports the failure. Then || rc=$? catches it into a variable without tripping set -e. When you need the status of each stage of a longer pipeline, bash keeps them in the PIPESTATUS array. But read it on the very next line, because every subsequent command overwrites it. And suspend -e around the pipeline (set +e … set -e) if you want to dissect a failure rather than die on it. One more classic to know: the while loop at the end of a pipe runs in a subshell. So variables it sets — a line counter, say — evaporate when the loop ends. Count in the parent, or count afterward with grep -c on the log.
Log File or Syslog? Use Each for What It's Good At
Everything so far wrote to a private log file. Linux also offers syslog — the system-wide logging service — and the logger command is the one-line bridge from any script into it:
logger -t push-orders -p local0.info "result=ok files=3 duration=8s" logger -t push-orders -p local0.err "result=fail rc=23 see /var/log/transfer/push-orders.log"
-t sets the tag — the source name the entry is filed under. -p sets facility and severity. The facility is a routing category (the local0 through local7 facilities are reserved for site use, ideal for jobs like these). The severity (info, warning, err) is how loudly it matters. The entry lands in the system log or journal with timestamps added for you and rotation already handled. It also gets — the real prize — a straight path into whatever central log collection your organization runs, a topic our log centralization guide takes further.
So which one? Both, for different jobs. Syslog is wrong for bulk detail — a hundred lines of per-file rsync output does not belong in the system journal. And a private file is wrong for visibility, because nothing else watches it. The pattern that works: the file gets the detail, syslog gets the verdict. Every run appends its full story to the job's own log, then emits exactly one syslog line summarizing the outcome. That one line is what dashboards count and log-based alerting triggers on, and its text tells the responder where the detail lives.
Recording Outcomes: The Summary Line
The single most valuable line a transfer job writes is its last one — the machine-parseable summary. Give it a fixed vocabulary of key=value pairs and never freestyle it:
Mar 14 02:17:09 push-orders[4711] INFO: summary result=ok files=3 bytes=52114 duration=8s Mar 15 02:17:44 push-orders[5290] ERROR: summary result=fail rc=23 files=1 duration=43s
Key=value survives every tool you will ever point at it — grep result=fail, an awk one-liner for average duration, a dashboard fed by a log shipper. The duration comes free from bash's SECONDS variable, which counts seconds since the script started. A duration that creeps up week over week is the early warning that a job is heading toward the overlap problem, visible long before anything fails. And the summary must appear on failure too — a log that simply stops mid-run answers nothing. The reliable way is an exit trap, which runs no matter how the script ends:
finish() {
local rc=$?
if [ "$rc" -eq 0 ]; then
log INFO "summary result=ok duration=${SECONDS}s"
logger -t "$JOB_NAME" -p local0.info "result=ok duration=${SECONDS}s"
else
log ERROR "summary result=fail rc=$rc duration=${SECONDS}s"
logger -t "$JOB_NAME" -p local0.err "result=fail rc=$rc see $LOG_FILE"
fi
}
trap finish EXIT
The first line banks the script's exit status before anything in the trap can disturb it. The mechanics of trap, and why EXIT is the right hook, are the centerpiece of the error-handling article. One caution the summary line cannot fix alone: a job that stops being scheduled writes no failure line at all. Pair result-line alerting with a freshness check that notices silence — the "expected a success line by 06:00 and saw none" alarm — which is the founding idea of our transfer job monitoring series.
Remember also that your log is only half of any dispute. The server on the other end keeps its own record, and correlating the two settles questions neither side can settle alone. Your log says the upload returned success at 02:17; does the server's log show the file landing? Where the far end is a Windows host running Sysax Multi Server, every session and transfer is written to its activity logs — to file and to a database. So the server-side half of the timeline is a filtered query away, matched against yours by account and timestamp.
Rotation: Logs That Don't Eat the Disk
An append-only log grows forever, and a transfer log that fills the disk eventually breaks the transfers it describes. The standard answer is logrotate, which ships with virtually every Linux distribution and takes a drop-in config per log family:
# /etc/logrotate.d/transfer-jobs
/var/log/transfer/*.log {
weekly
rotate 8
compress
delaycompress
missingok
notifempty
create 0640 xfer xfer
}
Line by line: rotate weekly, keep 8 generations (about two months of history). Use compress for old ones but delaycompress the newest so it stays instantly greppable. Don't complain if a log is missing or skip-rotate one that is empty. After rotating, create a fresh file owned by the service account with group-read permissions. Retention is a policy decision, not a technical one: keep logs as long as the questions keep coming — operational questions for days or weeks, audit questions often much longer.
One subtlety worth understanding rather than memorizing: rotation works by renaming the file, and any process still holding it open keeps writing into the renamed file. Jobs built on the log() function are immune — each call opens, appends, and closes, so the next line always lands in the current file. A job using whole-script exec redirection holds the file open for the run's duration. If rotation fires mid-run, its remaining lines follow the renamed file, which is untidy but loses nothing. Scheduling rotation away from job hours, or using per-run log files — run- plus a datestamp in the name, cleaned up with find ... -mtime +30 -delete — sidesteps it entirely. Per-run files trade easy single-run inspection against harder cross-run grepping; either choice is respectable, but make it deliberately.
What Not to Log
A log is a copy of information, with all the obligations copies carry. Three rules keep transfer logs from becoming a liability:
- No credentials, ever. Never log passwords, passphrases, or key material. And design jobs so there is nothing to leak: key-based authentication means no password exists in the job at all. Watch the sneaky path: URLs of the form
ftp://user:password@host/put a secret inside a "harmless" connection string; log the host and account, never the full URL. - Log decisions about data, not data. Record file names, sizes, counts, and checksums — not file contents. The moment payload rows appear in a log, the log inherits the payload's sensitivity, retention rules, and breach implications.
- Protect the log like it matters, because it does. Filenames, partners, schedules, and volumes sketch your business rhythms. Keep the log directory readable by the service account and the operations group only — which is exactly what the
create 0640 xfer xferline above enforces after every rotation.
Remember: the cheapest time to keep a secret out of a log is before it exists. If a job authenticates with keys and receives no sensitive arguments, there is nothing to redact on the worst day — no cleanup, no rotation of leaked passwords, no awkward disclosure conversation.
The Assembled Pattern
Everything above, in one runnable shape — strict mode and structure from the fundamentals article, stamped capture, an exit trap writing the summary, and a syslog verdict:
#!/usr/bin/env bash
set -euo pipefail
readonly JOB_NAME="mirror-reports"
readonly SRC_DIR="/data/reports/outgoing"
readonly REMOTE="reports@files.partner.example:incoming/"
readonly SSH_KEY="/etc/transfer/keys/reports_ed25519"
readonly LOG_FILE="/var/log/transfer/${JOB_NAME}.log"
log() {
local level="$1"; shift
printf '%s %s[%d] %s: %s\n' \
"$(date '+%b %d %H:%M:%S')" "$JOB_NAME" "$$" "$level" "$*" >> "$LOG_FILE"
}
finish() {
local rc=$?
if [ "$rc" -eq 0 ]; then
log INFO "summary result=ok duration=${SECONDS}s"
logger -t "$JOB_NAME" -p local0.info "result=ok duration=${SECONDS}s"
else
log ERROR "summary result=fail rc=$rc duration=${SECONDS}s"
logger -t "$JOB_NAME" -p local0.err "result=fail rc=$rc see $LOG_FILE"
fi
}
trap finish EXIT
log INFO "run started src=$SRC_DIR dest=$REMOTE"
rc=0
rsync -a --stats -e "ssh -i $SSH_KEY -o IdentitiesOnly=yes" \
"$SRC_DIR"/ "$REMOTE" 2>&1 |
while IFS= read -r line; do log RSYNC "$line"; done || rc=$?
if [ "$rc" -ne 0 ]; then
log ERROR "rsync failed rc=$rc"
exit 1
fi
log INFO "rsync completed"
Trace both endings. Success: rsync's statistics land as stamped RSYNC lines, the script reaches its natural end with status 0, and the trap writes matching ok-verdicts to file and syslog. Failure: pipefail surfaces rsync's exit code into rc, the script logs the specifics and exits 1, and the trap still fires. The summary line exists either way, which is the whole point. This is the Linux-side, hand-built version of what packaged automation provides out of the box. On the Windows half of a mixed estate, Sysax FTP Automation pairs its scheduled transfer tasks with email notifications, so the "how do I hear about failures" layer is configuration there rather than code.
Logs Are the Job's Memory
A transfer job's log is the only witness that shows up to every incident. Give it the four habits from this article and it will testify well. Use a log() function that stamps and labels every line, and capture tool output into the same timeline. Provide one parseable summary per run mirrored to syslog, and configure rotation before the disk forces the issue. None of it is glamorous, and all of it is the difference between reading the answer and reconstructing it.
Two natural next steps in the series: Error Handling in Bash deepens the trap machinery the summary line depends on and adds per-step failure accounting. Then the hardening article folds logging into the full production checklist a script must pass before it earns its schedule.
Frequently Asked Questions
Why printf instead of echo in the logging function?
Should my job log to a file or to syslog?
My cron job ran but the log file is empty. Why?
How do I get timestamps on rsync or sftp output?
How long should I keep transfer logs?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
