Home › Topics › Bash & Cron › Error Handling

Error Handling in Bash: Exit Codes and Traps

Left to its own devices, a bash script treats failure as a passing remark. The upload fails; the script shrugs and archives the files as if they were sent. The export step produces nothing; the script cheerfully uploads an empty directory and reports success. For interactive work this tolerance is harmless — you saw the error. For an unattended transfer job it is the root of the worst incident class there is: the job that limps on wounded. It does damage downstream while every dashboard stays green.

Good error handling in bash is not exotic. It rests on two primitives that fit in one hand. Exit codes are the one-number verdict every command leaves behind. And trap is the mechanism that guarantees your cleanup and your final status line run no matter how the script ends. Add a policy — fail loudly, retry only what deserves retrying, and tell the whole truth about partial success — and you have everything a production transfer job needs.

This article, part of our Bash & Cron transfer automation series, covers the primitives precisely. It then assembles them into a complete multi-step job with retries, per-file accounting, and cleanup that always runs.

Exit Codes: The Only Word a Process Leaves Behind

Every command that finishes hands its parent one integer between 0 and 255: its exit code (or exit status). Zero means success; anything else means some flavor of failure. Bash stores the most recent one in $?. That single number is the entire machine-readable interface between a command and your script — and between your script and cron, which is why it deserves respect from both directions.

A few conventions repay memorizing, because they turn mystery numbers into diagnoses:

Code Conventional meaning What it usually tells you
0 Success The command believes it did its job
1, 2 General failure; misuse Read the tool's stderr for the real story
126 Found but not executable Permissions or a directory where a program should be
127 Command not found PATH problems — the classic cron-environment symptom
128 + N Killed by signal N 130 = interrupted, 143 = terminated, 137 = killed outright

Transfer tools add their own vocabularies on top. rsync is unusually articulate — 23 means some files could not be transferred, 24 means source files vanished mid-run (often benign churn), 30 means a timeout. curl distinguishes could-not-resolve from could-not-connect from timed-out. The OpenSSH sftp and scp clients are blunter — essentially success or "1, something failed," with 255 reserved for the underlying SSH connection failing. So with those tools the exit code says whether and the captured stderr says why. That is one more reason the capture patterns from the logging article matter.

Finally, your script has an exit code of its own — it is the last command's, unless you say otherwise with exit N. That number is the only thing cron and your monitoring can see without reading logs. Define a tiny contract and document it in the header comment: 0 = ran clean (including "nothing to send"), 1 = something failed, 2 = bad configuration or usage. Three values cover nearly every job, and consistency across your estate is worth more than a richer scheme.

Checking Exit Codes Properly

The cardinal rule: $? is perishable. It refers to the single most recent command — and everything is a command, including echo and your log function. This classic bug ships constantly:

# WRONG -- the log line resets $?
sftp "${SFTP_OPTS[@]}" -b "$BATCH" "$REMOTE"
log INFO "sftp finished"
if [ $? -ne 0 ]; then          # tests the log command, which succeeded
    die "upload failed"        # ...so this never fires
fi

The fix is to test the command directly, which reads better anyway. The two idiomatic forms:

# form 1: branch on it
if ! sftp "${SFTP_OPTS[@]}" -b "$BATCH" "$REMOTE"; then
    die "upload failed"
fi

# form 2: compact guard for must-succeed steps
mkdir -p "$STAGING_DIR" || die "cannot create $STAGING_DIR"

# capturing the number itself (works under set -e)
rc=0
rsync -a "$SRC"/ "$DEST"/ || rc=$?
log INFO "rsync rc=$rc"

The third form deserves a word: under strict mode a failing command would normally end the script. But a command on the left of || is exempt, so || rc=$? both survives the failure and banks its code for inspection. That is the pattern for tools like rsync whose nonzero codes you want to interpret rather than die on. For example, you might treat 24, vanished source files, as a warning, while 23 stays fatal.

Two quieter traps complete the set. First, command substitution behind local: in local out=$(some_command), the exit status you see is local's — always success — so the substitution's failure is masked. Declare and assign on separate lines (local out, then out=$(some_command)) and the failure surfaces normally. Second, the one from the fundamentals article is worth repeating because it shapes design. Inside any function called as a condition (if send_one "$file"; then), set -e is suspended entirely. The function keeps executing past its own internal failures unless it checks them itself. The consequence: a function meant to be tested must do its own error handling and communicate through explicit return codes. The worked example below is built exactly that way.

Remember: read $? on the very next line or not at all — and prefer forms that never mention it: if ! cmd, cmd || die, cmd || rc=$?. Every one of those tests the command itself, leaving nothing to reset.

Pipelines: Where Status Goes to Get Lost

A pipeline reports the status of its last command, so generate | gzip hides a failed generate behind a successful gzip. Strict mode's pipefail fixes the common case: the pipeline reports the first failure instead. Two refinements matter in transfer jobs. When a nonzero status is a normal answer, accept it explicitly. grep exits 1 to mean "no matches," so a check like counting errors in a client log needs || true to say "empty is fine, keep going." And when you need each stage's verdict separately, bash's PIPESTATUS array holds them. But it is even more perishable than $? (any command overwrites it). And under -e a failing pipeline ends the script before you can look. The dissection pattern suspends strict mode for one pipeline only:

set +e
pg_dump_orders | gzip -c > "$WORK_DIR/orders.gz"
stages=("${PIPESTATUS[@]}")
set -e
log INFO "export stages: dump=${stages[0]} gzip=${stages[1]}"
[ "${stages[0]}" -eq 0 ] || die "export failed at the dump stage"

Copy the array on the first line after the pipeline, restore -e, then judge at leisure. Reserve this for pipelines whose middle genuinely matters; for the rest, pipefail plus || rc=$? is simpler and sufficient.

trap: Code That Runs No Matter What

Error handling has a second half: not just detecting failure but ending well — temp files removed, the final status line written, nothing half-finished left to confuse the next run. The naive approach puts cleanup at the bottom of the script, which is exactly where a failing script never reaches. The trap builtin fixes this by registering a handler that bash runs when an event fires. The event that matters here is EXIT. It fires on every way out — natural end, exit 1 from die(), a strict-mode abort, even termination signals like an operator's Ctrl-C or a system shutdown's TERM. (The one thing it cannot survive is kill -9, which by design gives a process no chance to do anything.)

WORK_DIR=""   # sentinel: empty until mktemp succeeds

cleanup() {
    local rc=$?                      # bank the script's exit status FIRST
    if [ -n "$WORK_DIR" ]; then
        rm -rf "$WORK_DIR"
    fi
    if [ "$rc" -eq 0 ]; then
        log INFO "summary result=ok duration=${SECONDS}s"
    else
        log ERROR "summary result=fail rc=$rc duration=${SECONDS}s"
    fi
}
trap cleanup EXIT

Three habits make trap handlers trustworthy. Bank $? on the first line — inside the handler it still holds the script's exit status, but only until the handler's own first command replaces it. Make cleanup idempotent and defensive: it must work when things it cleans never got created, which is what the empty-string sentinel and the -n test accomplish. Install the trap before creating anything it protects, and there is no window where a crash leaks. Keep it short: remove temp state, write the verdict, stop. A trap that does real work has real failures, in the one place with no one left to handle them.

Bash also offers an ERR trap that fires on failing commands. But it inherits every blind spot set -e has and needs an extra option (set -E) to reach inside functions. It is useful as a debugging aid, shaky as a foundation. The robust core is the combination above: explicit checks decide what happens, the EXIT trap guarantees how it ends. Note what you do not have to clean up: a lock held with flock releases itself when the process dies, kernel-guaranteed, as covered in the locking article — one less thing your trap can get wrong.

Fail Loudly, Retry Selectively

Policy, in three points. A transfer job that cannot do its whole job should stop, say so, and exit nonzero. A loud failure at 02:00 becomes a fixed job by 09:00, while a quiet half-success becomes a data investigation next week. Retries are for failures that time can fix: a dropped connection, a busy server, a DNS hiccup. Failures that time cannot fix — authentication rejected, unknown host key, no such directory — must not be retried. Retrying them is futile, and hammering a server with bad credentials can trip lockout defenses and turn one broken job into a locked account (see auth failures and lockouts).

A small retry loop with growing pauses covers the transient class honestly:

attempt=1
until sftp "${SFTP_OPTS[@]}" -b "$BATCH" "$REMOTE" >> "$LOG_FILE" 2>&1; do
    if [ "$attempt" -ge 3 ]; then
        die "upload failed after $attempt attempts"
    fi
    log WARN "attempt $attempt failed; retrying in $((attempt * 60))s"
    sleep $((attempt * 60))
    attempt=$((attempt + 1))
done

until repeats its body while the command fails, so success exits the loop immediately; three attempts with one- then two-minute pauses rides out blips without hiding real outages. This is deliberately the pocket version — retry budgets, exponential backoff with jitter, and when to retry the step versus the whole job are the territory of our retry and error handling series.

The Worked Example: A Multi-Step Job With Traps

Now the assembly. This job gathers export files, sends each with retries, and archives only what was actually sent. It writes a manifest of the batch and sends the manifest last as the "batch complete" signal. The trap guarantees the temp directory dies and the summary line gets written on every path out.

#!/usr/bin/env bash
set -euo pipefail

readonly JOB_NAME="send-daily-batch"
readonly EXPORT_DIR="/data/export/batch"
readonly ARCHIVE_DIR="/data/archive/batch"
readonly SSH_KEY="/etc/transfer/keys/batch_ed25519"
readonly KNOWN_HOSTS="/etc/transfer/known_hosts"
readonly REMOTE="batch@files.partner.example"
readonly LOG_FILE="/var/log/transfer/${JOB_NAME}.log"
readonly SFTP_OPTS=(-i "$SSH_KEY" -o IdentitiesOnly=yes
                    -o UserKnownHostsFile="$KNOWN_HOSTS"
                    -o StrictHostKeyChecking=yes -o ConnectTimeout=30)

WORK_DIR=""

log() {
    local level="$1"; shift
    printf '%s %s[%d] %s: %s\n' \
        "$(date '+%b %d %H:%M:%S')" "$JOB_NAME" "$$" "$level" "$*" >> "$LOG_FILE"
}

die() { log ERROR "$*"; exit 1; }

cleanup() {
    local rc=$?
    if [ -n "$WORK_DIR" ]; then rm -rf "$WORK_DIR"; fi
    if [ "$rc" -eq 0 ]; then
        log INFO "summary result=ok duration=${SECONDS}s"
    else
        log ERROR "summary result=fail rc=$rc duration=${SECONDS}s"
    fi
}
trap cleanup EXIT

send_one() {   # send_one FILE -- 0 = delivered, 1 = gave up
    local file="$1" name attempt
    name="${file##*/}"
    for attempt in 1 2 3; do
        if sftp "${SFTP_OPTS[@]}" -b - "$REMOTE" >> "$LOG_FILE" 2>&1 <<EOF
put "$file" "incoming/$name.part"
rename "incoming/$name.part" "incoming/$name"
EOF
        then
            log INFO "sent $name attempt=$attempt"
            return 0
        fi
        log WARN "attempt $attempt failed for $name"
        if [ "$attempt" -lt 3 ]; then sleep $((attempt * 30)); fi
    done
    return 1
}

main() {
    log INFO "run started"
    [ -d "$EXPORT_DIR" ]  || die "export directory missing: $EXPORT_DIR"
    [ -d "$ARCHIVE_DIR" ] || die "archive directory missing: $ARCHIVE_DIR"

    WORK_DIR=$(mktemp -d "/var/tmp/${JOB_NAME}.XXXXXX")
    local manifest="$WORK_DIR/manifest.txt"
    : > "$manifest"

    local file sent=0 failed=0
    for file in "$EXPORT_DIR"/*.csv; do
        [ -e "$file" ] || { log INFO "nothing to send"; return 0; }
        if send_one "$file"; then
            printf '%s\n' "${file##*/}" >> "$manifest"
            mv -- "$file" "$ARCHIVE_DIR"/
            sent=$((sent + 1))
        else
            failed=$((failed + 1))
        fi
    done

    log INFO "files sent=$sent failed=$failed"
    if [ "$failed" -gt 0 ]; then
        return 1                       # partial failure is still failure
    fi
    send_one "$manifest" || return 1   # manifest last: the all-clear signal
    log INFO "manifest sent"
}

main "$@"

The design decisions, made explicit:

  • The trap is armed before anything exists to leak. WORK_DIR starts empty, cleanup tolerates that, and from the moment mktemp succeeds, deletion is guaranteed on every exit path.
  • send_one trusts nothing implicit. Because it is always called inside if, strict mode is suspended within it — so it checks its one real command explicitly and speaks only through return. Per-file, it retries transient failures and gives up cleanly at three.
  • State changes only follow confirmed success. A file is archived, and listed in the manifest, strictly after its upload-and-rename succeeds. Failed files stay in EXPORT_DIR, so the next run naturally retries exactly the unsent remainder — recovery by design rather than by hand. (The -- in mv -- ends option parsing, so a filename starting with a dash cannot be mistaken for a flag.)
  • The manifest travels last, and only after a perfect run. Downstream, "manifest present" means "batch complete and trustworthy" — the marker-file discipline our partial-file safety series formalizes.
  • The exit contract holds. Nothing to send: 0. Everything sent: 0. Anything failed: 1, with sent= and failed= counts in the log telling the exact story — and the far end's records completing it. When the destination is a Windows server running Sysax Multi Server, its activity logs — written to file and database — show precisely which uploads and renames landed. That is the fastest way to reconcile "we sent 47 of 50" against what actually arrived.

Partial Failure Is a First-Class Outcome

Multi-file jobs force a question single-command scripts never face: what is the truth when 47 of 50 files made it? Two honest designs exist. All-or-nothing stops at the first failure and reports the batch failed — right when the files form one logical unit a consumer must see complete. Continue-and-account — the worked example's choice — delivers what it can, counts what it cannot, and reports failure if anything failed. That is right when files are independent and a late file is better than no files. The dishonest design is the default one: deliver some, lose track, exit 0. Whichever you choose, the summary must carry the numbers, and the exit code must not say "ok" when the truth is "mostly." Something must watch for the pattern. One failed file tonight is a retry; the same file failing five nights running is a poison file needing a human. That is a distinction our retry series and monitoring series divide between them.

It is also worth knowing when not to build this by hand. On the Windows side of a mixed estate, Sysax FTP Automation ships retry and error handling as configuration on its scheduled transfer tasks, with email notifications for failures. Those are the same policies this article implements in bash, expressed as settings. The concepts transfer intact; only the syntax changes.

Loud Failures Are a Feature

Everything here serves one operational virtue: when this job cannot do its whole job, everyone finds out immediately, and nothing is left half-done in the dark. Exit codes are checked at the command, statuses captured before they perish, and pipelines judged stage by stage where it matters. An EXIT trap ends every run — success, failure, or Ctrl-C — with clean state and a written verdict. Retries are reserved for failures that deserve them, and partial success is reported as exactly what it is.

From here, Hardening a Bash Transfer Job for Production adds the final layer — safe temp files, credentials, input validation, and the review checklist. The logging article makes sure the loud failures land somewhere someone will hear them.

Frequently Asked Questions

Why does $? show 0 right after a command I know failed?
Because something ran in between — a log line, an echo, anything — and $? always reflects the most recent command. Test the command directly with if ! cmd or cmd || rc=$?, and the problem disappears because there is no gap for another command to occupy.
Do I still need to check exit codes if I use set -e?
Yes. set -e only converts unchecked failures into script exits. It cannot express "retry this," "tolerate that," or "count this and continue." And it is suspended inside conditions and functions called from them. Explicit checks make the decisions; set -e is the safety net for the ones you missed.
Does a trap on EXIT run if my script is killed?
For ordinary termination signals like TERM and INT (Ctrl-C), yes — bash runs the EXIT trap on the way out. The exception is SIGKILL (kill -9), which by design terminates a process without giving it any chance to run cleanup. That is one reason cleanup should also tolerate leftovers from a previous run.
What exit codes should my own script use?
Keep it small and documented: 0 for a clean run (including "nothing to do"), 1 for any failure, 2 for configuration or usage errors. Schedulers and monitors mostly care about zero versus nonzero; the log carries the detail. Consistency across all your jobs matters more than a clever numbering scheme.
Exit code 127 keeps appearing in my cron runs. What is it?
127 means "command not found" — the script asked for a program the shell could not locate. Under cron that is almost always the minimal PATH problem: the tool lives in a directory cron's environment does not search. Use absolute paths or set PATH explicitly in the script.
Why not just retry every failure a few times?
Because some failures cannot succeed on retry — a rejected password, an unknown host key, a missing remote directory. Retrying them wastes the window, muddies the log, and can trigger account lockouts on the server. Retry what time can fix; fail immediately on what it cannot.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.