Home › Topics › Bash & Cron › Hardening Jobs

Hardening a Bash Transfer Job for Production

There is a moment in the life of every transfer script when it stops being an experiment and starts being infrastructure. It gets a schedule, other people start depending on its output, and its failures become incidents instead of curiosities. Most scripts cross that line without ceremony — nobody decides they are production; they just quietly become load-bearing. The hardening pass is the ceremony that should have happened.

By this point in the series, the big machinery is in place: a well-shaped script, a tested schedule, overlap protection, real logging, and error handling that fails loudly. What remains is the final layer of assumptions nobody has examined yet. Where do files get created and with what permissions? What happens when an input is hostile or missing? Where do the credentials actually live? Has anyone besides the author ever read the thing?

This article is that final pass, and the capstone of our Bash & Cron transfer automation series. It covers the remaining hardening topics one by one, then the full production checklist in copyable form. A worked review takes a real "works for me" script through it.

What "Production-Ready" Actually Means

A useful definition, because "it works" is not one: a production transfer job runs unattended without depending on anything it has not pinned down; stays safe when its assumptions break; can be read, reviewed, and re-run by someone who is not its author; and tells the truth about every outcome. Notice that three of those four clauses are about failure and strangers, not success — that is what distinguishes production thinking from getting a script to work.

Everything below attacks one category: implicit assumptions. The script assumes the working directory, the environment, the permission bits, the friendliness of filenames, the privacy of its command line. Each hardening step either makes an assumption explicit — written in the script, where it is enforced and reviewable — or removes it entirely.

Pin the Environment: Paths, PATH, and Locale

An unattended job should carry its environment with it rather than inherit whatever the scheduler provides. The cron article showed how minimal that inheritance is; hardening means depending on none of it:

  • Set PATH explicitly at the top of the script — export PATH=/usr/local/bin:/usr/bin:/bin — or call every binary by absolute path. Either way, the script now behaves identically from cron, from your shell, and from a colleague's shell.
  • Use absolute paths for every file and directory, held in the readonly configuration block. Never rely on the working directory. If a step genuinely requires changing directory, check it: cd /data/export || die "cd failed". The unchecked-cd-then-act sequence is one of scripting's most famous disasters.
  • Pin the locale where output is parsed. Tool messages, sort order, and number formatting can vary with language settings. export LC_ALL=C gives every run — regardless of who or what launched it — the same predictable, parseable output.
  • Pin the SSH details. Pin the key file, known-hosts file, and timeouts — either as explicit options (as this series' examples do) or collected in a dedicated Host block in an ssh_config file. The latter is a tidy pattern our guide to ssh_config for transfers walks through.

Files the Job Creates: umask and Safe Temp Space

Every file a script creates gets permission bits, whether anyone chose them or not. The chooser is umask — a per-process mask that removes permissions from newly created files. The common default of 022 removes only write access for group and others, which means your job's staging files, logs, and downloaded payloads arrive world-readable. Transfer payloads are usually somebody's business data; "every local account can read them" should be a decision, not an accident. Setting umask 027 near the top of the script yields files readable by owner and group only (640, directories 750); 077 restricts to the owner alone. The subtleties — inheritance, how it interacts with what servers create — live in our umask gotchas article.

Temporary files deserve equal suspicion. The traditional habit — a predictable name like /tmp/job.tmp — carries two real risks. One is collisions, where two runs or two jobs stomp each other's files. The other is mischief, because /tmp is world-writable. Any local account can create that exact name first — even as a symbolic link pointing somewhere you would never write. Your script, running as the transfer account, follows it. Both problems have one standard fix:

WORK_DIR=$(mktemp -d "/var/tmp/${JOB_NAME}.XXXXXX")   # e.g. /var/tmp/push-orders.k3Qf9w

mktemp -d creates a directory with an unpredictable suffix, owned by you, permissions locked to the owner, atomically — no window where someone else can claim the name. Pair it with the EXIT-trap cleanup from the error-handling article so it never outlives the run. Two placement notes: prefer /var/tmp or a job-owned staging directory for anything sizable. And when a temp file's destiny is an atomic rename into a final directory, create it on the same filesystem as that directory. A rename across filesystems silently becomes copy-then-delete, which reintroduces the half-written-file window the rename was meant to close.

Validate Inputs Before Trusting Them

A scheduled script's inputs arrive from arguments, configuration, and — most dangerously — the filesystem and remote listings, where other people choose the names. Validation is cheap, and the pattern is always the same: check early, fail with a named reason, exit with the usage code.

# arguments: accept only what you documented
case "${1:-}" in
    ""|--dry-run) ;;                              # allowed
    *) printf 'usage: %s [--dry-run]\n' "$0" >&2; exit 2 ;;
esac

# preconditions: named, early, fatal
[ -d "$SOURCE_DIR" ] || die "source dir missing: $SOURCE_DIR"
[ -r "$SSH_KEY" ]    || die "key not readable: $SSH_KEY"

# disk space: fail before filling the disk, not after
avail_kb=$(df -P "$STAGING_DIR" | awk 'NR==2 {print $4}')
[ "$avail_kb" -gt 500000 ] || die "less than 500MB free in $STAGING_DIR"

# filenames from outside: conservative allowlist, or reject
name="${file##*/}"
if ! [[ "$name" =~ ^[A-Za-z0-9][A-Za-z0-9._-]*$ ]]; then
    log WARN "rejecting suspicious filename: $name"
    continue
fi

The filename check is the one juniors skip and regret. Names arriving from partners can contain spaces, quotes, leading dashes, or worse. The allowlist above accepts only names that start with a letter or digit and contain nothing but safe characters. That is the conservative set that survives every system, per our file naming series. Note it works with, not instead of, the other defenses: quoting everywhere and -- before positional arguments (mv -- "$file" ...) mean even a name that slips through cannot be parsed as options. Rejected files should be logged and set aside for a human, not silently skipped — a naming problem is a conversation with the sender waiting to happen.

Credentials Sourced Safely

The rules compress to one sentence — no secret may appear in the script, the crontab, the process list, or the log. The process list is the one that surprises people. Every argument of every running command is visible to all local users via ps. So a password passed as --password secret123 is public for the duration of the transfer. From that rule, the practice follows:

  • Prefer key authentication with a dedicated service account. A per-job SSH key, file mode 600, owned by the transfer account, referenced by path — no secret string exists anywhere in the job. Our guides to service account hygiene and generating and storing keys carry the details, including rotation.
  • When a password is unavoidable (say, FTPS via curl), keep it out of the command line. curl reads credentials from a .netrc file (--netrc-file) or an options file (-K), both permission-locked to 600. The command line then names a file, not a secret.
  • Externalize secrets to a sourced file when the script needs them as variables. Use a root-owned, group-readable file such as /etc/transfer/push-orders.env, mode 640, sourced at startup. Then verify it: : "${API_TOKEN:?not set}" makes the script die immediately, with a clear message, if the file forgot to define it. Keep that file out of version control and backups that leave the security boundary.

Remember: the goal is not to hide secrets cleverly — it is to need fewer of them. Every integration moved to key authentication is one password that can never leak through a command line, a log, or a copied script. Hardening by subtraction beats hardening by encryption of things that did not need to exist.

A Dry-Run Flag Earns Trust

The single most confidence-building feature a transfer job can have is a mode that shows what it would do — against production configuration, with production files — while moving nothing. It turns "I think this change is safe" into "here is the list of actions the changed script would take." The implementation is a flag and a wrapper:

DRY_RUN=0
if [ "${1:-}" = "--dry-run" ]; then DRY_RUN=1; fi

run() {   # wrap every command that changes state
    if [ "$DRY_RUN" -eq 1 ]; then
        log DRYRUN "would run: $*"
        return 0
    fi
    "$@"
}

# mutating steps go through run(); read-only steps do not
run sftp "${SFTP_OPTS[@]}" -b "$BATCH_FILE" "$REMOTE"
run mv -- "$file" "$ARCHIVE_DIR"/

The wrapper executes its arguments verbatim ("$@" preserves each word exactly), or logs them instead when the flag is set. The craft is in what you wrap: only the mutating steps. Listings, existence checks, and validation should still run in dry-run mode — that is what makes the rehearsal honest, since the decisions are made against real inputs. Tools with a native dry-run mode do it even better: rsync's -n produces a faithful preview of exactly which files would transfer, a technique treated fully in rsync dry runs and verification. Reach for the dry run before the first scheduled night, after every edit, and during every review — including the one at the end of this article.

The Production Checklist

Here is the whole series compressed into a review instrument. A transfer script that can answer yes down this list has earned a schedule; print it, and check boxes with a colleague, not from memory.

  • Identity and credentials
    • Runs as a dedicated service account, not a person's account
    • Key-based authentication; key file mode 600, owned by the service account
    • No secrets in script text, crontab, command lines, or logs
    • Host keys pinned in a job-owned known-hosts file; strict checking on
  • Environment and paths
    • Bash shebang plus set -euo pipefail
    • PATH set explicitly (or absolute paths throughout); LC_ALL pinned
    • All file paths absolute, in a readonly configuration block
    • umask chosen deliberately; no world-readable payloads
  • Safety mechanisms
    • Overlap-protected with flock (skip or bounded wait — chosen, not defaulted)
    • Temp space via mktemp, cleaned by an EXIT trap
    • Deliveries atomic: temporary name, then rename
    • Inputs validated: arguments, preconditions, disk space, filename allowlist
  • Failure behavior
    • Exit-code contract (0 / 1 / 2) documented in the header
    • Retries for transient failures only, with a bounded count
    • Partial failure counted, reported, and reflected in the exit code
    • Every || true is deliberate and commented
  • Logging and monitoring
    • Timestamped log lines; transfer-tool output captured into the same log
    • One parseable summary line per run; verdict mirrored to syslog
    • Rotation configured; retention decided
    • Something alerts on absence of success, not just presence of failure
  • Documentation and review
    • Header states purpose, owner, schedule, and where the runbook lives
    • In version control; shellcheck clean
    • Dry-run flag exists and was exercised against production config
    • Tested cron-style (env -i, as the service account) before scheduling
    • Reviewed by someone who is not the author

A Worked Review: From "Works for Me" to Production

Theory is easy to nod along with, so here is the checklist applied to a script of the kind that quietly runs half the world. It was written in five minutes, it has "worked" for months, and its author just handed in their notice:

#!/bin/bash
# quick script to send the report
cd /data/reports
sftp -o StrictHostKeyChecking=no partner@files.partner.example <<EOF
put report.csv
EOF
echo "sent"

The review, finding by finding:

  1. No strict mode. Any failure is a passing remark; the script continues wounded. One line fixes it.
  2. Unchecked cd, then relative paths. If /data/reports is missing or unreadable, the script proceeds in whatever directory it started in and uploads the wrong report.csv — or a colleague's file of that name.
  3. Host-key checking disabled. StrictHostKeyChecking=no means the job will deliver the report to whatever machine answers that address — the exact protection against interception, switched off for convenience. The fix is a pinned known-hosts file, seeded and verified once.
  4. Whose credentials? No key is named, so the job leans on the launching user's keys and agent. It works only when the author runs it, which is why it "cannot be moved to cron." A dedicated service account and per-job key make it portable and reviewable.
  5. The output lies. echo "sent" prints whether or not the upload succeeded, and the script exits 0 regardless. There is no log, no timestamp, no record of what was sent when.
  6. Nothing prevents overlap, nothing bounds hangs. No lock, no ConnectTimeout — a stuck session and a second run can coexist indefinitely.
  7. Non-atomic delivery. The partner can read report.csv half-written; upload-then-rename costs two lines.
  8. No documentation, no review, no dry run. The header does not say what "the report" is, who owns the flow, or what schedule it should run on.

The rebuilt version is exactly the skeleton this series has been assembling. It has strict mode and a readonly config block naming the key, known-hosts file, and timeout options. It has flock before the first side effect, the batch's put-then-rename, and a checked exit with a logged summary. So it is not repeated here; the fundamentals article's skeleton plus this article's checklist reproduce it line for line. What is worth showing is the residue. The review recorded an owner, a schedule, and an exit-code contract in the header. The reviewer ran --dry-run and the env -i cron rehearsal before sign-off. And both names now sit in the header comment. Total cost, about an hour. The next departure in that team will transfer a job, not a mystery.

One more finding belongs in an honest review: the server side has controls of its own, and hardening only the client half leaves them unexamined. Suppose the far end is yours and runs on Windows — say Sysax Multi Server, serving SFTP, FTPS, FTP, and HTTPS as a Windows service. In that case, the review should confirm the job's service account exists there with least privilege and that authentication and encryption settings match policy. It should also confirm the server's activity logging (to file and database) is on. That way, the server-side record of this job exists independently of the script's own log.

Earning the Schedule

Hardening is finishing work: none of it changes what the job does on a good night, and all of it changes what happens on the bad one. The hostile filename is set aside instead of executed, and the full disk is caught before the half-written file. The leaked password never existed to leak, and the departed author's script explains itself. A transfer job that passes the checklist gets to be what production automation should be: boring, indefinitely.

If you arrived here mid-series, the load-bearing chapters are locking, logging, and error handling — this article assumed all three. And when the estate you harden includes Windows machines originating transfers, the same checklist thinking applies with different mechanics. There, Sysax FTP Automation covers scheduled and scripted transfers with wizard-built tasks, folder monitoring, OpenPGP encryption, and email notifications. That is configuration standing in for the bash you would otherwise write and review by hand.

Frequently Asked Questions

What umask should a transfer job use?
Start from 027 — files land readable by owner and group, invisible to everyone else — and tighten to 077 when only the service account needs access. The wrong answer is inheriting the default silently: whether other local accounts can read payload files should be a written decision.
Why is writing temp files straight into /tmp risky?
Because /tmp is world-writable and predictable names can be claimed first by any local account — including as a symlink that redirects your write somewhere harmful. And two runs using the same name corrupt each other. mktemp fixes both: unpredictable name, tight permissions, atomic creation.
How do I keep a password out of the process list?
Never pass it as a command-line argument — arguments are visible to every local user via ps for as long as the command runs. Put credentials in a permission-locked file the tool reads (a netrc or options file for curl, for example), or eliminate the password entirely with key authentication.
Should read-only steps run during a dry run?
Yes. Listings, precondition checks, and validation should execute normally so the dry run makes its decisions against real inputs — only state-changing commands get intercepted and logged. A dry run that skips everything proves nothing; one that skips only mutations is a genuine rehearsal.
Who should review a transfer script before it is scheduled?
Any competent peer who did not write it — the point is fresh eyes on the assumptions, not seniority. Give them the checklist and the dry run, and have them actually run the script cron-style as the service account. Record both names in the header so ownership survives staff changes.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.