Parsing File Names Reliably in Scripts
"The stamp field says orders." "That is not a date." "The script did not seem to mind." Everything else in this series is about producing good names. This article is about the other side of the contract: the scripts that read them, and what they do when the third field is a word. Parsing is where a naming convention either pays off — fields extracted in one line, decisions made confidently — or quietly fails. It fails when a script splits a name on faith and gets fields that are not what it assumed. The script routes, loads, or deletes based on garbage. The difference between the two outcomes is not cleverness. It is a habit: validate first, extract second, and reject what does not conform.
This closing article of our Naming & Datestamps series builds that habit into working code three times — in bash, PowerShell, and Python. Most transfer estates run at least two of the three. Along the way we cover the validation a regex cannot do. We cover the reject-to-error-folder pattern that keeps one bad name from stalling a pipeline. We also cover the tiny test set that keeps a parser honest for years. As throughout the series, examples use pattern tokens — YYYYMMDD where an eight-digit date appears in production. The code builds its stamps at runtime with date functions, so every snippet runs as written.
Validate First, Extract Second
A file name arriving in a folder is input, and input is untrusted — even when the producer is your own team. The producer has bugs, upgrades, and occasional humans dropping files by hand. The naive parser extracts on faith: split on underscore, take field three, call it the date. Feed it a nonconforming name and it does not fail — it succeeds wrongly. It hands back a "date" that is actually a sequence number or half a word. The rest of the script then acts on that. The damage lands two steps later, disconnected from its cause. The naive parser has never once reported a problem, which is the problem.
Watch that failure in slow motion. A partner adds a region field to their names — a change they consider harmless, since every file still ends in .csv. The split-on-faith parser keeps right on going:
name arrives: acme_eu_orders_YYYYMMDD_0001.csv (new field added) split on "_": [acme] [eu] [orders] [YYYYMMDD] [0001.csv] naive parser: src=acme type=eu stamp=orders seq=YYYYMMDD
No error is raised anywhere. The "datestamp" is now the word orders, and the "sequence" is an eight-digit date. Whichever job consumes those fields — an age comparison, a routing rule, a ledger key — misbehaves in a direction that points nowhere near the real cause. A validating parser turns the same event into one clean rejection with the offending name attached. The partner's change is discovered the day it ships instead of the month after.
The reliable parser inverts the order. First it asks a yes/no question: does this entire name match the convention's pattern? A regular expression — a compact language for describing text patterns — answers that in one operation. Only after a full match does the parser extract fields. At that point, the fields are guaranteed to be where the pattern says they are. Everything that fails the match takes a completely different path: not into the pipeline, not into a crash. It takes a reject path that quarantines the file and tells a human. The pattern comes straight from the convention document — the one designed in designing a naming convention. That document stores the regex precisely so that producers and parsers share one definition of "valid."
Our working contract for all three languages, in tokens: <source>_<type>_YYYYMMDD_<seq>.csv — for example acme_orders_YYYYMMDD_0042.csv — lowercase, underscore-separated, eight-digit big-endian stamp, four-digit zero-padded sequence.
Parsing in Bash
Bash gives you two tools worth knowing, and they compose: parameter expansion for trimming, and the [[ =~ ]] regex test for validation with capture. Parameter expansion strips pieces by pattern — ${name%.csv} removes the extension from the end, ${name##*_} keeps only what follows the last underscore. Handy for quick trims, but expansion never validates; the disciplined path runs the regex first. (Everything here follows the quoting and strict-mode habits from our bash and cron automation series — unquoted names are how spaces become outages.) Parameter expansion trims; it does not judge.
#!/bin/bash
# parse_name NAME -> sets src, typ, stamp, seq; returns 1 if invalid
NAME_RE='^([a-z][a-z0-9-]*)_([a-z]+)_([0-9]{8})_([0-9]{4})\.csv$'
parse_name() {
local name="$1"
[[ "$name" =~ $NAME_RE ]] || return 1
src="${BASH_REMATCH[1]}"
typ="${BASH_REMATCH[2]}"
stamp="${BASH_REMATCH[3]}"
seq="${BASH_REMATCH[4]}"
# the regex checks digits, not the calendar: delegate that
date -d "$stamp" >/dev/null 2>&1 || return 1 # GNU date
return 0
}
f="acme_orders_$(date -u +%Y%m%d)_0042.csv" # stamp built at runtime
if parse_name "$f"; then
echo "source=$src type=$typ day=$stamp part=$seq"
else
echo "REJECT: $f does not match convention" >&2
fi
Here are details that earn their keep: the regex lives in a variable. Quoting a regex inline inside [[ ]] changes its meaning in subtle ways, so the unquoted-variable form is the reliable idiom. BASH_REMATCH holds the captured groups after a successful match, index 1 onward. And the calendar check is delegated to date -d, because a regex happily accepts an eight-digit stamp with month thirteen. The GNU date on Linux systems rejects impossible dates for you. There is also a quicker splitting idiom for names already validated: IFS=_ read -r src typ stamp seq <<< "${f%.csv}" splits on underscores in one line. Use it only behind the regex. On a name with the wrong number of fields, a bare read shifts values into the wrong variables without a word of complaint.
Parsing in PowerShell
PowerShell's regex support includes named groups — captures you address by name instead of number, which keeps scripts readable when conventions evolve. Two platform habits matter here. First, -match is case-insensitive by default. Use -cmatch so that ACME_ fails validation exactly as it would on the case-sensitive systems downstream. That is the discipline argued in safe characters. Second, .NET's date parser will verify the calendar for you via TryParseExact. (For the wider scripting context — jobs, errors, exit codes — see our PowerShell transfer automation series.)
# Parse-TransferName NAME -> object with fields, or $null if invalid
$NamePattern = '^(?<src>[a-z][a-z0-9-]*)_(?<type>[a-z]+)_(?<stamp>\d{8})_(?<seq>\d{4})\.csv$'
function Parse-TransferName {
param([string]$Name)
if ($Name -cnotmatch $NamePattern) { return $null }
$day = [datetime]::MinValue
$ok = [datetime]::TryParseExact($Matches['stamp'], 'yyyyMMdd',
[cultureinfo]::InvariantCulture, 'None', [ref]$day)
if (-not $ok) { return $null } # digits, but not a real date
[pscustomobject]@{
Source = $Matches['src']
Type = $Matches['type']
Stamp = $Matches['stamp']
Seq = [int]$Matches['seq']
}
}
$f = 'acme_orders_{0}_0042.csv' -f (Get-Date -Format yyyyMMdd)
$parsed = Parse-TransferName $f
if ($parsed) { $parsed } else { Write-Warning "REJECT: $f" }
After a successful -cmatch, the automatic $Matches table holds each named group. Returning a real object rather than loose variables means the caller gets all fields or nothing. There is no state in which "half a parse" leaks into later logic. The quick-split alternative exists here too: $parts = $base -split '_', guarded by a $parts.Count check — and again belongs only behind the validation, never instead of it. Half a parse is a whole bug.
Parsing in Python
Python is where parsers grow up. When a flow needs a processed-files ledger, per-file accounting, or richer error reporting, scripts tend to land here. Our Python transfer automation series covers that migration. The pattern is the same, with two Python-specific strengths. One is re.fullmatch, which anchors the match to the entire string so a valid-looking prefix cannot sneak a bad name through. The other is datetime.strptime, which validates the calendar in the same call that parses it.
import re
from datetime import datetime, timezone
NAME_RE = re.compile(
r"(?P<src>[a-z][a-z0-9-]*)_"
r"(?P<type>[a-z]+)_"
r"(?P<stamp>\d{8})_"
r"(?P<seq>\d{4})\.csv"
)
def parse_name(name):
"""Return a dict of fields, or None if the name does not conform."""
m = NAME_RE.fullmatch(name)
if m is None:
return None
fields = m.groupdict()
try: # digits, but is it a date?
datetime.strptime(fields["stamp"], "%Y%m%d")
except ValueError:
return None
fields["seq"] = int(fields["seq"])
return fields
stamp = datetime.now(timezone.utc).strftime("%Y%m%d")
f = f"acme_orders_{stamp}_0042.csv"
parsed = parse_name(f)
print(parsed if parsed else f"REJECT: {f}")
Returning None for any failure — wrong shape or impossible date — gives callers a single honest signal. The calling code reads as policy: parse, and if the result is None, take the reject path. The str.split("_") idiom with a length check is fine for already-validated names, same caveat as the other two languages. Resist the temptation to be lenient here — fullmatch, not search — because every relaxation admits a family of names you never meant to accept. They will all arrive eventually, on the same night.
Validation Beyond the Regex
All three snippets already do one check the pattern cannot: the calendar. \d{8} accepts a stamp with month thirteen or day forty-one; only a date parser knows better. Depending on the flow's stakes, three more layers are worth their cost:
- Allowlists for coded fields. The convention's registered source and type codes are finite — check membership. Then
acmepasses and a typo likeamceis caught at the door instead of creating a brand-new "partner" downstream. One line per language: acase "$src" in acme|globex) ;; *) return 1;; esacin bash, orfields["src"] not in KNOWN_SOURCESagainst a set in Python. - Range and sanity checks. A sequence of
0000when numbering starts at0001, or a stamp days in the future, is legal to the regex and suspicious to a human. Encode the suspicion. The rules come from sequence numbers and uniqueness. - Content agreement. The name is a claim about the content — a
.csvthat begins with an archive's magic bytes is lying, and name-level parsing cannot see it. That deeper inspection belongs to the validation stage of the pipeline, covered in our pre- and post-processing series.
How far to go is a judgment call: a name check is cheap insurance everywhere. Full content validation is reserved for flows where a bad file is expensive. The principle stays constant — each check runs before anything irreversible happens. We learned the meaning of irreversible from a purge job.
The Reject Path: Error Folder, Not Exception
What happens to the name that fails? The wrong answers are common. One is to crash the job (one odd file stalls the whole batch). Another is to skip silently (the file waits forever and nobody knows). A third is to "helpfully" rename the file to fit (which hides a producer bug and manufactures a name the producer never sent). I have shipped all three wrong answers, in that order. The right answer is boring and mechanical:
- Move the file to an error folder — a quarantine area beside the inbox, on the same filesystem. That makes the move a single atomic rename that can never half-complete. It is the same rule that governs in-flight files in our partial-file safety series. The original name is preserved exactly: it is evidence.
- Log one line with the reason: the name, the check it failed, the folder it came from. "Rejected: bad stamp field" is actionable; a bare "error" is not.
- Alert once per file, not once per retry loop, and keep processing the rest of the batch. One nonconforming name should cost the pipeline one file, not the night's run. The alert-fatigue side of that is a story our retry and error handling series tells in full.
- Never delete, never auto-fix. The file is someone's data and the name is a clue. A human inspects, the producer gets fixed, and the file is re-submitted properly.
Kestrel Payroll ran the "helpfully rename" answer for two years without noticing it was an answer. When a partner's export began emitting an uppercase source code, the intake script lowercased it and carried on. The producer bug went unreported until the partner's own audit asked why their files had one name in their logs and another in ours. Nothing was lost; the two sets of logs no longer matched, which turned a one-minute reconciliation into an afternoon. The rename was replaced with a move to the error folder, and the partner fixed the export the same week. The first rejected file arrived within a month, exactly as designed.
Wired onto the bash parser from earlier, the whole intake pass is a dozen lines:
# intake pass: validate everything, quarantine what fails
inbox=/data/inbox; work=/data/working; err=/data/error
shopt -s nullglob # empty inbox = zero iterations
for f in "$inbox"/*; do
name="$(basename "$f")"
if parse_name "$name"; then
mv -n -- "$f" "$work/$name"
else
mv -n -- "$f" "$err/$name"
logger -t intake "REJECT $name: fails naming convention"
fi
done
Every character of ceremony here is a defense. nullglob makes an empty inbox loop zero times instead of once over a literal *. The -- tells mv that option parsing is over, so a hostile or accidental leading-hyphen name cannot masquerade as a flag. That is the hazard from the safe-characters article, defended in one keystroke. mv -n refuses to overwrite: if the error folder already holds a file by that name, the earlier evidence is preserved rather than clobbered. And because inbox, working, and error folders live on one filesystem, each mv is a single atomic rename — a half-moved file is impossible. logger drops the reason into the system log, where it joins the trail an investigator will follow. The layered troubleshooting method starts from exactly that trail. None of it is optional, and all of it looks optional.
That investigation is where good infrastructure shows. The activity logs of Sysax Multi Server record the file name of every transfer — to a log file and optionally a database. So a rejected name can be traced to the exact upload: when it arrived and from which account. That usually identifies the misbehaving producer in minutes. And since the intake folder is typically watched, Sysax FTP Automation pairs folder monitoring with pre/post processing and email notification on its tasks. Those are the natural slots, in a tool-managed flow, for exactly the validate-then-quarantine behavior this section describes. This whole reject design is the error half of the watch folder pattern, applied to names.
Remember: a parser has exactly two honest outputs — a full set of validated fields, or a rejection with a reason. Anything in between ("probably the date", "close enough") is a bug that has not happened yet.
The Tiny Test Set That Keeps the Parser Honest
A parser is code that makes decisions, and decisions deserve tests. You do not need a framework — just a dozen names with expected verdicts, run by a loop. Keep them next to the parser and run them after every change to it or to the convention. The set almost writes itself: one valid name, then one name for each rule the convention states, breaking exactly that rule:
| Test name (tokens) | Expected | Rule exercised |
|---|---|---|
acme_orders_YYYYMMDD_0001.csv | accept | the happy path |
ACME_orders_YYYYMMDD_0001.csv | reject | lowercase only (catches -match vs -cmatch) |
acme orders_YYYYMMDD_0001.csv | reject | no spaces |
acme_orders_YYYYMMDD.csv | reject | missing sequence field |
acme_orders_YYYYMMDD_001.csv | reject | sequence width is four |
acme_orders_99999999_0001.csv | reject | digits but not a calendar date |
acme_orders_YYYYMMDD_0001.csv.tmp | reject | in-flight suffix must not parse |
acme_orders_YYYYMMDD_0001_x.csv | reject | fixed field count |
And the runner, in bash, using the parse_name function from above — the stamp is generated, so this executes verbatim:
stamp="$(date -u +%Y%m%d)"
fails=0
while read -r want name; do
if parse_name "$name"; then got=accept; else got=reject; fi
if [ "$got" = "$want" ]; then echo "PASS $name"
else echo "FAIL $name (wanted $want, got $got)"; fails=1; fi
done <<EOF
accept acme_orders_${stamp}_0001.csv
reject ACME_orders_${stamp}_0001.csv
reject acme orders_${stamp}_0001.csv
reject acme_orders_${stamp}.csv
reject acme_orders_${stamp}_001.csv
reject acme_orders_99999999_0001.csv
reject acme_orders_${stamp}_0001.csv.tmp
reject acme_orders_${stamp}_0001_x.csv
EOF
exit $fails
Note the harness detail: read -r want name puts everything after the first word into name. So even the space-bearing test case survives the loop intact. The harness itself has to follow the quoting rules it is testing. Eight lines of test protect you at the two moments parsers break: when someone edits the regex, and when the convention gains a field. Run the set in both places — it is as much a test of the convention document as of the code. Eight lines is shorter than the incident report they replace.
The Contract, Closed
This article closes the loop the series opened. Producers promise a shape — fields, delimiters, a sortable stamp, safe characters, a unique sequence. The parser is the consumer's side of that promise. Verify the whole shape, extract with confidence, and validate what regexes cannot see. Send every violation to quarantine with its evidence intact. Written this way, a parser is not defensive boilerplate; it is the mechanism that turns a naming convention from documentation into an enforced interface. And the next time the stamp field says orders, the script will mind.
If you arrived here without the foundations, here is the path back through the series. The article designing the convention defines what your regex should be. The article safe characters explains why the character class is so strict. And why file names make or break automation is the story of what happens when nobody does any of this.
Frequently Asked Questions
Why use a regex instead of just splitting on underscores?
Can a regex check that the date in a name is real?
What should my script do with a file name it can't parse?
Why did my PowerShell parser accept an uppercase name my bash parser rejected?
Is a handful of test names really worth maintaining?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
