Home › Topics › Pre & Post Processing › Validation

Validating Files Before They Leave

The worst place to discover a bad file is inside somebody else's system. A feed runs cleanly for months. Then one night the source application dies halfway through its export, or a software update quietly renames a column. The transfer job — which checks nothing — dutifully delivers the damage. The upload succeeds. Every dashboard is green. Two days later the partner escalates: their batch loaded garbage, or refused to load at all. Now the conversation is about data repair, re-sends, and why nobody on your side noticed.

Validation is the pipeline stage that prevents that conversation. It is the set of checks that decides, before a file is sent, whether the file is actually fit to send. The checks look for the right structure, right size, contents that agree with themselves. The governing principle is cost. A bad file caught in your own pipeline costs one alert and one rerun. The same file caught by the partner costs their staff time, your credibility, and a cleanup project. Every stage a bad file survives multiplies the price of its eventual discovery.

This article builds the validation stage in layers, from checks that cost microseconds to checks that require knowing your data. Those layers are envelope checks, structure checks, reconciliation against a trailer record, and business sanity checks. Then the article covers the part most flows skip — what happens to the file that fails. This article is part of our pre- and post-processing series. The validation stage slots directly into the pipeline skeleton from the anatomy of a transfer pipeline, right after the acquire stage.

Validation Is Not Integrity Checking

One distinction up front, because the two get conflated constantly. Integrity checking asks "is this file the same bytes it was?" — checksums, size comparisons, the proof that a transfer or a disk did not alter anything. Validation asks "are these bytes correct in the first place?" A truncated export can transfer with perfect integrity: the checksum of the damaged file matches the damaged file beautifully. Integrity verification is essential and has its own series — see verifying transfers end to end. But integrity verification runs after the transfer and answers a different question. Validation runs before, and it is the only stage that can catch a file that is intact, complete, delivered — and wrong.

You Cannot Validate Without a Contract

Every validation check compares the file against a definition of "good," so the first job is making that definition exist outside somebody's head. Call it the file contract: a short written statement of what a valid file for this flow looks like. Name pattern. Encoding. Delimiter. The exact header row. Column count. Which fields may be empty. Whether there is a trailer record — a final line carrying totals — and what it contains. Expected size and row-count ranges.

Flat files like CSV need this most, because unlike a database, nothing enforces their structure. A CSV file is just text that everyone hopes is shaped consistently. Schema thinking for flat files means writing the schema yourself, even though no software demands it. It does not need tooling; a plain text file in the flow's folder does the job:

# contract: daily sales feed to processing partner
name:       sales_YYYYMMDD.csv          # one per business day
encoding:   UTF-8, no BOM
line ends:  CRLF
delimiter:  pipe (|), fields never quoted
header:     store_id|txn_id|txn_date|sku|qty|amount
columns:    6 on every row; txn_date = YYYYMMDD; qty and amount numeric
trailer:    T|<row count excluding header and trailer>|<sum of amount>
typical:    9,000-16,000 rows; 2-6 MB
empty ok:   sku only; all other fields required

If no spec exists — common with inherited flows — reverse-engineer one from a run of known-good files, then confirm it with whoever consumes the data. The moment the contract is written down, two things improve. Your validator has something objective to enforce. Format changes become negotiations ("the contract changes on the first of next month") instead of surprises. If the only place the rules live is the partner's parser, you are validating blind and the partner's production run is your test environment.

Layer 1: Envelope Checks — Cheap and First

Envelope checks look at the file from the outside, cost almost nothing, and catch a surprising share of real problems. Run them first, so the expensive checks never open a file that fails the basics:

  • The name matches the contract's pattern. A misnamed file is either the wrong file or a sender-side change you have not heard about yet; both deserve a stop. Parsing and enforcing names is covered in the file naming and datestamping series.
  • The size is plausible. Not just nonzero — within the contract's historical band. A 40 KB file where 4 MB is normal is a truncated export wearing a valid name. A 400 MB file is a backfill or a runaway query. Either way, a human should look first.
  • The file is complete. A file still being written passes every content check on the part that exists. Settle checks, temp-name conventions, and the other arrival defenses live in the partial-file safety series — validation assumes they already ran.
  • You have not processed it before. A resent or re-exported file may be legitimate or a mistake; detecting the repeat is the subject of the duplicate detection series. The validation stage is a natural place to consult that ledger.

Layer 2: Structure Checks — Does It Parse the Way the Contract Says?

Structure checks open the file and test its shape against the contract. These are the checks that catch the renamed column and the export that switched delimiters — mechanical failures with mechanical detection:

  • Encoding is what the contract says. If the contract says UTF-8, attempt a strict decode. An export that slipped into a legacy single-byte code page will contain byte sequences that are simply invalid UTF-8. A strict decoder refuses them. This one check prevents the classic "accented names turned to gibberish" incident downstream — more on conversions in transforming files between systems.
  • The header row matches exactly. Spelling, order, and count. Compare as a literal string against the contract. A renamed or reordered column is invisible to every other check — row counts and totals still reconcile. It silently maps data into the wrong fields downstream.
  • Every row has the right field count — every row, not just the first few. An unescaped delimiter inside a value shifts one row's field count; a broken export shifts thousands. Per-row counting finds both and names the line numbers.
  • Line endings are consistent and match the contract. A file that mixes CRLF and LF endings has usually been through a hand edit or a careless transform. In that case, strict parsers count rows differently than you do.
  • Quoting is balanced. In quoted CSV, one unclosed quote makes many parsers swallow everything to the next quote — sometimes half the file — as a single field. If the contract says fields are never quoted, a quote character appearing at all is worth a reject.

None of this needs special software. Windows admins can do the same in PowerShell — see the PowerShell automation series. On any system with standard Unix-style tools, three commands cover the core checks:

# strict decode: exits non-zero if the file is not valid UTF-8
iconv -f UTF-8 -t UTF-8 sales_YYYYMMDD.csv > /dev/null

# header row: compare line 1 to the contract, byte for byte
head -n 1 sales_YYYYMMDD.csv | diff - header.expected

# field counts: list every row that does not have 6 pipe-separated fields
awk -F'|' 'NF != 6 { print NR ": " NF " fields" }' sales_YYYYMMDD.csv

Each command exits zero on success and non-zero (or prints offending line numbers) on failure. That is exactly the shape a pipeline wants, because the job wrapping these checks can abort on any non-zero result.

Layer 3: Reconciliation — Does the File Agree With Itself?

Reconciliation checks use redundancy the producer built into the file. The most common form is the trailer record. The source system writes, as the last line, how many data rows it produced and what key columns sum to — something like T|12417|8551023.44. The file now carries its own receipt, and validation cashes it:

  • Row count versus trailer count. Count the data rows; compare to the trailer's claim. This is the single best truncation detector in existence. An export that died midway produces fewer rows than its trailer promises — or no trailer at all, which is equally damning. A missing trailer should always be a hard reject.
  • Column totals versus trailer totals. Sum the amount column; compare to the trailer's figure. This catches subtler damage that row counting misses. That includes duplicated blocks of rows, a numeric column shifted by a delimiter fault, corruption in the middle of an otherwise complete file.

If the producer cannot write a trailer, the same redundancy can travel as a control file. This is a tiny sidecar (sales_YYYYMMDD.ctl) holding the counts and totals, produced by the same run that produced the data. Sidecar conventions, and how they extend to checksums and multi-file manifests, are covered in checksum files and manifests. Either way, the principle is the same: the producer states its intent, and validation refuses any file whose contents disagree with the statement.

Layer 4: Business Sanity — Structurally Perfect Nonsense

A file can pass every check so far and still be garbage. Every row parses, the trailer reconciles — and every amount is zero, because a source table was empty. Or the row count is 214 on a day that normally produces 12,000, because an upstream filter broke. Or the dates inside the file are yesterday's, because the export re-ran against a stale snapshot. Structure checks cannot see this; only expectations about the business can. The practical sanity set:

  • Volume within its historical band — row count and total value each between, say, half and double their recent norms. The band is written into the contract rather than guessed each time.
  • Dates inside the file match the business date in the file name. A mismatch means a stale or misdated export — arguably the most damaging bad file of all, because it looks completely healthy.
  • Required fields are actually populated — not just present but non-empty in at least the expected proportion of rows.
  • Nothing is leaving that should not. Validation is the natural checkpoint for data-protection rules: are there columns in this file that the destination has no need for? Trimming and disguising personal data is its own discipline — see data minimization for transfers and masking and pseudonymization. But the check that the rules were applied belongs here, before the file leaves.

One honest complication: sanity checks, unlike structure checks, can be legitimately violated. A holiday sale really does triple the row count. So decide per check whether a breach is a hard failure (stop, reject) or a warning (send, but flag for a human to glance at). A useful default: reconciliation and date checks are hard failures. Volume-band checks start as warnings and get promoted to hard failures once you trust the bands. What you must not do is let warnings scroll by unread — a warning nobody reads is a check you do not have.

Reject Handling: Where Bad Files Go

A validator that only says "no" is half-built. The other half is what happens to the file that failed, and the rules are strict because everything downstream depends on them:

  • The file leaves the flow. Move it — do not copy it — to the flow's error folder, out of the path of the next run. A rejected file left in the inbox will be retried forever or, worse, picked up by a colleague "just clearing the backlog."
  • Nothing gets sent. No best-effort delivery of the rows that looked fine. Partial deliveries create reconciliation puzzles on both sides that cost far more than a late, complete file.
  • The reject leaves a report. Next to the quarantined file, write a small plain-text report saying exactly what failed and what was expected. The person who fixes the file is usually not the person who runs the pipeline. This report is the ticket you are pre-writing for them.

A reject report worth copying looks like this:

REJECT REPORT   sales_YYYYMMDD.csv
when:      Mar 14 02:11
stage:     validate (after acquire; nothing was sent)
moved to:  error\sales_YYYYMMDD.csv

FAILED  header row
        expected: store_id|txn_id|txn_date|sku|qty|amount
        found:    store_id|transaction_id|txn_date|sku|qty|amount

FAILED  field count
        4,211 of 12,408 rows have 7 fields (expected 6)
        first offending line: 3

PASSED  name pattern, size band, encoding (UTF-8),
        line endings (CRLF), trailer reconciliation

action: feed owner notified by mail; rerun is safe after
        the source export is corrected

Notice the last line. Because validation failed before anything was sent, the rerun after the fix is completely safe — there is no half-delivered state to untangle. Also notice what a reject is not: a reason to retry. A network timeout deserves a retry; a wrong header row will be exactly as wrong on the tenth attempt. In the vocabulary of the retry and error handling series, validation failures are permanent failures. The correct response is quarantine plus a human, never the retry loop.

Remember: a validation failure is a permanent failure. Retrying cannot fix the file — it can only delay the alert. Quarantine the file, write the report, notify the owner, and stop.

Wiring Validation Into the Job

Validation belongs after acquire and before any packaging. That ordering matters more than it looks. Once a file is zipped or encrypted, its contents are opaque — a checker cannot count columns inside ciphertext. So content checks must run on the plain file, with compression and encryption applied only to files that passed. (The packaging sequence is covered in the compression article). Validate what you will send, as the last look inside before the box is sealed.

Mechanically, a validator is just a script with an exit code: zero for pass, non-zero for reject, report written as a side effect. That shape drops into any automation. In Sysax FTP Automation, a wizard-generated transfer task can be extended in the script editor — with line-by-line debugging to test the addition. That way, your validation script runs as a pre-transfer step. The transfer only proceeds on success, and the tool's email notification carries the failure to the feed owner. The checks stay yours, because only you know your data; the automation contributes the scheduling, the sequencing, and the loud failure.

And validation is not only a sender's discipline. When you are the receiving side, the same layers apply to what partners send you — ideally the moment a file arrives. If your transfer server is Sysax Multi Server, event triggers in the Pro and Enterprise editions can launch your intake-validation script as soon as an upload completes. That way, a bad inbound file is quarantined minutes after arrival instead of discovered by the nightly import.

The Pre-Flight Checklist

The full stage, compressed into the list to pin above your desk. For each outbound file, in order, stopping at the first hard failure:

  1. Name matches the contract pattern; business date in the name is the expected one.
  2. File is complete (settle check passed) and size is within the contract's band.
  3. Not a duplicate of a file already processed.
  4. Encoding survives a strict decode in the contract's encoding.
  5. Header row matches the contract exactly.
  6. Every row has the contract's field count; line endings are consistent.
  7. Row count matches the trailer or control file; totals match.
  8. Sanity: volumes in band, dates coherent, required fields populated, nothing leaving that minimization rules forbid.
  9. On any failure: move to error, write the reject report, notify the owner, send nothing.

The Version to Tell a Colleague

Validation is the cheapest moment you will ever have to catch a bad file, because nothing has been spent on it yet. That means no bandwidth spent, no partner attention, no downstream damage. Write the file contract down, then enforce it in layers: envelope, structure, reconciliation, sanity. Reject loudly, quarantine cleanly, leave a report a stranger can act on, and never retry a file that failed on content. A transfer pipeline with real validation almost never has to apologize to a partner — and that, not the checks themselves, is the point.

Where to go next in this series: the pipeline anatomy shows where this stage sits among its six siblings. The article on transforming files between systems covers fixing the fixable instead of rejecting it. The reject-folder pattern connects to the wider error design in our watch folders series.

Frequently Asked Questions

Should I validate before or after compressing and encrypting?
Before. Once a file is zipped or encrypted its contents are opaque, so column counts, encodings, and totals can no longer be checked. Validate the plain file, then package only the files that passed, then use integrity checks (sizes and checksums) to confirm the packaged bytes traveled intact.
What is a trailer record?
A trailer record is a final line the producing system writes into the file. The record states how many data rows the system produced and often what key columns sum to. Validation counts and sums the actual rows and compares; any disagreement means truncation, duplication, or corruption. If the producer cannot write a trailer, the same figures can travel in a small sidecar control file.
My partner never gave me a file spec. What do I validate against?
Write the contract yourself from a run of known-good files. Capture the header, delimiter, encoding, typical sizes, and row-count ranges, then ask the partner to confirm it. Even an unconfirmed contract beats none — it turns "the file looks weird" into specific, checkable rules, and it flushes out undocumented assumptions fast.
Should a failed sanity check block the file or just warn?
Decide per check, in advance. Reconciliation and date mismatches should always block, because they indicate real damage. Volume-band breaches can start as warnings — legitimate business spikes happen — and be promoted to hard failures once the bands are trusted. The one rule: a warning someone must read, or it is not a check at all.
How do I validate very large files without slowing the pipeline down?
Order the checks by cost: name, size, and duplicate checks are instant and reject the worst cases before the file is ever opened. The content checks — encoding, field counts, totals — can all be done in a single streaming pass that reads the file once. For most feeds, that takes seconds. Only after that single pass does the file earn compression, encryption, and bandwidth.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.