Home › Topics › ASCII & Encoding › Detection

Detecting Mode and Encoding Corruption Early

The worst place to discover a mangled file is the last one. That is the accounting import that fails at month end, or the archive that will not open the morning a customer needs it. Or it is the loader that quietly inserts José into the master record. By then the transfer that caused it is hours or days old, and the source file may be gone. The only evidence is a downstream error message written for a different audience. Every member of the silent-mangler family — ASCII mode, line endings, encodings, byte-order marks — is cheap to detect at the moment a file arrives. It is expensive to detect anywhere else.

This article is about moving detection to that moment. It lays out the checks in order of cost, from a byte count that takes no time at all to a cryptographic hash. It explains what each one catches and what it misses, and shows the commands on Linux and Windows. It ends with a complete intake script and the question of what a job should do when a check fails. It is part of our ASCII vs Binary and Encoding Corruption series; it assumes you have read the ASCII/binary trap, which explains the damage these checks are looking for.

The Detection Ladder: Cheapest Tell First

Four kinds of check are available, and they form a ladder. Each rung costs more than the one below it and tells you more. The trick is to run the cheap ones on every file, and to reserve the expensive ones for the files and flows that justify them.

Check Cost Catches Misses
Size delta Free — a directory entry ASCII-mode conversion, truncation, empty files, and — with one extra count — which of those it was Anything that keeps the byte count: encoding mix-ups, flipped bytes
First and last bytes Microseconds — a few bytes read Wrong file type, ASCII-mangled signatures, unexpected BOM, UTF-16, truncated archives and PDFs Damage in the middle of the file
Content sanity and format validation One pass over the file Mixed line endings, invalid UTF-8, mojibake, corrupt archives, malformed CSV, JSON, and XML Valid files with wrong data
Hash comparison One pass, plus a hash from the sender Any change to any byte, definitively Damage present before the sender hashed; says "changed", not "how"

The ladder does not start with hashing, even though hashing is the strongest check. A hash needs cooperation from the sender, and when it fails it says only that the file is different. The lower rungs are free, need nothing from the other side, and — in the case of the size delta — tell you what went wrong. Use them first and always.

Size Deltas: The Free Check

Every transfer has an expected size. The sending job logged it; an FTP SIZE command returns it; a manifest or control file lists it. At worst, the source file is still there to ask. Compare it with the exact byte count of what arrived — not the rounded "1.0 MB" a file manager shows, but the number:

# Linux
$ stat -c %s report.zip
1052679

# Windows PowerShell
PS> (Get-Item report.zip).Length
1052679

Three outcomes cover almost every case. Equal sizes rule out ASCII-mode conversion and truncation — the two most common mangles — though not encoding problems, which keep the byte count. A size that is much smaller, zero, or a suspiciously round number means truncation: the transfer stopped early. That is a connection problem that belongs to the size-stability and settle-check family of fixes. A size that is slightly different — a few bytes per kilobyte, larger or smaller — is the fingerprint of ASCII mode. One more count turns suspicion into proof:

# On the sending side: how many line-feed bytes does the original contain?
$ tr -cd '\n' < report.zip | wc -c
4103

# Sent 1048576 bytes; received 1052679.  1052679 - 1048576 = 4103.
# The file grew by exactly its line-feed count: every 0A became 0D 0A.
# Verdict: transferred in ASCII mode and received on Windows.

The reverse case — a Windows-origin file received on Linux — shrinks by the number of 0D 0A pairs in the original, which grep -c $'\r$' file counts. Building this arithmetic into an intake check means the alert can say "ASCII mode suspected: delta equals line-feed count" instead of "size mismatch". The person who reads it knows which client setting to go and fix.

First Bytes: Signatures and Marks

Most binary formats begin with a fixed signature, often called a magic number: a few bytes that identify the type. Reading them costs nothing and catches three things at once. It catches a file that is not the type its name claims, or a file whose signature has been mangled by ASCII mode. It also catches a text file that arrived with a byte-order mark or in UTF-16 when the contract said plain UTF-8. The ones worth memorizing:

Bytes (hex)                 Type                      Note
89 50 4E 47 0D 0A 1A 0A     PNG image                 contains CR LF and LF on purpose
50 4B 03 04                 ZIP (also docx, xlsx, jar)  "PK"
1F 8B                       gzip
25 50 44 46 2D              PDF                       "%PDF-"
FF D8 FF                    JPEG
37 7A BC AF 27 1C           7-Zip archive
4D 5A                       Windows executable        "MZ"
EF BB BF                    UTF-8 byte-order mark     text follows
FF FE                       UTF-16 little-endian      every other byte will be 00

# Read the first eight bytes
$ xxd -l 8 logo.png
00000000: 8950 4e47 0d0d 0a1a                      .PNG....
PS> Format-Hex logo.png -Count 8

# Or ask file(1), which knows hundreds of signatures
$ file logo.png report.zip customers.csv
logo.png:      data
report.zip:    Zip archive data
customers.csv: UTF-8 Unicode (with BOM) text, with CRLF line terminators

Look closely at the xxd output: the fifth and sixth bytes are 0d 0d where a healthy PNG has 0d 0a. That is an ASCII-mode transfer received on Windows, visible in eight bytes. It is why file no longer recognizes the image and reports plain data. The PNG signature was designed to fail this way. Other formats are less obliging, which is where the size delta and the format validators come in.

The end of a file is worth a glance too. A zip archive keeps its directory at the end, marked 50 4B 05 06 shortly before the final byte; a PDF ends with %%EOF. A file that starts correctly but ends abruptly is truncated, and tail -c 32 file | xxd shows it in one line. The signature check is also the cheap defense against a file that is not what it claims to be at all, which intake validation and safety covers from the security side.

Byte-Level Sanity Checks for Text Files

Text files have no signature, but they have expectations, and each of the silent manglers leaves a byte pattern that violates one. An intake check for a flow that expects "UTF-8, no BOM, LF endings" can test all of the following in a single pass:

  • Byte-order mark present: the first three bytes are EF BB BF. Reject or strip according to the contract.
  • Invalid UTF-8: iconv -f utf-8 -t utf-8 file >/dev/null exits non-zero and names the offending position. Almost always a legacy code page.
  • Replacement characters: the bytes EF BF BD anywhere in the file. Some earlier step met bytes it could not decode and substituted a placeholder; the original characters are already gone.
  • Mixed line endings: the count of lines ending in CR is neither zero nor equal to the total line count.
  • Lone carriage returns: the total number of CR bytes exceeds the number of lines ending in CR. A CR sits in the middle of a line, which is data in a fixed-width file or corruption in anything else.
  • NUL bytes: any 00 in a text file means UTF-16, a binary file with the wrong name, or a truncated write padded with zeros.
  • Mojibake tells: à or  followed by a symbol such as ©, ¼, or ¤ almost never occurs in genuine text. It almost always means UTF-8 was decoded as a single-byte code page. Treat it as a warning rather than a hard reject, since a few languages use those letters legitimately.
f=customers.csv
[ "$(head -c 3 "$f" | od -An -tx1 | tr -d ' ')" = "efbbbf" ] && echo "BOM present"
iconv -f utf-8 -t utf-8 "$f" >/dev/null 2>&1 || echo "not valid UTF-8"
LC_ALL=C grep -q $'\xEF\xBF\xBD' "$f" && echo "replacement characters present"
cr_lines=$(grep -c $'\r$' "$f"); lines=$(wc -l < "$f"); cr_total=$(tr -cd '\r' < "$f" | wc -c)
[ "$cr_lines" -ne 0 ] && [ "$cr_lines" -ne "$lines" ] && echo "mixed line endings"
[ "$cr_total" -gt "$cr_lines" ] && echo "lone CR bytes present"
[ "$(tr -cd '\000' < "$f" | wc -c)" -gt 0 ] && echo "NUL bytes present"

Each line prints a message only when its test trips, so a clean file produces no output. In PowerShell the same tests are a few lines against [IO.File]::ReadAllBytes, and the PowerShell file operations guide has the byte-array idioms. The line-ending arithmetic is explained in line endings across systems, and the encoding tests in character encodings for admins.

Structural Validation: Let the Format Check Itself

Many formats carry their own integrity checks, and the tools that read them will run those checks on request without extracting or importing anything. This is the most thorough test short of a hash, and unlike a hash it needs nothing from the sender:

$ unzip -t report.zip                 # tests every member's checksum
    testing: orders.csv              OK
No errors detected in compressed data of report.zip.

$ unzip -t report_ascii.zip           # the same archive after an ASCII-mode transfer
  End-of-central-directory signature not found.  Either this file is not
  a zipfile, or it constitutes one disk of a multi-part archive. ...

$ gzip -t export.gz && echo ok         # gzip has a trailing checksum
$ tar -tzf backup.tgz >/dev/null      # lists the archive; fails on damage

$ xmllint --noout invoice.xml         # silent when well-formed
$ python3 -m json.tool payload.json >/dev/null
Expecting ',' delimiter: line 2 column 1 (char 7)

# CSV: every row should have the same number of fields
$ awk -F, '{print NF}' orders.csv | sort -u
3

The CSV check deserves a note because CSV has no validator of its own. Listing the distinct field counts across all rows should produce exactly one number. Two or more mean a broken row, an unquoted comma, or a line-ending problem that glued rows together. Where the sender includes a trailer record with a row count, compare it to wc -l as well. Those conventions belong in the flow's file interface contract. The discipline of checking a file before acting on it is the subject of validating files before sending.

Hashes: The Definitive Check

A hash is a short fixed-length fingerprint computed from every byte of a file; change one byte and the fingerprint changes completely. How hashing works, which algorithms to use, and how to distribute hash files and manifests are covered in hashing explained and checksum files and manifests. Here the point is narrower: as a detector of mode and encoding damage, a hash comparison is perfect and blunt. The five-line file from the first article in this series shows both qualities:

# Sender (Linux) publishes the hash of the 65-byte original
$ sha256sum orders.csv
f3fb57150b3dfc9c9a4fe032b5fd53c5b08f88de008e10c436e7d316dc935abc  orders.csv

# Receiver (Windows) hashes the 70-byte copy that arrived in ASCII mode
PS> (Get-FileHash orders.csv -Algorithm SHA256).Hash
5AEB673708FA854103CEF43A411C25EAA0CFC47CAB545A7B8B3CE8A017757CE7

# Or let the tool compare against the sender's list
$ sha256sum -c orders.sha256
orders.csv: FAILED
sha256sum: WARNING: 1 computed checksum did NOT match

Five added bytes, a completely different fingerprint: that is the perfect part. The blunt part is that the hash says nothing about why. Pair it with the size delta and it does. Then "hash mismatch, and the file is larger by exactly its line-feed count" is a diagnosis, where "hash mismatch" alone is a ticket. There are two limits to keep in mind. A hash computed by the sender after the damage was done will match a damaged file happily. So the sender must hash the finished file it meant to send. And a flow that deliberately converts line endings must hash on the same side of that conversion at both ends. Integrity in automation works through where the hash step sits in a job.

Where the Checks Live and What Failure Does

The checks belong in the transfer job, not in the downstream application. Use a post-download step on inbound files, before anything else touches them, and a pre-upload step on outbound files. A scheduled-transfer tool with pre- and post-processing hooks and error handling, such as Sysax FTP Automation, lets a script like the one below run as part of the job that fetched the file. The tool lets a non-zero exit stop the job rather than pass the file along. On the server side, an activity log — the kind Sysax Multi Server keeps — lets you match a failed check back to the session and client that delivered the file.

What should a failed check do? Three rules that avoid the common mistakes:

  1. Stop, and keep the evidence. Move the file to a quarantine folder with its original name, never delete it, and never let a later run of the job overwrite it. The poison-file and dead-letter pattern is the model.
  2. Alert with the tell, not the symptom. "ORDERS-IN: size grew by 4103 bytes = line-feed count, ASCII mode suspected, session from 203.0.113.7 at 02:10" leads straight to the fix. "Validation failed" leads to a meeting.
  3. Never auto-repair. Stripping a BOM or converting line endings automatically hides the fact that the sender's configuration is wrong. The next file will be wrong in a way the repair does not cover. Fix the source; the check exists to tell you it needs fixing.

Remember: a check that runs after the downstream system has already consumed the file is a post-mortem. Put the checks in the job that moves the file, and make them fail loudly. Make the alert name the mangler — size delta equal to line-feed count, BOM present, invalid UTF-8. That way, the fix is obvious to whoever reads it.

A Complete Intake Check

The script below combines the rungs of the ladder for a flow that expects a UTF-8 CSV without BOM and with LF line endings, delivered with a sender-side hash. It exits non-zero on the first failure and prints a one-line reason, which is all an alerting system needs:

#!/bin/bash
# intake-check.sh FILE EXPECTED_SIZE [SHA256]
f="$1"; expected="$2"; want_hash="$3"
fail() { echo "REJECT $f: $1"; exit 1; }

[ -s "$f" ] || fail "empty or missing"
actual=$(stat -c %s "$f")
if [ "$actual" -ne "$expected" ]; then
    lf=$(tr -cd '\n' < "$f" | wc -c); crlf=$(grep -c $'\r$' "$f")
    [ $((actual - crlf)) -eq "$expected" ] && fail "size +$crlf = CRLF count: ASCII mode suspected"
    [ $((actual + lf))  -eq "$expected" ] && fail "size -$lf = LF count: ASCII mode suspected"
    fail "size $actual, expected $expected"
fi
[ "$(head -c 3 "$f" | od -An -tx1 | tr -d ' ')" = "efbbbf" ] && fail "BOM present"
iconv -f utf-8 -t utf-8 "$f" >/dev/null 2>&1 || fail "invalid UTF-8"
[ "$(grep -c $'\r$' "$f")" -eq 0 ] || fail "CRLF line endings"
[ "$(awk -F, '{print NF}' "$f" | sort -u | wc -l)" -eq 1 ] || fail "inconsistent field count"
if [ -n "$want_hash" ]; then
    [ "$(sha256sum "$f" | cut -d' ' -f1)" = "$want_hash" ] || fail "hash mismatch"
fi
echo "OK $f"

Adapt the middle per flow: drop the CSV line for other formats, add unzip -t for archives, invert the line-ending test for a flow that legitimately uses CRLF. The structure — size, then bytes, then format, then hash — stays the same, and so does the habit of naming the mangler in the rejection message.

The Version to Tell a Colleague

Detect at intake, not downstream. Compare the exact byte count first: a small difference equal to the file's line-feed count is ASCII mode, a large one is truncation. Read the first eight bytes for a signature, a BOM, or UTF-16's zero bytes. Run the format's own validator — unzip -t, xmllint, a field count for CSV. Run the text sanity checks for invalid UTF-8, replacement characters, and mixed line endings. Use a hash to settle the question definitively, paired with the size delta to explain it. Quarantine on failure, alert with the tell, and never repair automatically.

The last article in this series, never get bitten again, turns these detectors into a prevention checklist for clients, servers, jobs, and partners. For the broader layered method — connectivity, authentication, permissions, protocol, and finally content — see the systematic troubleshooting series.

Frequently Asked Questions

If I already verify hashes, do I need the other checks?
A hash tells you the file changed, not why, and it depends on the sender having hashed the right thing. The size delta and first-byte checks are free, need nothing from the sender, and name the cause. Run them first; keep the hash as the final word.
How can a size difference tell me it was ASCII mode?
ASCII mode changes exactly one byte per line ending: an LF becomes CR LF or the reverse. So the received size differs from the original by precisely the number of line feeds (or CR LF pairs) the original contained. Count them with tr -cd '\n' | wc -c and compare; an exact match is proof.
Can I detect encoding problems without knowing the source encoding?
Partly. You can prove a file is or is not valid UTF-8 and spot a byte-order mark without any outside knowledge. You can also detect replacement characters and NUL bytes without that knowledge. Telling one single-byte code page from another needs information about the producer, which is why the contract should state the encoding.
Should the intake check fix the file automatically?
No. Automatic repair hides a misconfigured sender, and the next file may be damaged in a way the repair does not handle. Quarantine the file, alert with the specific tell, and fix the client, job, or partner setting that caused it. Deliberate, documented conversion steps are different — those are part of the flow's design.
Where should these checks run — sender or receiver?
Both, when you control both. The sender validates before upload so it never sends a bad file; the receiver validates after download so it never processes one. When you control only one side, run the checks there and put the expectations — encoding, line endings, hash — into the partner agreement.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.