Home › Topics › ASCII & Encoding › Prevention

Never Get Bitten Again: A Prevention Checklist

Every problem in this series has been solved at least once in your organization already. Someone found the ASCII-mode transfer, someone stripped the byte-order mark, someone renamed the file with the question mark in it. The trouble is that each fix lived in one person's memory, one script, or one ticket. The next new flow, new partner, or new laptop started from zero. The silent manglers keep biting not because they are hard to understand but because the understanding never gets turned into defaults.

This article is that conversion. It collects the prevention measures from the rest of our ASCII vs Binary and Encoding Corruption series into a checklist. The checklist is organized by where the setting lives — the client, the server, the job, and the partner agreement. It closes with a small test file built to expose every mangler at once. That way, a new flow can prove it preserves bytes before it carries anything that matters. Each item names the article that explains it, so the list stays short.

The Principle Behind Every Item

Four ideas generate the whole checklist, and if you remember only these, you can rebuild the rest.

  1. The transport copies bytes. Every transfer runs in binary mode or over a protocol with no text mode, so the file that arrives is byte-for-byte the file that left. Nothing is ever converted "on the way".
  2. Every conversion is a named step. If a flow genuinely needs a line-ending or encoding change, it happens in a step with a name, a place in the job, and a line in the log. It never happens as a side effect of a client setting, an editor, or a shell idiom.
  3. Every expectation is written down. Transfer mode, encoding, byte-order mark, line endings, filename rules, and hashing are stated per flow, in the job definition and in the partner agreement. That way, nobody guesses.
  4. Every flow is tested with a file designed to break. Before a new flow, client, server, or partner carries real data, it carries a canary file whose hash proves the path is clean.

Clients: Defaults That Cannot Be Wrong

Most incidents in this family start on a client, because the client chooses the transfer mode and the filename encoding. The goal is a client configuration where the safe choice is the default and the unsafe one takes deliberate effort:

  • Binary mode is the default, not "auto". Graphical clients that pick ASCII or binary per file from an extension list should be set to a fixed default of binary. The list will be wrong for someone's file eventually; a fixed setting cannot be. See the ASCII/binary trap.
  • Text mode is off, even over SFTP. Some clients can convert line endings locally before or after an SFTP transfer. Turn that off and do conversions as a job step.
  • The Windows ftp.exe client is retired or wrapped. It defaults to ASCII, has no UTF-8 filename support, and sends passwords in the clear. If a script must still use it, the command file begins with binary. In that case, the script is on the list of things to replace, as the honest look at built-in clients argues.
  • Command-line and library defaults are left alone. curl and lftp default to binary; nobody adds --use-ascii or -a. Library code uses the binary call — storbinary, not storlines, in Python's ftplib — and a code review checks for it.
  • UTF-8 filenames are forced on. Where the client has a charset setting for FTP or SFTP names, it is set to UTF-8 rather than "autodetect". That way, a server that supports UTF-8 without advertising it still gets the right bytes. See filename encoding problems.
  • The configuration is deployed, not described. A documented setting that each person applies by hand is a setting that half the machines lack. Ship the client with these values preset, as described in deploying standard client configurations. Choose a client that can be preset in the first place — choosing the standard client covers that selection.

For scripted FTP, a two-line habit makes the mode visible in every log. Ask for verbose output and look for the TYPE I line before the first transfer; if it is missing or says TYPE A, the script is wrong:

$ curl -v -T report.zip -u transfer_svc ftp://ftp.example.com/inbound/ 2>&1 | grep -E '^> (TYPE|STOR)'
> TYPE I
> STOR report.zip

Servers: Make the Safe Path the Easy Path

A server cannot stop a client from asking for ASCII mode — the protocol gives the client that choice. But it can remove the opportunity, make names safe by default, and leave the evidence that turns a mystery into a five-minute fix:

  • SFTP is offered and preferred. SFTP has no transfer mode in practice, so a flow moved from FTP to SFTP cannot be bitten by ASCII mode again. Where a server such as Sysax Multi Server offers FTP, FTPS, SFTP, and HTTPS side by side, new flows are put on SFTP or HTTPS by default. In that setup, FTP is kept only for documented exceptions, in line with the FTP retirement plan.
  • Where FTP remains, UTF-8 names are advertised. The server lists UTF8 in its FEAT reply and honors OPTS UTF8 ON, so modern clients negotiate correct names automatically.
  • The service runs under a UTF-8 locale. On Linux, an SFTP or FTP service started under a legacy locale can store names in that locale's code page. Check the environment the service actually starts with, not the one in your login shell.
  • The platform's name rules are documented for partners. A Windows server refuses colons and reserved names and folds case; a Linux server counts name length in bytes. Write the rules that apply to your server into the partner agreement so that rejections are expected rather than surprising.
  • Activity logging is on and retained. When an intake check rejects a file, the session log is what links it to a client, an account, and a time. Logging is the difference between "someone sent us a bad file" and "the nightly job on host X still runs the old script".
  • Inbound folders have a quarantine sibling. A rejected file needs somewhere to go that is not the folder downstream jobs read from and not the recycle bin.

Jobs: A Template Every New Flow Copies

Automation multiplies whatever it is given. A job created by copying an old one inherits the old one's assumptions, which is how a single ASCII-mode script becomes a dozen. The fix is a template that new flows start from, with every relevant setting present and explicit, even when the value is the obvious one:

Flow name:        ORDERS-IN
Direction:        inbound, partner -> our server
Protocol:         SFTP (no transfer mode); if FTP/FTPS: mode = binary, always
Content encoding: UTF-8, no byte-order mark
Line endings:     LF; CRLF files converted by step "normalize-eol" (receiver side)
Filename rule:    ^[a-z0-9_-]+\.csv$ ; non-conforming names -> quarantine
Conversion steps: normalize-eol (dos2unix -n), after intake-check passes
Validation step:  intake-check.sh (size, BOM, UTF-8, line endings, field count, hash)
Hash:             SHA-256 from partner manifest, checked before conversion
On failure:       stop job, move file to /srv/quarantine/orders-in/, alert with reason
Canary verified:  YYYYMMDD by <name>; re-run after any client, server, or job change

Three items in the template deserve emphasis. The conversion step has a name and a side, so nobody wonders whether the partner or the receiver converts. The validation step runs before the conversion, so the hash is checked against the bytes the partner actually sent. And the canary line records that the flow was proven clean, and when. That is the evidence that lets a later incident be narrowed to "what changed since".

Inside the scripts the template calls, the same explicitness applies. Every file open in PowerShell or Python states its encoding rather than trusting the platform default. A reviewer rejects Get-Content | Set-Content on data files because it rewrites line endings as a side effect. And conversions call dos2unix or iconv by name so they appear in the log. Line endings across systems and its companion article on character encodings give the commands. Suppose the job runs in a scheduled-transfer tool with pre- and post-processing steps and error handling — Sysax FTP Automation is one. Then the validation and conversion steps become part of the job definition rather than a separate script someone has to remember. In that setup, a failed check stops the job rather than being logged and ignored. The pipeline shape those steps fit into is described in anatomy of a transfer pipeline.

Partners: The Clauses in the Agreement

You do not control a partner's client, server, or job template. What you control is the agreement, and five short clauses cover this whole family. They belong next to the naming, timing, and format rules in the flow's file interface contract. They are exchanged during onboarding, before the first real file, following the partner onboarding runbook:

1. Transfer mode.  All transfers use SFTP. If FTP or FTPS is used, every
   transfer is made in binary mode (TYPE I). Files transferred in ASCII
   mode are considered corrupt and will be rejected.

2. Content encoding.  Text files are UTF-8 without a byte-order mark.
   Files in any other encoding, or with a byte-order mark, are rejected.

3. Line endings.  Text files use LF line endings only. (Or: CRLF only.)
   Files with mixed line endings are rejected.

4. Filenames.  Names use only lowercase ASCII letters, digits, underscore,
   hyphen, and a single dot before the extension, and are at most 64
   characters. Names outside this set are rejected.

5. Integrity.  Each delivery includes a SHA-256 manifest computed over the
   files as sent. A file whose hash does not match is rejected and the
   sender is notified with the reason.

The clauses are deliberately blunt: they define rejection, not preference. A partner who reads "should be UTF-8" sends whatever their export tool produces; a partner who reads "will be rejected" tests before they send. The partner exchange standards article discusses how to present such clauses so they read as shared protection rather than demands. The same five items, phrased as expectations, apply to internal flows between teams.

The Canary File: Proof the Pipeline Preserves Bytes

Settings can drift and agreements can be misread; a test is the only thing that proves a path is clean. A canary file is a small file built so that every mangler in this series changes it in a detectable way. This one is twenty-four bytes long:

Offset  Bytes                     What it tests
00      89 50 4E 47 0D 0A 1A 0A   PNG signature: CRLF, a DOS end-of-file byte, and LF
08      63 61 66 C3 A9 0D 0A      "café" in UTF-8, then CRLF
0F      63 61 66 E9 0A            "café" in Windows-1252 (invalid UTF-8), then LF
14      0D 00 FF 1A               lone CR, NUL, the highest byte value, another 1A

Size:    24 bytes
SHA-256: 0c2ead7572ff0c04b4421607c08326128d1372992fe444e11c59f4bde1a708a8

# Create it on Linux
$ echo 89504e470d0a1a0a636166c3a90d0a636166e90a0d00ff1a | xxd -r -p > canary.bin
$ sha256sum canary.bin

# Create it in PowerShell
PS> $hex = "89504e470d0a1a0a636166c3a90d0a636166e90a0d00ff1a"
PS> $bytes = [byte[]](0..($hex.Length/2-1) | ForEach-Object { [Convert]::ToByte($hex.Substring($_*2,2),16) })
PS> [IO.File]::WriteAllBytes("canary.bin", $bytes)
PS> (Get-FileHash canary.bin -Algorithm SHA256).Hash

Send the canary through the flow exactly as a real file would go — same client, same account, same job, same folders, same post-processing — and hash what arrives. A matching hash proves the whole path preserved every byte. A mismatch tells you which mangler is present. A size of 28 bytes means ASCII mode received on Windows. A size of 22 bytes means ASCII mode received on a Unix-family system. A still-24-byte file with a different hash means something rewrote content without changing length, which points at an encoding conversion step. Run it at partner onboarding, after any change to a client, server, or job, and on a schedule for flows that matter. A quarterly canary through every production path is cheap insurance. Testing changes against a staging copy of the flow before they reach production is the subject of the testing and staging series.

The canary tests content, not names. To test the filename path as well, send a second copy under a name that contains an accented character — canary_café.bin. Check that the name arrives byte-for-byte, using the ls -b or PowerShell code-point listing from filename encoding problems. If your policy forbids such names in production, that copy should be rejected by the intake check, which is itself a useful test.

Remember: a canary that passes today proves today's path. Record the date and the hash in the job template, and treat any change to a client, server, or job step as a reason to send it again. Most recurrences of this family come from a change nobody thought was related.

The Checklist

The whole article in one copyable list, for the runbook, the onboarding template, or the review meeting.

CLIENTS
[ ] Default transfer mode is binary; "auto" / extension-based mode disabled
[ ] Client-side text mode off, including over SFTP
[ ] ftp.exe retired; any survivor's command file starts with "binary"
[ ] Scripts and libraries use binary calls; no --use-ascii, -a, or storlines
[ ] Filename charset forced to UTF-8
[ ] Settings deployed centrally, not applied by hand

SERVERS
[ ] SFTP or HTTPS is the default for new flows; FTP only by documented exception
[ ] FTP server advertises UTF8 and honors OPTS UTF8 ON
[ ] Service runs under a UTF-8 locale (check the service's environment)
[ ] Platform name rules (case, forbidden characters, length) documented for partners
[ ] Activity logging on and retained long enough to trace a rejected file
[ ] Quarantine folder exists beside every inbound folder

JOBS
[ ] Every job starts from the template; every field filled, none inherited
[ ] Conversion steps named, placed on one side, and visible in the log
[ ] intake-check (size, signature, BOM, UTF-8, line endings, format, hash) before processing
[ ] Hash checked against the bytes as delivered, before any conversion
[ ] Failure stops the job, quarantines the file, alerts with the specific reason
[ ] Scripts state encodings explicitly; no Get-Content | Set-Content on data files

PARTNERS
[ ] Five clauses in the agreement: mode, encoding, line endings, filenames, hash
[ ] Clauses define rejection, not preference
[ ] Canary exchanged at onboarding, hash confirmed both ways

PROOF
[ ] Canary sent through every new or changed path; date and hash recorded
[ ] Periodic canary through production flows that matter

The Version to Tell a Colleague

The silent manglers are prevented by defaults, not by vigilance. Clients transfer in binary and send UTF-8 names because they were shipped that way. Servers offer SFTP, advertise UTF-8, log every session, and have a quarantine folder. Jobs are copied from a template that names every conversion and validates before it processes. Partner agreements say what will be rejected. And every path carries a twenty-four-byte canary before it carries anything real. The hash is written down so the next change can be checked against it.

If this article is your entry point to the series, the mechanism behind every checklist item is in the ASCII/binary trap. The checks the job template refers to are built step by step in detecting corruption early. For the fundamentals of the hash that anchors the canary test, see hashing explained.

Frequently Asked Questions

If we move everything to SFTP, can we skip the rest of the checklist?
SFTP removes the ASCII-mode item and most filename encoding trouble, which is a big win. It does nothing about content encodings, byte-order marks, line endings, or client-side text modes, because those happen in the programs that write and read files. Keep the job template, the partner clauses, and the canary.
Why is the canary file so small?
It only needs to contain each byte pattern that a mangler would change. Those are CR LF, a lone LF, a lone CR, the DOS end-of-file byte, valid and invalid UTF-8, a NUL, and the highest byte value. Twenty-four bytes cover all of them, and a tiny file can be sent through every path in seconds and compared by hand if needed.
Our partner will not sign up to rejection clauses. What then?
Keep the clauses in your own job template and intake check anyway, phrased internally as what you accept. Run the canary with the partner during onboarding — most agree to a quick test even when they resist paperwork. Send them the specific rejection reason the first time a file fails. Concrete evidence changes minds faster than a clause.
Should the conversion step be on the sender's side or the receiver's?
On whichever side you control and can test, and on exactly one side. If you receive, convert after the intake check so the hash is verified against the bytes as sent. If you send, convert before hashing so the manifest describes the file the partner will receive. Write the choice into the template.
How often should we re-run the canary?
Run it after any change to a client, server, job, or partner configuration, without exception. Also run it on a fixed schedule for flows that carry important data — quarterly is common. Record the date and hash each time so that when a mangler reappears, the search space is only what changed since the last clean run.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.