Safe Characters: Names That Survive Every System
"Is that an underscore or a hyphen after 'report'?" "Neither. It is a space, and then an ampersand." A file in an automated flow does not live in one place. Over its lifetime, its name will be stored by a filesystem, quoted by a shell script, and matched by a glob. It will be embedded in a URL, written into log lines, packed into an archive, and pasted into a ticket. And, as above, it will be read aloud over the phone during an outage. Each of those environments has its own rules about which characters are allowed, which are special, and which two names count as "the same." The rules disagree. A name that is perfectly legal where it was created can be unrepresentable, ambiguous, or actively dangerous three hops later. The phone call is usually hop four.
The safe response is not to memorize every system's rules. It is to name files within the small, boring intersection that every system treats identically — the character set that has never caused a ticket anywhere. This article, part of our Naming & Datestamps series, walks through each family of hazards honestly — case, spaces, Unicode, reserved names, special characters, length. That way, you understand why the conservative set is drawn where it is. The article closes with that set as a copyable rule. As throughout the series, example names use pattern tokens like YYYYMMDD in place of real dates.
Why "Legal" Is Not the Same as "Safe"
Legality is local, and pipelines are not. Every filesystem publishes rules about legal names, and it is tempting to treat those rules as the standard to meet. Think of a file name as a traveler. It does not just need a valid passport at home. It needs to clear every border on the itinerary. The itinerary for a typical transfer flow is longer than it looks:
- the producing system's filesystem, then the producer's scripting shell;
- the transfer protocol and the server on the other side — possibly a different operating system with different rules;
- the consuming system's shell, scheduler, and parsing scripts;
- the side channels: URLs if files are fetched over HTTP, archive formats if they are zipped, and email if names appear in notifications. These also include log files and databases where names are recorded, and spreadsheets where someone pastes a listing.
A name is safe only if it passes every border unchanged and unambiguous. That is a much stricter test than "my filesystem accepted it." It is the test that matters, because the failure always happens at the border you were not thinking about. The incidents in why file names make or break automation all share that shape: legal at home, broken abroad.
Case Sensitivity: The Quiet Difference
The deepest disagreement between mainstream systems is whether Report.csv and report.csv are the same name. Windows filesystems are case-insensitive but case-preserving. They remember the capitalization you used, display it, and then ignore it for every comparison. You cannot have both spellings in one folder, and either spelling opens the file. Typical Linux filesystems are fully case-sensitive. The two spellings are unrelated names, both can exist side by side, and a pattern like report_* matches one of them only. Windows remembers your capitals; it just declines to act on them.
Every mixed pipeline therefore contains a silent translation problem. A case change made on Windows is invisible on Windows — every local test passes. It only fails on the case-sensitive hop, where the consumer's pattern quietly matches nothing. No error, no crash: the job "succeeds" while doing no work. This is the single most common naming incident in cross-platform flows, and the defense is not technical but conventional:
- Lowercase everything. When only one case is ever legal, case cannot drift. Lowercase wins over uppercase for legibility in listings and because most tooling examples assume it.
- Compare case-sensitively everywhere, even on systems that would forgive you. Write globs, regexes, and equality checks as if every system were case-sensitive. Then behavior never depends on which OS a script happens to run on.
- Never rely on case to distinguish files. If
ORDERS_andorders_mean different flows in your design, the design fails the moment those files land on a Windows server. There, they collide.
Spaces and the Quoting Tax
Spaces are legal on every modern filesystem, and they are still the most reliable script-breaker in the business. The reason is that the space is the fundamental separator of command-line computing. Shells split unquoted text into arguments at whitespace. So a name containing a space arrives at a program as two names, neither real. Quoting fixes it — "$file" in shell scripts. But quoting is a tax levied on every script, every command, every log parser, forever. It must be paid perfectly each time. One forgotten quote in one inherited script is a latent failure. A space in a file name is a life choice; the shell disagrees with it.
The tax collectors do not stop at the shell. In URLs, a space must be encoded as %20, and different tools disagree about doing that automatically. In log files, a space-bearing name is ambiguous: does the line record one file or two? In CSV exports of listings, spaces flirt with delimiter confusion. None of these problems has any payoff attached — the space adds readability that an underscore provides equally well at zero risk. This is why every serious pipeline convention bans spaces outright: monthly_report_YYYYMM.csv, never monthly report_YYYYMM.csv.
The Honest Unicode Story
Modern systems can, in principle, put almost any character in a file name — accented letters, non-Latin scripts, symbols. For documents humans manage by hand, that is a good thing. For pipeline files, the honest advice is different, and it rests on how non-ASCII names actually behave in transit.
First, some vocabulary. ASCII is the small, ancient character set — unaccented letters, digits, basic punctuation — that every system in existence represents identically. Unicode is the modern standard covering essentially all human writing, stored as bytes through an encoding. Our primer on character encodings for admins covers the ground in more depth. The catch is that Unicode sometimes offers more than one byte sequence for the same visible text. An accented letter like é can be stored as one composed character or as a plain e followed by a combining accent mark. The two forms — called different normalization forms — render identically on screen and compare as different strings byte for byte. Operating systems disagree about which form they store and whether they convert. A name created on one platform can arrive on another normalized differently. At that point, "the same name" no longer matches itself. A glob that worked on the producer fails on the consumer. The two files in a listing that look identical are, to the machine, unrelated. The listing shows one name twice; the machine sees two strangers.
Transport layers add their own hazards. Older transfer protocol paths and clients assume ASCII and can mangle other bytes in transit. Archive formats have a messy history of storing name encodings inconsistently. So a ZIP created on one system can unpack with garbled names on another. Terminals, logs, and monitoring tools vary in what they display. The mechanics of each mangling are traced in filename encoding problems. Every one of these failures is invisible at creation time and expensive at diagnosis time. The name looks right everywhere while comparing unequal somewhere.
The pipeline rule is pragmatic, not ideological: machine-read file names stay ASCII. Human-language information — customer names, localized titles, descriptions in any script — belongs inside the file or in accompanying metadata. There, encodings are declared and handled properly. That information does not belong in the name that routing and matching depend on. Pipelines built this way never meet the normalization problem.
Remember: two file names can look identical on screen and still be different byte sequences. If a "missing" file is plainly visible in the listing, suspect case on one axis and Unicode normalization on the other. Prevent both forever by keeping pipeline names lowercase ASCII.
Windows Reserved Characters and Names
Windows forbids a specific set of characters in file names. Any cross-platform flow must treat that set as globally forbidden, because a name using them cannot land on a Windows disk intact. The reserved characters are:
< > : " / \ | ? * and control characters
Several of these are tempting. The colon is the natural time separator. That is exactly why sortable time stamps are written HHMMSS instead, as covered in datestamp formats that sort. The forward and backward slashes are path separators on Linux and Windows respectively, so both are structurally unusable inside a name. Question mark and asterisk are wildcard characters in both shells and many transfer tools. A literal ? in a name turns every later pattern match into a puzzle.
Windows also reserves a set of device names inherited from its earliest ancestry: CON, PRN, AUX, NUL, and the numbered COM and LPT names. A file called con.csv or aux.txt — the reservation applies even with an extension — cannot be created normally on Windows. One created on Linux and transferred over becomes a booby trap for every Windows tool that touches the folder. The names look innocent; a flow that derives file names from user input or upstream data can produce one by accident. Finally, Windows disallows names ending in a dot or a space. Linux happily accepts these trailing characters, and they are invisible in most listings. They produce files that exist, display normally, and cannot be opened or deleted by ordinary Windows tools.
Acme met the trailing space in a partner feed that had run cleanly for a year. A change on the partner's Linux side left one space at the end of every name. No listing anyone looked at showed it. The Windows transfer server on Acme's side silently trimmed it on write, so the files landed under the name everyone expected. The partner's own confirmation step then listed the folder and looked for the name it had sent, space included. It found nothing and resent the file every ten minutes until someone asked why the same orders file had arrived forty times. Nothing was loaded twice, because the receiver refused to overwrite and quarantined the repeats. The partner agreement gained one line about trailing spaces and dots. The export was fixed the next day. The confirmation step now compares trimmed names, which is what the server had been doing all along.
If your receiving side runs on Windows — a Windows transfer server such as Sysax Multi Server, for instance — these rules are not trivia. In that case, they are the landing conditions for every inbound name. A partner producing names on Linux has no local reason to notice them. Put the forbidden list in the partner agreement, a practice our guide to trading partner onboarding treats as a standard onboarding item.
Legal Everywhere, Trouble Everywhere
Beyond the outright forbidden, there is a band of characters that every filesystem accepts and almost every tool mishandles somewhere. These are the "creative punctuation" incidents of the opening article, and each has a specific mechanism:
- Ampersand (
&): unquoted in a shell, it backgrounds the command. In batch files it separates commands; in URLs it separates parameters. In HTML and XML logs it must be escaped. Four environments, four different misbehaviors, one character. - Quotes and apostrophes: they terminate quoted strings early, so a name like
o'brien_report.csvbreaks the very quoting that was protecting the script. - Dollar sign and backtick: inside double quotes, shells expand these — a name can trigger variable expansion or, with backticks, command execution. This is how a bad name graduates from nuisance to security problem.
- Parentheses, semicolons, exclamation marks: command separators, history expansion, and grouping in various shells; all require careful quoting to pass through unharmed.
- Percent (
%): variable syntax in Windows batch files and the escape prefix in URLs —report%20final.csvmeans something entirely different to a web server than to a filesystem. - Hash (
#): starts a comment in most config files and scripts; a name pasted into a config silently truncates. - Leading hyphen: the special one. A file named
-for--forcelooks like a command-line option to every tool that receives it as an argument. Tools accept--to mark "end of options," and paths can be prefixed with./. But those defenses must be present in every script that will ever touch the file. A convention that starts every name with a letter costs nothing and removes the whole class.
The unifying lesson: each of these characters is manageable with perfect discipline. Pipelines do not run on perfect discipline. They run on the accumulated scripts of years, written by many hands. Characters that require discipline are exactly the characters to design out. Discipline is a fine thing to have and a poor thing to depend on.
Length Limits: The Path Is the Real Constraint
Name length feels like a non-issue — every modern filesystem allows names far longer than anyone types. The practical limit is somewhere else: the full path. Systems and tools limit the total length of the folder-chain-plus-name string. Several widely used Windows APIs and utilities enforce a much shorter total than the filesystem itself supports. A file with a modest name at the bottom of a deep folder tree can exceed the working limit of some tool in the chain. At that point, it exists, appears in listings, and fails to open, copy, or delete depending on which program tries. The disk is fine with it; the tool asking is not.
Transfer flows hit this in predictable ways. Archives unpack into nested folders and push depth suddenly. Sync and backup tools add their own staging prefixes to paths. Migration jobs copy trees into deeper trees. Bulk-copy work, where long paths are a classic wrecker, is covered in our robocopy and rsync bulk move series. The practical guidance is stated generally because exact numbers vary by tool and configuration. Keep pipeline file names comfortably under about a hundred characters, and keep transfer folder trees shallow. Treat the name-plus-path budget as shared. A verbose name spends headroom the folder structure may need. Names built from a handful of short fields — source, type, stamp, sequence — land far inside every limit with room to spare. We learned the path budget the slow way, from an archive that unpacked six folders deep.
The Conservative Character Set
All of the above compresses into one small ruleset. This is the intersection that survives every filesystem, shell, protocol, URL, archive, and log format a transfer flow realistically meets:
THE CONSERVATIVE FILE NAME RULE
Allowed characters: a-z 0-9 _ - .
Case: lowercase only
First character: a letter
Dots: exactly one, before the extension
Spaces: never
Length: keep name under ~100 chars; budget the full path
Extension: present, honest, lowercase
As a validation regex:
^[a-z][a-z0-9_-]*\.[a-z0-9]+$
Examples that pass: sales_YYYYMMDD.csv
backup_YYYYMMDD_HHMMSS.zip
ord_YYYYMMDD_0042.txt
Examples that fail: Sales_YYYYMMDD.csv (uppercase)
sales YYYYMMDD.csv (space)
sales_YYYYMMDD.final.csv (two dots)
-sales_YYYYMMDD.csv (leading hyphen)
A few notes on the deliberate choices. The underscore is the field separator and the hyphen is allowed within fields — or the reverse, if your convention prefers. What matters is that the separator character never appears inside field values. The reasoning for that rule belongs to convention design. The single-dot rule keeps "split on dot, take the last piece" a valid way to find the extension. It also prevents the double-extension confusion that both parsers and humans fall for. Requiring a leading letter kills the leading-hyphen hazard and keeps names shell-argument-safe. Nothing in the set needs quoting, encoding, or escaping anywhere — which is the entire point.
Enforcing it at the border
A character rule that lives only in a document is a suggestion. Make it executable: validate every arriving name against the regex at the pipeline's front door. Route nonconformers to an error folder for a human instead of letting them into the machinery. This is the reject-don't-guess pattern detailed in parsing names in scripts and used throughout well-designed watch folder workflows. Enforcement has a second audience, too: your own logs and tools stay clean. Every transferred file's name ends up recorded. A server like Sysax Multi Server writes the name of each transfer into its activity logs. Names drawn from the conservative set keep those records unambiguous, searchable, and safe to paste into any script or query during an investigation.
A Name That Survives Everything
Cross-platform naming is one of those subjects where the full story is long and the conclusion is short. The systems disagree about case, so use one case. Shells split on spaces, so use none. Unicode normalization makes identical-looking names unequal, so keep machine names ASCII. Windows forbids characters and words that Linux permits, so honor the stricter set everywhere. Special characters demand perfect quoting forever, so choose characters that demand nothing. Paths, not names, hit length limits, so spend characters frugally. Follow the conservative set and every one of those sentences becomes someone else's problem. Nobody will ever again have to read an ampersand aloud at two in the morning.
From here, continue with designing a file naming convention, which gives the safe characters a structure — fields, delimiters, and a written standard. Also continue with the series article on sequence numbers for keeping those well-formed names from colliding. If you want the motivation refreshed, the incident stories in why file names make or break automation show each hazard from this article doing real damage.
Frequently Asked Questions
Is it really wrong to use spaces in file names?
Why avoid accented characters if my systems all support Unicode?
What are the Windows reserved names I should know?
Underscores or hyphens — which should I use?
How long is too long for a file name?
A file shows in the listing but scripts say it doesn't exist. How?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
