Sequence Numbers, Uniqueness, and Collision Avoidance
Two files want the same name on the same day, and only one of them is going to get it. Every naming convention starts life with a comfortable assumption: one file per flow per day. The datestamp makes each day's name unique, the folder stays tidy, and nobody thinks about uniqueness at all. Then the business grows, or a retry fires, or a partner splits a big extract into parts, and the assumption ends without notice. What happens next depends entirely on whether anyone planned for it. Unplanned, the outcome is one of three bad surprises. It might be a silent overwrite that destroys data or a rejected transfer that stalls the pipeline. Or it might be an automatic rename that quietly breaks every script expecting the agreed pattern. None of the three asks first.
This article, part of our Naming & Datestamps series, is about engineering that surprise out of existence. We will cover the mechanics of sequence numbers that sort correctly and compare the three practical sources of uniqueness with their honest tradeoffs. We will dissect the retry bug that manufactures duplicates out of thin air. And we will finish with the collision policy every flow should write down before it is needed. As throughout the series, examples use pattern tokens — YYYYMMDD where an eight-digit date would appear — and small sequence values like 0042. That way, everything stays evergreen and copy-safe.
When One File Per Day Stops Being True
A name collision is two different files claiming the same name in the same place. Collisions arrive by predictable routes, and recognizing them in advance matters. Each one looks like an anomaly the first time and a pattern the tenth:
- Volume growth. The nightly extract becomes an hourly extract. The daily stamp that made names unique now produces the same name up to two dozen times a day.
- Batch splitting. A file grows too large and the producer splits it into parts — same flow, same day, several files. There is nothing in
orders_YYYYMMDD.csvto tell them apart. - Corrections and resends. The morning file was wrong; the corrected afternoon file carries the identical name. Whether the second should replace the first is a business question — but the name alone can no longer distinguish them.
- Retries. A transfer times out and runs again. Depending on how the name is generated, the retry may collide with the original — or worse, fail to collide when it should. We return to that subtlety below.
- Multiple producers. A second warehouse, a second server, a second region comes online, producing files for the same flow on the same days. Each is unaware of the other's names.
Once any of these is on the horizon, the name needs a uniqueness field beyond the date. The conventional choice is a sequence number — a counter appended after the stamp: orders_YYYYMMDD_0001.csv, orders_YYYYMMDD_0002.csv, and so on, restarting each day. Before choosing where the numbers come from, two pieces of mechanics have to be right: what collisions do, and how the numbers must be written.
What a Collision Actually Does
When a second file arrives bearing an existing name, one of three things happens, and none of them is good if it is unplanned.
- Silent overwrite. Most upload operations replace an existing file without comment. The morning file is gone — not archived, not versioned, gone — and no error was raised anywhere. So the loss is discovered only when someone needs the original. This is the worst outcome for data and the most common default.
- Rejection. Some servers and some carefully written scripts refuse to replace an existing file. The data is safe, but the transfer fails, the retry fails identically, and the flow is stalled until a human intervenes. Rejection is the right behavior — but only when paired with an alert and a procedure, otherwise it is an outage with good intentions.
- Automatic rename. Some tools sidestep the conflict by inventing a new name — appending a suffix or a copy marker. The transfer "succeeds," but the file now violates the naming contract: patterns stop matching it, parsers reject it, and downstream jobs skip it silently. The pipeline is broken in the quietest possible way — the failure mode our opening article on names calls failing at a distance.
The common thread: the system picked the outcome, not you. The entire discipline of this article is moving that decision from accident to policy. First, make legitimate names never collide. Then decide deliberately what happens if they do anyway. The system will decide for you; it has just never decided well.
Zero-Padding: Making Numbers Sort Like Numbers
Sequence numbers live inside names, and names are sorted as text — character by character, left to right. Text comparison only agrees with numeric comparison when every number occupies the same number of characters. Write sequences unpadded and the listing betrays you as soon as the count passes nine:
unpadded — sorted as text zero-padded — sorted as text orders_YYYYMMDD_1.csv orders_YYYYMMDD_0001.csv orders_YYYYMMDD_10.csv orders_YYYYMMDD_0002.csv orders_YYYYMMDD_11.csv orders_YYYYMMDD_0003.csv orders_YYYYMMDD_2.csv orders_YYYYMMDD_0010.csv orders_YYYYMMDD_3.csv orders_YYYYMMDD_0011.csv
On the left, part 10 sorts between part 1 and part 2. That is because the comparison sees 1 versus 2 in the first digit position and never looks further. Any script that processes "in order" processes out of order; any human skimming the listing miscounts the parts. On the right, zero-padding to a fixed width restores the alignment. Every number is four characters, so the character comparison and the numeric comparison give the same answer. The same principle makes big-endian datestamps sort, as explained in datestamp formats that sort correctly. Part ten cut in line, and nothing in the listing objected.
How wide to pad is a capacity question: choose a width with an order of magnitude of headroom above the realistic daily maximum. A flow peaking at a few dozen files a day is comfortable at four digits. Padding wider than you need costs nothing but two characters, while padding too narrow creates an overflow. Overflow behavior must be decided, not discovered. A generator that reaches 9999 and emits 10000 breaks both the sort order and every parser expecting four digits. One that wraps to 0001 manufactures collisions. The right answer is almost always: pad generously, and make the generator fail loudly if it ever exceeds the width. A flow that has outgrown its sequence field by a factor of ten has outgrown other assumptions too. Two extra zeros are the cheapest insurance in this series.
Remember: a sequence number is part of the sorted identity of the file. Fixed width, zero-padded, always — 0042, never 42 — and an explicit, alarming failure if the width is ever exceeded.
Three Sources of Uniqueness
A sequence field is only as good as the mechanism that guarantees no two files get the same value. There are three practical sources, and they trade off against each other in ways worth knowing before you pick.
Source one: a counter
The producer keeps state — a counter file, a database row — increments it for each file, and resets it daily. Counters produce the friendliest names: 0001, 0002, 0003 reads naturally, sorts perfectly, and carries meaning ("part three of today's run"). A gap in the numbers even tells consumers something went missing — a property none of the other sources offer.
The cost is the state itself. The counter must survive restarts, which means it lives on disk or in a database. That means two processes incrementing it concurrently can read the same value and mint the same number — the classic race. A counter is trustworthy only with a single writer, or with locking around the increment. The lock-file techniques in our bash and cron automation series apply directly. Counters also stay honest only per producer: two servers with two counter files will both emit 0001 at dawn. Multi-producer flows need the producer's identity in the name as well (below), or a different source entirely.
Northgate Retail's counter collided after a server rebuild, and the rebuild itself went perfectly. The export server was reimaged one Tuesday at noon from a clean build. The counter file was not on the restore list, because nobody had thought of it as data. The morning run had already sent parts 0001 through 0003. The afternoon run on the rebuilt server started counting from 0001 again. The receiver refused to overwrite, so the afternoon's 0001 went to quarantine with an alert. Then 0002 and 0003 followed it there. Sorting out which was which took an hour with the server's activity log. The counter file is now on the list of things a rebuild restores, and the restore drill checks for it. The convention document notes that a counter is state, which means it is backed up like state.
Source two: the clock
Extend the datestamp to the second — orders_YYYYMMDD_HHMMSS.csv — and the time itself becomes the uniqueness field. The clock is stateless: nothing to store, nothing to lock, nothing to lose in a restart, and it works identically on every machine. For flows where files arrive minutes apart, time-to-the-second is the cheapest uniqueness there is, and it keeps the name sortable and human-readable.
Its honest limits: two files generated within the same second collide. Bursts are exactly the condition that breaks it, and batch splitters that write parts in a tight loop produce bursts by design. Two machines with skewed clocks can also collide or misorder files. And a local-time clock repeats one hour each year when daylight-saving time ends, making stamp collisions possible across that hour. Stamping in UTC eliminates the repeat, as covered in the datestamp article's time-zone discussion. Where bursts are possible, the standard reinforcement is clock plus a short counter within the second — or a different source. Same-second is not an edge case; it is what a loop does.
Source three: an identifier
Sometimes uniqueness already exists and the name just needs to carry it. A batch number, order number, or run ID from the producing business system is the best uniqueness money can't buy. It is guaranteed distinct by the system of record. It is meaningful to humans ("that's batch 0042 — the reprocessed one"). And it lets anyone join the file to the business event that produced it: ORD_YYYYMMDD_0042.txt. When a business identifier exists, use it.
When it does not, generated identifiers fill the gap. A UUID — a universally unique identifier, a long random value designed never to repeat anywhere — makes collisions effectively impossible with zero coordination. That is why multi-producer systems lean on it. The tradeoffs are real, though: random identifiers carry no order (sorting them is meaningless), no meaning, and considerable length. In file names they work best truncated to a short random suffix used as a tiebreaker after a sortable stamp, not as the whole identity. Finally, for multi-producer flows, the simplest identifier is often the producer's own name as an extra field — orders_wh1_YYYYMMDD_0042.csv versus orders_wh2_YYYYMMDD_0042.csv. This converts a shared namespace into per-producer namespaces where local counters are safe again. A UUID never repeats and never means anything.
Choosing a Source: The Decision Table
The right source follows from the flow's shape. This table is the whole tradeoff in one place:
| Uniqueness source | Needs state? | Sorts in order? | Burst-safe? | Multi-producer safe? | Best fit |
|---|---|---|---|---|---|
Daily counter (0042) |
Yes — file or DB, locked | Yes | Yes | No — one writer only | Single producer, parts must count up |
Time to the second (HHMMSS) |
No | Yes | No — same-second collides | Mostly — clock skew risk | Files spaced minutes apart |
Business ID (ORD_..._0042) |
No — system of record owns it | Usually | Yes | Yes | Files that map to business events |
| Random suffix / UUID | No | No — random | Yes | Yes | Many uncoordinated producers; tiebreaker after a stamp |
Producer field + counter (wh1_..._0042) |
Yes — per producer | Yes, within producer | Yes | Yes — namespaces split | Known set of producers, ordered parts |
A useful habit when reading the table: ask which failure you can live with. A counter's failure is duplicate numbers under concurrency; the clock's failure is same-second collision; an identifier's failure is depending on another system's discipline. Pick the failure you are best equipped to prevent. I usually pick the counter, because I trust a lock more than a clock.
The Retry That Regenerates the Name
The subtle bug is the one that produces duplicate content under distinct names, sails past every collision defense, and quietly double-loads data downstream.
Picture a job that builds a name from the current time, then uploads: report_YYYYMMDD_HHMMSS.csv. The upload times out. The retry logic — sensibly — runs the job again. But the job's first step is generate the name, and the clock has moved. So the retry uploads the same content as report_YYYYMMDD_HHMMSS.csv with a later time. If the first upload actually succeeded, the destination now holds two names, one content. Timeouts often mean "the response got lost," not "the transfer failed." Our retry and error handling series examines that distinction. Every name-based duplicate check passes. The order file loads twice.
The fix is a design rule: a file's name is decided once, at creation, and never again. Generate the name and write it down — in the job's state, in a manifest, in the file itself sitting in an outbox. Make retries reuse the recorded name rather than re-running the generator. The general form of that habit is checkpointing in automation. With a stable name, one of two things happens. The retry overwrites the identical file it already sent (harmless). Or it is refused by an existence check that can then verify sizes match (better). Either way, the destination converges on one file. This one rule is most of what makes a transfer job idempotent — safe to run twice. It is the naming half of the story told in our duplicate detection and idempotency series. Note what it implies for counters, too: the counter increments when a file is created, never when a transfer is attempted.
The rule in one line: generate once, retry with the same name. If your retry path calls the name generator, you have a duplicate factory waiting for its first timeout.
A Collision Policy Decided in Advance
Even with a sound uniqueness source, collisions remain possible — a misconfigured producer, a manual re-upload, a restored backup replayed into a live folder. The difference between an incident and a non-event is a collision policy. It is a short written answer, per flow, to the question "what happens when a name already exists?" It belongs in the naming convention document — the one built in designing a naming convention — and it needs only three decisions:
- The outcome. Almost always: reject, quarantine, alert. The colliding file is refused or moved to an error folder, the original is untouched, and a human is notified. Silent overwrite is acceptable only where the design explicitly makes later files replace earlier ones — a "latest snapshot" flow. In that case, the convention should say so in writing. Deliberate replacements deserve a visible marker instead of a quiet clobber: a correction file named
orders_YYYYMMDD_0042_r01.csvannounces itself, where an overwrite hides. - The enforcement point. Belt and suspenders: the producer checks for the name before sending (cheap, catches its own bugs). The receiving side avoids blind replacement where the workflow allows. Uploading to a temporary name and renaming into place is the natural moment for that existence check. Our partial-file safety series details the pattern. The rename is where the final name is claimed.
- The evidence. A collision is a symptom; the investigation needs to see both events. This is where server-side logging earns its keep. Sysax Multi Server records the file name of every transfer in its activity logs — written to a log file and optionally to a database. So a search on the disputed name shows each upload, when it happened, and which account sent it. On the automation side, Sysax FTP Automation pairs its scheduled and folder-monitoring tasks with retry and error handling and can send email notifications. So a refused transfer becomes a message naming the file instead of a silence in a log nobody reads.
Write the three answers down while everything is calm. A collision handled by policy is a log line and a ticket; a collision handled by improvisation at month-end close is a war story. This series has enough of those already.
Uniqueness Rules to Take Away
The whole article compresses to a short list. Names must be unique per file within a flow. The date alone stops guaranteeing that the moment volume, splitting, corrections, retries, or a second producer arrives. Sequence fields are fixed-width and zero-padded — 0042 — with overflow a loud failure, never a wrap. Uniqueness comes from a locked single-writer counter, the clock to the second, or an identifier — chosen by the table above. The producer's name is added whenever more than one producer shares a flow. A file's name is generated exactly once and survives retries unchanged. And every flow has a written collision policy: reject, quarantine, alert — with the server's activity log as the referee.
From here, the natural next read is datestamp formats that sort correctly if you arrived without the date field settled. Read designing a naming convention to give the sequence field its official place in the pattern. Also read parsing names in scripts for the consuming side — including validating that a sequence field really is four digits before trusting it.
Frequently Asked Questions
How many digits should a sequence number have?
Why not just use the time as the sequence?
What should happen when a file with the same name already exists?
Do gaps in sequence numbers matter?
Are two files with different names but identical content a problem?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
