Home › Topics › Testing & Staging › Test Files

Synthetic Test Files and Test Data

The first file through any new transfer job is called test.txt. It contains the word "test", or nothing. It proves that a file can be moved from one folder to another, which was never in doubt. A staging environment with nothing better to move is an empty stage with one very small prop. Every real test of a transfer change needs files. They need a known size so you can check nothing was truncated. They need known content so you can check nothing was corrupted. They need awkward names so you can check the job handles them. They need to look like the real thing so the job's filters and parsers are actually exercised. The tempting shortcut, copying last night's production files into staging, is the one thing you must not do.

This article builds a test data set that earns its place. It explains what synthetic data is and why production data is banned from staging. It covers which classes of test file catch which mistakes. It shows how to create files of exact sizes and known hashes on Windows and Linux, and how to generate the awkward filenames. It shows how to write a small generator that produces realistic-but-fake CSV files in every line ending and encoding you need. It explains how to organize the set so anyone can regenerate it. It is part of our Testing and Staging Transfer Changes series. It feeds the staging environment described in building a transfer staging environment.

I have a test.txt of my own from my first year in this work. I am fairly sure it is still sitting in a partner's inbound folder, proving nothing. This article is how to do better than I did.

Synthetic Data, and Why Production Files Are Banned

Synthetic data is data you generate for testing. It has the shape of real data — the same columns, the same file names, the same sizes. But none of the content is real. A synthetic customer file has customer IDs that belong to nobody, names drawn from a list, and addresses in towns chosen at random. Its cousin is masked data: real files with the sensitive fields scrambled or replaced. Masking is a legitimate technique with its own rules, described in masking and pseudonymization. For transfer testing synthetic data is simpler and safer, because there is nothing in it to leak.

Copying production files into staging is banned for four reasons, and the practical one is as strong as the legal ones:

  • Personal data. Production files routinely contain names, addresses, account numbers, and payroll or health details. Privacy law generally requires such data to be used only for the purpose it was collected for. "Testing a transfer job" is not that purpose. Recognizing personal data shows how much of it hides in ordinary batch files.
  • Contract data. Partner files are often covered by confidentiality clauses that say where the data may be stored. A staging server is rarely on that list.
  • Staging is less protected. By design, more people have access, experiments happen, and logs are read casually. A production file in staging is a production file with weaker guards around it.
  • Production files do not contain the edge cases. Last night's files are all well-formed, medium-sized, and correctly named — the ones that were not have already been fixed. A test needs the empty file, the huge file, and the file with a space in its name. Those are exactly what a production copy lacks.

Remember: nothing that came from production goes into staging. Not a file, not a row, not a filename copied from a real partner's feed. If you need a file that looks like the real thing, generate one. If a test genuinely needs real structure, mask a copy under the rules in our personal-data series and treat the result as production data anyway.

The Test File Catalog

A good test set is small — a few dozen files — but every file is there for a reason, and none of them is called test.txt. Each class below catches a particular kind of mistake, and the table is the checklist for building the set.

Class Example What it catches
Empty TEST_empty.csv, 0 bytes Jobs that skip, crash on, or wrongly "succeed" with a zero-byte file; partner loaders that reject it
Tiny 1 byte; a header row only Off-by-one size checks, "file has no data rows" handling
Typical A generated CSV of a few hundred kilobytes The happy path; parsers, filters, post-processing
Boundary Exactly 1048576 bytes; one byte over a configured size limit Size limits and quotas that are off by one, "greater than" versus "at least"
Huge Larger than your biggest real file; larger than 4 GiB if any tool in the chain might be 32-bit Timeouts, disk space, resume, size fields that overflow
Many small Two thousand files of 1 KB Per-file overhead, listing limits, jobs that stop after N files
Awkward names Spaces, accents, leading dash, two hundred characters Quoting bugs in scripts, encoding mismatches, length limits on the far side
Content variants Same CSV with CRLF, LF, a byte-order mark, a legacy encoding Text-mode transfers, parsers that choke on a BOM, mangled accented characters
Binary Random bytes; every byte value present Any transfer that alters bytes, compression that assumes text
Decoy A file the job should not pick up: wrong extension, wrong prefix, a temp name Patterns that match too much — the wildcard incident

Two classes deserve a word. The huge file is about limits, not realism, so make it just larger than whatever limit worries you. Genuinely large transfers have their own strategies in our Large File Strategies series. The many small class exists because per-file overhead is a real failure mode, explained in the many-small-files problem. A job that handles ten files perfectly can behave quite differently with two thousand. Two thousand is where a job discovers its personality.

Files of Exact Size and Known Hash

The two properties every test file needs are an exact size and a known hash. A hash is a short fingerprint computed from the file's bytes. If even one byte changes in transit, the fingerprint changes too. Hashing explained covers the idea in depth. Here you only need to know that SHA-256 is the usual choice and that every operating system can compute it.

Windows

The built-in fsutil creates a file of any exact size almost instantly. The file is filled with zeros — fine for size and limit tests. It is useless for anything involving compression (zeros compress to nothing) or corruption detection (a zero replaced by a zero is invisible). For random content, PowerShell fills a buffer from the cryptographic random generator and writes it out. For a large file, it writes the buffer repeatedly.

:: exact sizes, zero-filled (fast)
fsutil file createnew C:\testdata\TEST_empty.bin 0
fsutil file createnew C:\testdata\TEST_1MiB-zeros.bin 1048576
fsutil file createnew C:\testdata\TEST_limit-plus-one.bin 10485761

:: a huge file that takes no disk space: mark sparse, then set its length
fsutil file createnew C:\testdata\TEST_10GiB-sparse.bin 0
fsutil sparse setflag C:\testdata\TEST_10GiB-sparse.bin
fsutil file seteof C:\testdata\TEST_10GiB-sparse.bin 10737418240
# PowerShell: random content, 1 MiB and 100 MiB
$rng = [System.Security.Cryptography.RandomNumberGenerator]::Create()
$buf = New-Object byte[] 1048576
$rng.GetBytes($buf)
[System.IO.File]::WriteAllBytes('C:\testdata\TEST_1MiB-random.bin', $buf)

$fs = [System.IO.File]::Create('C:\testdata\TEST_100MiB-random.bin')
for ($i = 0; $i -lt 100; $i++) { $rng.GetBytes($buf); $fs.Write($buf, 0, $buf.Length) }
$fs.Close()

# record the hashes
Get-FileHash -Algorithm SHA256 C:\testdata\*.bin | Format-Table Hash, Path -AutoSize
# or, from a command prompt: certutil -hashfile C:\testdata\TEST_1MiB-random.bin SHA256

Linux

On Linux the same three jobs are one line each. head -c takes an exact number of bytes from the random device. dd does the same in blocks for larger files. truncate makes a sparse file of any length. And sha256sum writes a manifest that can later verify the whole set in one command.

mkdir -p ~/testdata && cd ~/testdata
: > TEST_empty.bin                                   # zero bytes
head -c 1048576 /dev/urandom > TEST_1MiB-random.bin      # exact size, random content
dd if=/dev/urandom of=TEST_100MiB-random.bin bs=1M count=100 status=none
truncate -s 10G TEST_10GiB-sparse.bin                  # sparse: no disk used until written
sha256sum *.bin > SHA256SUMS                             # the manifest
sha256sum -c SHA256SUMS                                  # verify everything later

The manifest is the point of the exercise. After a test run, use the same command on the receiving side — sha256sum -c, or Get-FileHash compared against the stored list. It tells you in seconds whether every file arrived intact. The format and habits around manifests are described in checksum files and manifests.

Watch out: a sparse file is a lie the filesystem tells for your convenience. It occupies no disk space locally. But the moment a transfer tool reads it, every one of those ten gigabytes of zeros goes over the wire. It lands fully allocated on the far side. That is exactly what you want when testing limits and timeouts. It is exactly what you do not want to discover on a staging server with eight gigabytes free.

The Awkward Names

Filename bugs are quoting bugs, encoding bugs, and length bugs, and they hide until a real file trips them. The script below creates the classic troublemakers in one go. Run it on the sending side and see which ones the job refuses, mangles, or silently skips.

mkdir -p ~/testdata/names && cd ~/testdata/names
echo test > "TEST_report with spaces.csv"
echo test > "TEST_résumé-données.csv"              # accented, UTF-8
echo test > "TEST_日本語.csv"                       # non-Latin script
echo test > "-TEST_leading-dash.csv"                # looks like an option to careless scripts
echo test > ".TEST_hidden.csv"                      # hidden on Linux, ordinary on Windows
echo test > "TEST_UPPER.CSV"                        # case: distinct file on Linux, same on Windows
echo test > "TEST_many.dots.in.name.csv"
echo test > "TEST_$(printf 'a%.0s' {1..200}).csv"   # two hundred a's
echo test > "TEST_semicolon;and(parens).csv"

Two names are deliberately missing. A name with a trailing space or trailing period cannot be created on Windows at all — the system strips it. So a Linux sender can produce a file a Windows receiver cannot store. That is worth knowing but not worth putting in the standard set. And names containing characters Windows forbids outright (< > : " | ? *) are the same story. The full cross-platform rules are in safe characters across platforms. The test set should include the cases that are legal everywhere but fragile, which is what the list above is.

Northgate Retail's test set was, for two years, one file called test.txt. The week they built the awkward-name set, the store returns job silently skipped TEST_report with spaces.csv in staging. The pre-processing script passed the name to a command line without quotes. The shell went looking for three files that did not exist. The fix was two quotation marks. The following Monday a new regional office began sending files named Store 12 returns.csv, spaces included. The job took them without comment. Nobody at the regional office ever learned how close their first feed came to being reported as a success with nothing in it.

Notice the TEST_ prefix on every name. It is not decoration. It is the rule that a test file can never be mistaken for a production one. The rule is enforced on the receiving side as well as the sending side. Partner test windows and test endpoints explains how the prefix is agreed with partners and how production automation is taught to reject it.

A Generator for Realistic-but-Fake CSV

Size-and-hash files test the transfer. To test the job, you need files with the right structure and plausible content. You need to test its filters, its validation step, its post-processing, and the partner's parser. The way to get an endless supply of files is a small generator. The Python script below produces shipment files with fake names, fake cities, and customer IDs in a reserved range that production validation is configured to reject. Everything is driven by a seed, so the same arguments always produce byte-identical output. That means the generated files have known hashes too.

# make_testdata.py - deterministic fake shipment CSVs for transfer tests
import csv, random, sys
from datetime import date, timedelta

FIRST = ["Aiko", "Bram", "Chloé", "Dmitri", "Esme", "Farid", "Greta", "Hugo"]
LAST  = ["Okafor", "Lindqvist", "Marchetti", "Nakamura", "O'Brien", "Petrov"]
CITY  = ["Leeds", "Tromsø", "Zürich", "Porto", "Kraków", "Cork"]

def make(path, rows, base, seed=42, newline="\r\n", encoding="utf-8"):
    rng = random.Random(seed)            # same seed + same base date = identical file
    with open(path, "w", newline="", encoding=encoding) as f:
        w = csv.writer(f, lineterminator=newline)
        w.writerow(["shipment_id", "customer_id", "name", "city",
                    "ship_date", "qty", "weight_kg"])
        for i in range(rows):
            w.writerow([
                f"TEST-{i:07d}",
                900000 + rng.randint(0, 99999),   # reserved range: never a real customer
                f"{rng.choice(FIRST)} {rng.choice(LAST)}",
                rng.choice(CITY),
                (base - timedelta(days=rng.randint(0, 365))).isoformat(),
                rng.randint(1, 500),
                round(rng.uniform(0.1, 950.0), 2),
            ])

if __name__ == "__main__":
    rows = int(sys.argv[1]) if len(sys.argv) > 1 else 1000
    # pass a fixed base date (YYYY-MM-DD) so the output hash never changes
    base = date.fromisoformat(sys.argv[2]) if len(sys.argv) > 2 else date.today()
    make("TEST_shipments_crlf.csv",   rows, base)
    make("TEST_shipments_lf.csv",     rows, base, newline="\n")
    make("TEST_shipments_bom.csv",    rows, base, encoding="utf-8-sig")
    make("TEST_shipments_latin1.csv", rows, base, encoding="latin-1")
    make("TEST_shipments_header-only.csv", 0, base)

Run it as python make_testdata.py 1000 YYYY-MM-DD with a real base date in place of the placeholder. You get five files. Four contain the same thousand rows: with Windows line endings, with Unix line endings, with a UTF-8 byte-order mark, and in a legacy single-byte encoding. The last is a header-only file with no data rows. The names contain accents on purpose. The cities are chosen so that every one of them survives the legacy encoding. That is what makes the fourth file a fair test rather than a crash.

Three design choices are worth copying into any generator you write:

  • Reserved identifiers. Use customer IDs from a range production never issues, and shipment IDs prefixed TEST-. That way, if a test file ever does reach a production loader it is rejected on content, not just on filename.
  • Determinism. A fixed seed and a fixed base date make the output reproducible. So the hash of TEST_shipments_crlf.csv can be written down once and checked forever. Change the seed when you want a different file, and record the new hash.
  • Variants from one source. Line endings and encodings are parameters, not separate scripts, so the four variants differ in exactly one property each. When a parser rejects the BOM file and accepts the others, you know precisely why.

Line Endings, Encodings, and Other Content Traps

The content variants exist because the most common "the file arrived but it is wrong" failures are not corruption in the dramatic sense. One is a text-mode FTP transfer that rewrote every line ending. Another is a parser that treated the three-byte marker at the start of a UTF-8 file as part of the first column name. Another is a legacy system that turned Zürich into Z├╝rich. Each variant in the set is there to provoke one of those. Zürich has been mangled by more transfer jobs than any city deserves. If your flows carry fixed-width or other flat-file formats, add a variant of those too. Our article on flat file formats lists the shapes to cover. Our ASCII, Binary, and Encoding Corruption series explains the mechanisms behind each trap.

A useful trick: record the hash of each variant and its byte count. A line-ending rewrite changes the size by exactly the number of lines; a stripped byte-order mark changes it by exactly three. When a test fails, the size difference often names the culprit before you open the file.

Organizing and Regenerating the Set

The test set lives in one folder, under version control if you have it, with a fixed layout and a manifest. Nobody should have to remember how it was made, because nobody will.

testdata/
  README.txt            what each file is for, and the exact commands used
  make_testdata.py      the generator (seed and base date recorded in README)
  make_sizes.sh         the size-and-hash commands from this article
  SHA256SUMS            manifest for every file below
  sizes/                TEST_empty.bin, TEST_1MiB-random.bin, TEST_100MiB-random.bin ...
  names/                the awkward-name files
  csv/                  TEST_shipments_crlf.csv, TEST_shipments_lf.csv, ...
  decoys/               TEST_shipments.tmp, TEST_shipments.csv.part, notes.txt
  many-small/           generated on demand: two thousand 1 KB files

Regeneration is a script, not a memory. The huge file and the many-small folder are generated on demand rather than stored, because they are cheap to make and expensive to keep. The manifest is regenerated whenever the generator changes, and the change is recorded. A regression suite — see regression testing transfer jobs — will compare outcomes against these exact hashes for months.

Where the set is stored matters too. It never lives inside a production source folder, an inbound drop, or any path a production job watches. The TEST_ prefix and the reserved IDs are defense in depth, not an invitation to keep test files next to real ones. The prefix is a seatbelt, not a reason to park in the inbound folder. In staging, the test itself copies the set into the source folders fresh for each run, so every test starts from a known state.

Putting the Set to Work

The everyday use of the set is a staging run. Copy the relevant files into the staging job's source folder. Run the job against the partner simulator, then verify the result on the simulator. A tool with a folder-compare feature makes the last step easy. Sysax FTP Automation, for instance, can compare a local folder against a remote one to show which files are missing, extra, or different in size. That is a quick first check before the hash comparison confirms that the contents match byte for byte.

Beyond the staging run, the same files serve three other jobs. The size-and-hash files are the payload for a post-change smoke test. The whole catalog, one class at a time, becomes the case list for the regression suite. And a small agreed subset — one typical CSV and one hash file — is what you exchange with a partner during a test window. The partner knows the names and hashes in advance.

The set also answers the question people ask when a test fails: "was it the file or the transfer?" Every file's size and hash are known before it moves. So a mismatch on the far side is always the transfer. Our article on verifying transfers end to end shows where along the path it happened.

Build the Set Once, Use It for Years

A good synthetic test set is a few dozen files, a two-page generator, a manifest of hashes, and a README that says how it was made. It contains no production data, because it needs none: the edge cases production files never have are precisely what it is for. Build it once, keep it under version control, and regenerate it by script. Every staging run, smoke test, regression run, and partner test can draw on the same known-good files. And test.txt, having proved that files can move, can finally retire.

The next articles in this series put it to use. Read partner test windows and test endpoints for the files you exchange with the other side. Read regression testing transfer jobs for the suite that runs the whole catalog after every change.

Frequently Asked Questions

Can I use a production file if I delete the sensitive columns first?
That is masking, and it is only safe under a proper process. That means knowing every sensitive field and replacing rather than deleting where structure matters. It means treating the result as still sensitive. For transfer testing it is almost always simpler to generate a synthetic file with the same columns and fake content.
Why do I need random content? Is a zero-filled file not enough?
Zero-filled files are fine for size and limit tests. They fail for anything involving compression, because zeros compress to almost nothing. They also fail for corruption tests, because a corrupted zero is often still a zero. Use random content whenever the test cares about what is in the file rather than how big it is.
What is a sparse file?
A file whose unwritten regions take no disk space; the filesystem returns zeros for them on read. It lets you make a ten-gigabyte test file instantly. But a transfer reads and sends all ten gigabytes, and the receiving side stores them fully.
How big should the "huge" test file be?
Make it just larger than the limit you are testing. That could mean bigger than your largest real file or bigger than a configured size cap. Or it could mean bigger than 4 GiB if any tool in the chain might use 32-bit size fields. Realism is not the goal; crossing the boundary is.
Why does the generator take a seed and a base date?
So that the same arguments always produce a byte-identical file. That gives the generated CSV a known hash, which a regression suite can check against for as long as the generator is unchanged. Change the seed for a different file, and record the new hash.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.