Masking and Pseudonymization Before Sending
"It's fine, it's test data — we changed the names." The file was last quarter's customer export with every name replaced by Test1 through Test4400. Every email address, birthdate, and postcode was left exactly as it was. Sometimes the file has to carry person-level records. The analytics partner needs to follow repeat customers across months. The test team needs realistic-looking data. The processor needs to work case by case. You have already applied minimization and cut everything the purpose does not need — and what remains still points at real people. This is where the disguise toolkit comes in: techniques that keep records useful while making them harder to connect to the individuals behind them. Changing the names is not on the list.
This article walks the toolkit in plain words — redaction, masking, tokenization, pseudonymization, hashing, and aggregation — and it is deliberately honest about the limits. Several of these techniques are routinely oversold, hashing most of all. An admin who believes "we hashed it, so it's anonymous" is one file share away from an unpleasant discovery. You will come out knowing what each technique actually protects against, where each one fails, and where in your transfer pipeline to apply it. It is part of our Personal Data in File Flows series.
First, the Honest Frame
Nothing in this article makes data anonymous, with the partial exception of careful aggregation. These techniques narrow the audience that can identify people. It shrinks from "anyone who reads the file" down to "someone holding the key, the token table, or enough outside data to re-link it." That is genuinely valuable: it shrinks the blast radius of a misdirected file, limits what a nosy insider learns, and reduces harm when a recipient's laptop goes missing.
But disguised data whose disguise can be reversed is pseudonymized, not anonymized. What privacy laws typically expect is that pseudonymized data is still personal data, with obligations attached. Whether a specific dataset has crossed the line into genuinely anonymous territory is a judgment for your privacy or legal team. It depends on what other data exists in the world, not just on what you did to the file. Your role is to describe honestly what transformation you applied and what could reverse it. This article gives you the vocabulary for that conversation; it is education, not legal advice.
Remember: the question is never "is this data masked?" but "who could undo it, with what effort?" Every technique below is an answer to that question, and the answers differ enormously.
Redaction and Masking: Hiding in Plain Sight
Redaction removes a value outright — the field is blanked or deleted. It is minimization's twin, applied at the value level, and it is the right choice whenever the recipient needs the row but not that field. Masking is redaction's gentler sibling: part of the value is obscured, part left visible. This is usually so a human can still recognize or verify the item without seeing the whole thing. The familiar pattern is showing only the last few digits of an account number. Here is a synthetic customer row, before and after a masking pass — every value invented:
# fictional sample - all values invented # before C-104,Alex Sample,alex@example.com,555-014-990,EX-88410-330 # after masking for a support-tools export C-104,A. S.,a***@example.com,***-***-990,****-330
What masking protects against: casual reading, over-the-shoulder exposure, screenshots in tickets, and the wrong recipient learning full values from a misdirected file. Support and verification workflows love it — "can you confirm the reference ending 330?" works without the agent ever seeing the full number.
Where it fails, and it fails often:
- The row is still identified. In the example above,
C-104survived untouched, so anyone with customer-system access still knows exactly who this is. Masking one column does nothing about the others — the lesson of the combination problem from recognizing personal data. - Partial values can still be unique. A masked phone number plus a postcode may match one person. The visible fragment carries more information than it seems.
- Bad masks leak. Masking the middle but keeping first and last characters of a short value leaves little to guess. Length-preserving masks reveal length; format-preserving masks reveal format.
- It is cosmetic against joins. Anyone who holds the full value can match it to the visible fragment and confirm a hit.
Verdict: masking is a presentation-layer control. Excellent for limiting what eyes see; nearly worthless as the sole protection on a file that leaves your control. Asterisks are not a cipher.
Tokenization: Swap the Value, Keep the Key at Home
Tokenization replaces a sensitive value with a stand-in — a token — and records the pairing in a mapping table that never leaves your side. The file that travels contains T-88071 where the account number used to be; only your mapping table knows which account that is. When results come back from the recipient keyed by token, you re-link them internally.
Done properly, tokens are random — generated with no mathematical relationship to the original value. That is tokenization's great strength: there is nothing to crack. A recipient, or a thief, holding only the file can stare at T-88071 forever; the answer is not derivable, because it exists only in your table. Do not derive the token from the value. Not a hash of it, not a checksum, not the account number backward.
The strength relocates the risk rather than abolishing it:
- The mapping table becomes crown jewels. Whoever reads it can undo every tokenized file you ever sent. It needs the tightest access control you have, and its own logging.
- Consistent tokens build profiles. If the same person gets the same token in every file, a recipient can accumulate a rich history about "T-88071" without knowing the name. That history can include purchases, complaints, locations. That accumulation is itself the kind of profiling privacy teams care about, and one strong outside clue re-attaches the name to the entire profile at once.
- Scope tokens per recipient. If two partners receive the same token for the same person, they can pool their files and join on it. Different token spaces per recipient prevent cross-partner joins you never intended.
Pseudonymization: The Umbrella Term
Pseudonymization is the general name for replacing identifying fields with consistent stand-ins — pseudonyms. That way, records stay linkable to each other but not, directly, to a person. Tokenization is one implementation; so is hashing, discussed next. The consistency is the point and the product. An analytics recipient can see that customer P-4471X returned three times and churned in month six. That is exactly what they need, with no name in sight.
Keep two facts glued together in your mind, because vendors and colleagues will regularly separate them. First: pseudonymization is a genuinely strong risk reducer, and what privacy laws typically expect is precisely this kind of safeguard for data that must flow. Second: it remains personal data, because a key exists, the quasi-identifier columns riding alongside may re-identify on their own, and outside data can finish the job. Pseudonymization changes how carefully data must be handled; it does not change whether privacy rules apply.
Hashing Identifiers — and Why It Is Weaker Than It Looks
Now the technique most often oversold. A hash function turns any input into a fixed-length fingerprint, with two famous properties. The same input always yields the same fingerprint, and you cannot run the function backward. (Our hashing explained article covers the mechanics.) The tempting logic follows: replace every email address with its hash, and you get consistent pseudonyms nobody can reverse — anonymization by math. The logic fails. I know because I once wrote it into a design document, and I would like that document back.
You cannot run a hash backward, but you do not need to when you can run it forward. Identifiers live in small, structured spaces. Phone numbers in a country follow a known format with a countable number of possibilities. A computer can hash every possible number and compare the results against your file. At that point, every hashed phone number is unmasked. This works because hashing is deterministic and the input space is enumerable. Attackers do not even have to do the computing per file. Precomputed lookup tables of hashes — rainbow-style tables, in the jargon — already exist. They cover phone-number formats, common passwords, and vast lists of real email addresses harvested from breaches. Matching your "anonymized" column against such a table is an eyeblink, not an attack.
The rule of thumb: hashing hides a value only when the value was unguessable to begin with. Identifiers are the opposite of unguessable — they are drawn from public, structured, enumerable sets. Hashing a random secret protects it; hashing a national ID number merely re-encodes it.
Does salting fix it? A salt is extra secret data mixed into every hash. With a strong secret salt, an outsider can no longer precompute or enumerate — they would need the salt. That genuinely helps, but notice what you have built: a keyed system whose security rests on a secret you hold. That is pseudonymization with a key, functionally like tokenization but weaker. If the salt ever leaks, every dataset hashed with it becomes reversible retroactively. Unlike a token table, you cannot delete "knowing the salt" from the world. Also, salted or not, identical inputs still produce identical outputs within a dataset, so linkage — and the profiling that comes with it — works exactly as before.
Honest verdict: hashed identifiers are pseudonymized at best, trivially re-identifiable at worst, and never anonymous. Treat "we hash the emails before sending" as the beginning of a risk discussion, not the end of one.
Bluewater Bank came within a sign-off of shipping exactly that. An analytics extract for an outside firm had its email column replaced by unsalted hashes, and the design note described the file as anonymized. The privacy officer, who had read notes like it before, asked for a demonstration rather than an argument. It took an analyst eleven minutes to hash the bank's own customer email list, match it against the extract, and put a name back on every row. The feed went live three weeks later with per-recipient tokens and a mapping table locked down on the bank's side. The word "anonymized" did not appear in the revised agreement. Eleven minutes is the figure I quote whenever someone tells me hashing is enough.
Aggregation: The Only Real Exit
The one technique that can genuinely leave personal-data territory is aggregation — replacing individual records with group statistics. It is covered in depth in data minimization. No row about a person exists in the output at all. Even here the caveats matter. Small groups betray their members (a "group" of one is a person). Enough overlapping aggregates published over time can be intersected to reconstruct individuals. Sensible minimum cell sizes and a privacy-team sign-off keep aggregation on the right side of its own promises.
The Toolkit at a Glance
One table, techniques against threats, with the honest failure modes attached:
| Technique | What it does | Protects against | Where it fails | Still personal data? |
|---|---|---|---|---|
| Redaction | Removes the value entirely | Any use of that value | Other columns still identify the row | Row usually yes, via remaining fields |
| Partial masking | Obscures part of the value | Casual reading, screenshots, shoulder surfing | Fragments stay unique; joins confirm matches | Yes |
| Tokenization | Random stand-in; mapping kept at home | Recipient or thief recovering values from the file | Mapping table breach; profile-building on consistent tokens | Yes — you hold the key |
| Hashing, no salt | Deterministic fingerprint of the value | Only accidental glances | Enumerable inputs; rainbow-style lookup tables | Yes — often trivially reversible |
| Hashing, secret salt | Keyed fingerprint | Outsider enumeration and lookups | Salt leak reverses everything, retroactively | Yes — keyed pseudonyms |
| Aggregation | Group statistics, no individual rows | Identification generally | Small cells; intersecting many releases | Can be no — with care and sign-off |
Where in the Pipeline to Apply the Disguise
Technique chosen, the next question is placement. The principle is simple: disguise as early as possible, always before the file reaches shared storage or leaves the machine. The diagram shows the pipeline with the right and wrong places marked. (The wrong place is the popular one.)
Three placement rules carry most of the value:
- Mask before staging. The moment an unmasked export lands in a folder that operators, sync agents, and backup jobs can reach, you have multiplied the copies you must protect. Generate the export, run the disguise, and let only the disguised output touch the transfer path. In Sysax FTP Automation, this maps onto pre-transfer processing. The scheduled job runs your masking or tokenization script first, then transfers only the script's output. If the script exits with an error, the transfer step never runs and the job's email notification tells you. Fail closed is the only acceptable failure mode for a masking step. The product supplies the sequencing and the alarm. The masking logic itself is and should be your script.
- Masking and encryption are different answers to different questions. Encrypting the file — say with OpenPGP before sending, as our encrypt-before-send guide describes — protects it from everyone except the intended recipient. The intended recipient decrypts and reads everything inside. Masking protects the data from the recipient too. A payroll feed to the payroll processor needs encryption, not masking; an analytics extract needs masking, then encryption on top for the journey. Most serious flows want both, in that order.
- Log around the exception. If some flow genuinely must stage unmasked data even briefly, isolate it in its own folder with its own accounts. In that case, make sure the server records every touch. Per-account access control plus activity logging to file and database in Sysax Multi Server answers the question that follows any incident. Exactly which accounts downloaded the unmasked file, and when?
The Gap Between Pseudonymized and Anonymized
The word "anonymized" gets used loosely, and the gap between the loose usage and the real thing has embarrassed enough organizations to deserve its own section. The recurring pattern starts with a dataset released or shared with names removed and identifiers replaced. Someone joins it against outside information — public records, social posts, another leaked dataset — using the quasi-identifiers that survived. Named individuals fall out. The failure is never the substitution itself; it is the dates, places, and rare attributes that rode along, exactly the combination problem from earlier in this series. The names were removed; the names came back.
Genuine anonymization means degrading those quasi-identifiers — coarsening, suppressing, adding statistical noise — until individuals blur into crowds no realistic effort can resolve. Doing that properly usually destroys the record-level utility that made the recipient want the data, which is why the honest real-world resolution is rarely "we anonymized it." It is "we pseudonymized it, minimized what rides along, encrypted it for the journey, restricted and logged access, and put contract terms around re-identification". That is a defensible package of partial measures, honestly described. When a colleague wants to write "anonymized" in a data-sharing agreement, route the word past the privacy team; it is a legal claim wearing a technical costume.
Choosing a Technique: A Quick Decision Path
For each field that survived minimization, walk this order:
- Can the field be dropped or coarsened after all? Then that — minimization outranks every disguise.
- Must the recipient contact or identify the person? Then the field stays clear by necessity; protect the file with encryption, tight access, and logging instead.
- Must records link across files or runs without identifying anyone? Tokenize, with per-recipient token spaces — or salted hashing if you accept its retroactive-leak risk and the linkability that remains.
- Must data merely look realistic, as for testing or demos? Mask or substitute with synthetic values; nothing real needs to leave at all.
- Does nobody need individuals? Aggregate, with minimum cell sizes.
Write the chosen technique into the flow's record — the transfer inventory, if you keep one — next to its field list. The disguise decision is part of the feed's design. The next admin should inherit it as documentation, not archaeology. The follow-on article, privacy by design for batch jobs, shows where that record lives in a well-built job.
Disguise Honestly, Then Protect Anyway
The toolkit is real. Redaction and masking limit what eyes see, and tokenization keeps the key on your side of the wall. Pseudonymization preserves linkage without names, and aggregation can leave personal data behind entirely. The overclaiming is real too, and hashing is where it concentrates. Enumerable inputs and ready-made lookup tables mean a hashed identifier is often one script away from being an identifier again. So disguise early in the pipeline, describe what you did precisely, and let the privacy team make the legal calls. Keep the boring protections — encryption, least privilege, logging, short retention — wrapped around the disguised file as if the disguise were not there. When it fails, and sometimes it will, those are the controls that turn a story into a non-event. And when someone tells you the test data is fine because they changed the names, ask what they did with the postcodes.
Frequently Asked Questions
Is hashing email addresses enough to share a customer list safely?
What is the difference between tokenization and encryption?
Is masked or pseudonymized data still personal data?
Can we describe our shared data as "anonymized" in a contract?
Where should the masking script run — on the source system or the transfer server?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
