When Access and Deletion Requests Meet File Transfers
The email arrives from the privacy team: "We've received a request from a customer named Alex Sample. Please confirm what personal data about them exists on the transfer servers, and be ready to delete it." The database teams answer in an hour — their data lives in tables you can query. Then everyone turns to the file transfer estate, and the room goes quiet. Which of the thousands of CSVs in the partner folders mention this person? What about the staging areas, the archives, the zipped batches from two winters ago? I have been the person the room turned to. My honest answer that day was "give me a week."
This article is about making that moment routine instead of terrifying. It covers what subject rights requests are in plain words and why transfer estates are genuinely the hard part of answering them. It provides a where-to-look checklist and honest search techniques with their blind spots. It covers how deletion actually works when copies are scattered and the hygiene that turns future requests into small tasks. It is part of our Personal Data in File Flows series.
Subject Rights in Plain Words
Most privacy regimes give individuals — data subjects, in the jargon — enforceable rights over data about them. The recurring set includes a right of access (tell me what you hold about me, and give me a copy). It includes a right of correction (fix what is wrong). It also includes a right of deletion, sometimes called erasure (remove my data, subject to exceptions). GDPR is the famous example, but variations of these rights appear across many laws. They typically come with a deadline: the organization must respond within a fixed period. It is long enough to do real work but short enough that you cannot start from zero.
Division of labor matters here, and it protects you. The privacy or legal team owns the request. They verify the requester's identity, interpret which rights apply, decide what falls inside scope, and apply the exceptions the law allows. The administrator owns the finding and the acting: searching the estate, compiling what exists, and deleting what the privacy team says to delete. You should never be deciding on your own whether a deletion request is valid. This series, as always, is education about what privacy laws typically expect, not legal advice about your obligations.
Why the Transfer Estate Is the Hard Part
Databases hold one authoritative copy of a record, indexed and queryable. Transfer estates hold the opposite: point-in-time copies of exported data, scattered across folders, in whatever format each flow happened to use. One customer's information might exist as a row in last night's courier feed and ten rows across a month of nightly marketing extracts. It might appear as a mention in a support-ticket attachment, a line item inside a zipped batch archive, and a name embedded in a filename. All describe the same person, none aware of the others. Each copy is quietly convinced it is the only one.
Three properties make the search genuinely hard:
- Duplication without an index. Nobody tracks which files mention which people. The information is there; the lookup structure is not.
- Format opacity. Plain CSVs are searchable. Zipped batches, spreadsheets (which are compressed XML inside), and PGP-encrypted files are invisible to naive text search. We will return to this point, because assuming your search saw everything is the classic failure.
- Identifier drift. The same person appears as
Alex Sample,SAMPLE, ALEX,asample@example.com, and customerC-104in different flows. These are the direct and indirect identifiers cataloged in recognizing personal data. A search for one variant finds a fraction of the truth.
To feel the scale, trace one file. A nightly customer extract is written by the database server, then copied to the automation machine's staging folder. It is uploaded to the partner's folder on the transfer server, archived to a sent tree, and swept into that night's backup. That is five copies from one run, before the partner makes theirs. Multiply by nightly runs and by every flow that mentions the same person. Now "find everything about Alex Sample" means combing a population of files that no one has ever counted. That is not a reason for despair; it is the reason the rest of this article runs map first, search second, and habits last. Nobody has counted the files, and nobody is going to start tonight.
The Where-to-Look Checklist
When a request lands, work through the estate systematically rather than by memory. This is the map; adapt the paths to your environment and keep the adapted version in your runbook:
- Partner and user folders on the transfer server. The delivered files themselves — per-account directories, pickup and drop-off areas. Check both directions: files you sent and files received about the person.
- Staging and temp areas. Wherever exports rest between production and transfer, including the automation machine's working folders and any
temp,work, orfaileddirectories where broken runs park files. - Automation job folders. Each scheduled job's source and output paths — read them from the job definitions rather than guessing. Note jobs whose configs mention the subject's email address directly (notification recipients are personal data too).
- Archive and retention folders. The
sent,processed, and dated archive trees where completed transfers accumulate, often for years. - Per-user home directories and download areas. Staff who downloaded exports to "just check something" created copies outside every controlled path.
- Transfer logs. Both as evidence (which files moved where) and as data — logs mention usernames, addresses, and filenames. If the subject is a staff member or account holder, log entries themselves are their personal data.
- Backups. Every location above, multiplied by the backup schedule. You rarely search backups file by file; you account for them by policy, as covered under deletion below.
Knowing your flows is what makes this list answerable at all. If you maintain a file flow census or a transfer inventory with a personal-data verdict per flow, you can shortlist the dozen flows that carry customer data. You can search those thoroughly, instead of grepping the entire estate and hoping. The census is the index the file system never had.
Kestrel Payroll's first subject access request took nine working days, and eight of them went on finding out where the copies were. The person turned up in fourteen monthly client feeds and a failed folder nobody had emptied. They also appeared in two staff download directories, a zipped year-end batch, and — on day seven — the filename of a ticket attachment. Nine places, none of them indexed, each discovered by someone remembering it existed. The privacy officer got an honest answer on day nine. Kestrel wrote its flow census on day ten, and the next request took an afternoon.
Searching Honestly: What Text Search Can and Cannot Find
For the searchable portion of the estate, simple tools go a long way. Collect the subject's identifier variants from the privacy team first — email addresses, customer IDs, name spellings, phone number. Then sweep with something like:
REM quick sweep for one identifier across text-like files (Windows) findstr /s /i /m "asample@example.com" D:\Transfers\*.csv D:\Transfers\*.txt # PowerShell: several identifiers at once, listing matching files Get-ChildItem D:\Transfers -Recurse -Include *.csv,*.txt,*.log | Select-String -List -Pattern "asample@example.com","C-104","555-014-990" | Select-Object -ExpandProperty Path
Then respect the blind spots, because they decide whether your answer to the privacy team is true:
- Compressed and container formats. Zip batches and spreadsheets will not match plain text search. Expand archives to a scratch area and search inside, or use tooling that reads containers. List the formats you could not open rather than pretending they were covered.
- Encrypted files. A PGP-encrypted export is unreadable by design — that is the point of encrypting before sending. For archives your organization holds keys for, decrypt copies into a controlled scratch area and search there. A scripted decrypt step (the OpenPGP support in Sysax FTP Automation can do this as a processing step in a job) beats hand-decrypting forty files. Files encrypted to a partner's key are answerable only via logs and manifests: you can prove what batch went where, even when you cannot read it back.
- Pseudonymized and tokenized data. A file where the subject appears as token
T-88071will never match a search for their name. You must first consult the token mapping. This is one more reason the mapping table from masking and pseudonymization is a governed asset, and a reminder that disguised data is still their data. - Free-text mentions. "Spoke to Mr. Sample's daughter" matches no identifier list. Accept that text search finds identifiers, not meaning, and say so in your report.
Narrow by time before you sweep; it multiplies every other technique. If the subject became a customer in a known month and closed their account in another, flows that only carry current customers cannot mention them outside that window. In that case, date-stamped archive trees outside it drop off the search list. The archives inside it get the careful treatment, expansion and decryption included. Reading delivery dates out of your transfer logs, as described in reading transfer logs, is often the fastest way to establish that window. It is also often the fastest way to prove afterward that the skipped ranges were skippable for a reason, not from fatigue.
Remember: a search proves presence, never absence. Report findings as "these locations contained matches; these formats could not be searched and are accounted for by policy". That honest sentence is worth more to your privacy team than a confident "nothing else exists" that nobody can stand behind.
Answering an Access Request
For access requests, your product is a compilation of the matching files or extracts. Label each with where it lived, which flow produced it, and the date range it covers. Hand the compilation to the privacy team — they assemble the formal response, redact other people's data from shared files, and deal with the requester. Two operational cautions. First, the compilation is itself a concentrated file of one person's data — the most sensitive artifact of the whole exercise. So move it through a controlled, encrypted path with named access, never as a casual email attachment. Second, give it a deletion date of its own. The answer to a privacy request should not become the newest scattered copy. (It has happened. It was found during the next request.)
Correction Requests: The Quiet Middle Case
Files in your estate are snapshots — accurate, or not, as of their export moment — and that gives correction requests a transfer-specific wrinkle. The correction itself happens in the source system, owned by whichever team runs it. Your part comes in two questions. First: will the fix propagate? If the feed sends full refreshes, the next run carries the corrected value automatically. If it sends deltas, confirm the correction triggers a delta, or the partner keeps the wrong value forever. Second: do any already-delivered files need chasing? Usually the answer is no — historical snapshots are understood to be historical. But when the error is harmful (a wrong flag that affects decisions about the person), the privacy team may ask you to act. They may ask you to send a correction file or ask the partner to update their copy. Knowing which flows are refresh-style and which are delta-style, per your flow records, turns that conversation from research into recall.
Answering a Deletion Request
Deletion is where scattered copies bite hardest, and where discipline matters most. The sequence that works:
- Get scope in writing. The privacy team tells you what to delete and — just as important — what to keep. Deletion rights typically have exceptions: records under legal hold, data needed to meet other retention duties, security logs. Never free-lance the scope; deleting a file that was under hold is its own incident.
- Delete from live locations first. Partner folders, staging, archives, home-directory strays — the copies your search found. Remember that "delete" on most storage means "stop referencing". For genuinely sensitive material, follow your organization's secure-deletion practice, a subject our retention and deletion series treats properly.
- Handle backups by policy, not surgery. Editing individuals out of backup sets is usually impractical and risky. The widely accepted operational pattern is documented aging. The deleted data disappears from restorable history as backup generations expire. If a restore happens meanwhile, a documented procedure re-applies the deletion. Whether that pattern satisfies a given request is — again — the privacy team's call to make and defend. Your part is stating the backup retention facts precisely.
- Record what you did without re-creating the data. Keep a deletion log: file paths, flow names, timestamps, who performed it — not the contents you deleted. The log proves compliance; contents would resurrect the problem.
The Hygiene That Makes Requests Small
Every practice this series recommends quietly shrinks the subject-rights problem, and it is worth seeing why as a system:
- Short retention is pre-answered deletion. Consider staging purged by the job after delivery and archives aged out on schedule. Each purge that already ran is a folder you no longer search and a copy you no longer delete. The staging-cleanup habit from privacy by design for batch jobs earns its keep here.
- Minimized feeds produce fewer hits. A courier file without birth dates and emails is a file that may not even fall in scope. Less data sent is less data found.
- Structured flows beat sprawl. One folder per flow per partner, no personal data in filenames, no ad-hoc side copies. On the server, per-account folder jails — the per-account access control model in Sysax Multi Server — mean each flow's files live where its account lives. So "where could this feed's files be?" has one answer. The server's activity logging, written to file and to a database, adds the searchable half. Query the log database for a filename and you get every account that downloaded it and when. This is exactly the "who else has copies?" question a deletion request raises.
- The flow census is your index. Keep the personal-data verdict per flow current, and the shortlist step of every future request is already done.
Run the Drill Before the Real Request
The difference between a calm response and a scramble is one rehearsal. Invent a subject — a made-up customer with a made-up email. Seed a few test files, and run the runbook end to end against the clock:
SUBJECT REQUEST RUNBOOK - TRANSFER ESTATE 1. Receive scope + identifier variants from privacy team (in writing) 2. Shortlist flows carrying personal data (flow census) 3. Sweep live folders, staging, archives with identifier list 4. Expand containers; decrypt archives we hold keys for; search again 5. Check token/pseudonym mappings for indirect appearances 6. Query transfer log database: deliveries + downloads of matched files 7. Compile findings with locations; note unsearchable formats + policy 8. Access: hand compilation to privacy team via secure path, set its deletion date 9. Deletion: confirm scope/exceptions in writing, delete, log actions 10. Note the time each step took; fix the slowest before next time
The drill always finds something — an archive tree nobody remembered, a format the sweep missed, a mapping table only one person can read. Finding it on a quiet afternoon, with no deadline running, is the entire point. Mine turned up a folder called keep; it had been obeyed for six years.
From Dread to Routine
Subject rights requests feel threatening when the estate is unmapped and copies are everywhere. They become routine when flows are inventoried, retention is short, disguises are documented, and logs can be queried. Answer requests with honest search plus honest statements about the limits of search. Let the privacy team own scope and exceptions. Invest the lessons of each request back into hygiene. One companion piece to read next is data minimization — the less you send, the less you will ever have to find. The other is privacy by design for batch jobs, where the purge steps and flow records that make this article easy are built. The next time the room goes quiet, let it be because everyone is reading the runbook.
Frequently Asked Questions
How long do we have to answer a subject request?
Do we have to delete a person from backups right away?
What about files encrypted to a partner's key that we cannot read?
Should I start deleting as soon as a deletion request arrives?
Are transfer logs in scope, and can we keep them?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
