Home › Topics › Retention & Deletion › Retention Basics

Data Retention Basics for File Transfer Admins

The folder is called outbox\payroll_old. It holds every payroll extract the server has ever sent, and its only documentation is a text file left by an administrator two managers ago: "keep for now." Nobody has opened it in years. Nobody will delete it either, because deleting feels like a decision and keeping feels like none. Keeping is the decision. It is a decision to secure that data forever, and to lose it in any breach you ever have. It is a decision to hand it over when a court, regulator, or data subject asks what you hold.

Retention is the discipline that replaces "keep for now" with an actual answer. This data stays for this long, for this reason, and then it goes. The good news for administrators is that most of the hard part is not your job. Someone else decides how long; you make the decision real on disk. The skill you need is smaller and very learnable: understand the vocabulary, know who owns which decision, and read a retention schedule without a law degree. (The law degree is genuinely optional. The vocabulary is not.)

That is this article. It is part of our Retention & Deletion series. It sits between the mapping exercise that finds your data and the automation that cleans it up: the "what should happen" layer in the middle.

What Retention Actually Means

Strip the jargon and retention is just the answer to one question, asked per kind of data: how long do we keep this? A retention period is that answer expressed as a duration — ninety days, six years, "until the contract ends plus three years." When the period expires, disposition happens: the planned end-of-life action. Usually that is deletion, but it can also mean moving data to an archive or transferring it to another system.

Two more terms carry most of the weight in real schedules. A record class is a category of information that shares one retention rule — "invoices," "payroll registers," "job applications." It is defined by what the data is, never by where it happens to be stored. And a trigger event is the moment the retention clock starts. Sometimes it is creation ("keep for two years from creation"), often a business event ("keep for six years after the account closes"). Miss the trigger concept and every event-based rule in a schedule will confuse you.

One distinction matters more to a transfer administrator than any other: the difference between a record and a copy in transit. The record is the authoritative version of the information, living in its system of record — the payroll application, the accounting system, the document management platform. What crosses your transfer server is almost always a copy made for conveyance: an extract, an export, a feed. The record class rules bind the record wherever it lives, but they do not require the conveyance copy to live equally long. The cleanest estates keep transfer copies only briefly, precisely because the system of record already satisfies the long obligation.

Why "Keep Everything Forever" Is a Liability

Storage is cheap, so the cost of keeping everything looks like zero. It is not. The costs are indirect, they arrive at the worst possible moments, and they arrive together.

Every retained file widens your breach. When an attacker gets into a transfer server, they take what is there. If the outbox holds one week of payroll extracts, the incident involves one week of data. If it holds five years, the incident involves five years — five years of employees to notify, regulators to face, and headlines to survive. Retention is the only security control that shrinks the blast radius of a breach you have not had yet. Our threat-modeling series calls this the exposure window: an attacker's haul is bounded by what you kept. The security-side view of cleaning up server-side data is covered in server-side at-rest protection.

Everything you hold is discoverable. In a lawsuit, each side can demand the other's relevant documents — the process lawyers call discovery. The demand does not stop at the system of record; it covers relevant data wherever it exists, including forgotten folders of transfer copies. Holding a decade of miscellany means paying people to search a decade of miscellany. It may mean handing an opponent something you had no obligation to still possess.

Privacy laws set maximums, not just minimums. GDPR-style regimes include a storage limitation principle: personal data may be kept no longer than its purpose requires. For personal data, "forever, just in case" is not a conservative choice — it is the non-compliant one. The same logic appears in PCI DSS, which expects cardholder data storage to be minimized and time-limited. Retention duties cut both ways: some rules say "at least this long," others say "no longer than needed," and a defensible schedule respects both.

Hoards hide things. A server holding only what it needs is a server where anomalies stand out and access reviews are quick. A server holding everything since forever is one where nobody can say what is sensitive or what is duplicated. Nobody can say what that seventy-gigabyte folder even contains. The article on why transfer servers fill up explains how it got that big. The folder is called misc, and that is the whole of its documentation.

Deleting too early is also a failure, and honesty requires saying so. Destroying data you were required to keep — or data under a legal hold — is far worse than over-retention. The goal is not minimal retention; it is deliberate retention, with each period chosen on purpose and applied consistently.

Who Decides, Who Implements

Setting retention periods is not the administrator's job, and that one sentence takes most of the weight off your shoulders. Retention periods encode legal obligations, regulatory expectations, and business judgment. Deciding them is the work of legal counsel, records management, or a compliance function, together with the business owners of the data. Articles like this one can explain the concepts. But what your organization must keep and for how long is a question for those teams. That division of labor is not bureaucracy, it is protection for you.

What the administrator owns is everything after the decision, and that half is substantial:

  • The inventory. Nobody can set periods for folders they do not know exist. Your map of where transferred files accumulate is the input the deciders need, and producing it is squarely your craft.
  • Feasibility. You know whether "delete on contract end" is implementable when files carry no contract metadata. You can push a rule toward something enforceable — "age since the file landed" — before it is signed off.
  • Faithful implementation. Turning each approved period into an automated job that runs without a human remembering, which is the whole subject of automated purge policies.
  • Evidence. Being able to show what was deleted, when, by which job — so compliance is demonstrable rather than asserted.

The handshake between the two halves should be written. A one-line email can say "Finance confirms: payroll extracts in outbox\payroll may be deleted ninety days after upload." That converts a risky unilateral cleanup into an implemented business decision. Never purge on your own authority alone; never let "nobody answered my email" silently become "so I kept it forever," either. Escalate until each folder has an answer. I once kept a folder for three years on the strength of one unanswered email, which is not a retention period. It is a stalemate.

Kestrel Payroll ran a tidy transfer server with one exception: outbox\clients_archive. Every client's payroll extract had been copied there "temporarily" since the server was built. The administrator had asked Finance for a retention period once, received no reply, and moved on to things that did reply. Six years later an audit walkthrough asked what the folder held. The honest answer took two people three days and a spreadsheet. Nobody could say which clients, which pay periods, or whether any of it was still under contract. Finance settled on ninety days in a single meeting once someone showed them the folder's size. The period had never been hard to decide. It had only never been asked twice.

How Retention Periods Get Chosen

You do not choose the numbers, but understanding where they come from makes schedules stop looking arbitrary. Four forces set most periods:

  • Legal minimums. Tax and corporate law require certain records — invoices, ledgers, contracts — to be kept for fixed spans, commonly six or seven years depending on jurisdiction. These produce the "at least" rules.
  • Regulatory frameworks. Regimes like HIPAA and SOX require certain documentation and records to survive for defined periods. Audit regimes expect logs and evidence to be available for review. These also produce "at least" rules, often aimed at proof rather than payload.
  • Privacy maximums. Storage limitation under GDPR-style laws, and data minimization expectations in PCI DSS, produce the "no longer than" rules for personal and cardholder data.
  • Business need and dispute horizon. How long might we genuinely need to re-send, reconcile, or defend this? Statute-of-limitation thinking lives here: keep contract records while a claim about the contract is still possible.

A given record class can sit under several forces at once, which is why schedules are built by people who can weigh them. Your takeaway is simpler: every period in a schedule has a reason. When a period seems strange, asking "what drives this one?" is a legitimate administrator question that usually gets an illuminating answer. (Keep the answer. It is the one file nobody will argue about retaining.)

Reading a Retention Schedule Without a Law Degree

A retention schedule is the document where all those decisions live: a table of record classes, each with its period, trigger, and disposition. Schedules look intimidating because they are long, but every row has the same anatomy. Here is a worked excerpt of the kind you might be handed, trimmed to the rows that touch a transfer server:

Record class Retention period Trigger Disposition
Customer invoices Seven years End of fiscal year issued Delete from all systems
Payroll registers Six years Pay period end Delete; system of record: payroll app
Job applications (not hired) One year Position filled Delete everywhere, incl. copies
Transfer conveyance copies Ninety days Date file landed Automated purge, logged
Server activity logs Two years Log entry date Delete after archive

Read any row in four moves. First, what is it? The class describes content, so ask which of your folders receive that content — this is where your accumulation map earns its keep. Second, how long? The period, written here in words, becomes a number your automation can enforce. Third, since when? Creation-triggered rules translate directly to file age. Event-triggered rules ("position filled") need a pragmatic translation on a transfer server, usually a conservative age-based stand-in agreed with the owner. Your file system does not know when a position was filled. Fourth, then what? Disposition tells you whether expiry means delete, archive elsewhere, or hand off.

Notice what the last two rows do. The schedule explicitly names conveyance copies as their own class with a short period. That one row is your authority to clean the transfer server without touching the seven-year duty, which the system of record carries. And it treats logs as a record class too, because deciding how long to keep evidence is itself a retention decision. On a server such as Sysax Multi Server, activity can be logged to file and to a database with rollover into dated files. That makes "archive then delete after two years" an implementable rule rather than a wish. What belongs in those logs in the first place is covered in what to log.

Remember: the retention schedule binds the information, not the folder. If invoices must live seven years, that duty is satisfied by the accounting system — not by your outbox. The transfer server's job is to hold copies briefly and verifiably, then let them go.

The Transfer Server Is Not the System of Record

The transfer server is not the system of record, and that one principle simplifies almost everything. When the authoritative copy lives in an application, the transfer server's copies need to survive only long enough to serve the transfer. That means redelivery if the partner asks again, troubleshooting if something arrived corrupted, and reconciliation if counts disagree. For most flows that is somewhere between thirty and ninety days. That is long enough that any dispute about a given exchange has surfaced, short enough that the server never becomes a shadow archive. Nobody plans a shadow archive. It is just an outbox nobody emptied.

There are honest exceptions. Sometimes the transfer archive is the record. That might be a regulated submission where you must prove exactly what was sent. Or it might be a flow where the receiving system transforms data and the original file is the only evidence of the input. Those folders take the record's full retention period. They deserve stronger evidence around them: hashes, receipts, and the kind of documentation our evidence pack article describes. The point is not that transfer folders are always short-lived. It is that "conveyance copy, short" should be the default. And "this folder is the record" should be a labeled, justified exception.

A useful test when someone insists a transfer folder must be kept for years: ask what question the folder would answer that the system of record cannot. If there is a real answer, write it into the schedule as an exception with its own row. If the answer is "well, just in case," you are looking at fear, not a requirement. I have asked that question in a dozen meetings and heard a real answer twice; both got their row.

From Schedule Row to Server Reality

The path from a signed schedule to a clean server is short, and it is the same five steps for every folder:

1. Inventory   List the folder, what lands in it, and its owner.
2. Classify    Match the folder's contents to a schedule row
               (owner confirms in writing).
3. Translate   Express the rule as the server can enforce it:
               "delete files older than ninety days, daily run."
4. Automate    Build the purge as a scheduled job with logging
               and a dry-run first.
5. Evidence    Keep the job's logs; review quarterly that it
               still runs and still matches the schedule.

Step four is where tooling enters. A purge is nothing exotic — a scheduled task that finds files past their age and removes them. An automation product like Sysax FTP Automation covers the mechanics with scheduled tasks that perform file and folder operations and send email notifications with each run's results. So deletion happens on calendar time and leaves a trail. The design details — grace periods, dry runs, and the monitoring that catches a stalled job — are the subject of the next article in this series.

Some folders resist classification: mixed content, unknown origin, an owner who left. Do not let them stall the whole program, and do not quietly delete them either. The pragmatic pattern is a written interim rule. Under that rule, the folder is frozen (no new writes), and its contents are reviewed by the closest thing to an owner you can find. Anything unclaimed after an agreed review window is treated as expired conveyance data. Even that modest rule should be someone's signed decision, not your improvisation. An unclassifiable folder handled by a documented process is a finding being fixed. The same folder deleted on instinct is a story you may have to tell a lawyer.

When Deleting Is Forbidden

One instruction overrides every schedule row and every purge job: a legal hold. This is the notice that litigation or an investigation requires certain data to be preserved untouched. When a hold arrives, deletion of anything in its scope must stop — including automated deletion. That is the part transfer admins forget until it is urgent. Holds have their own article in this series, legal holds and exceptions. For now, plant the flag: every purge you automate must have a way to be paused, per folder, quickly.

The Short Version

Retention is a deliberate answer to "how long do we keep this?" — per record class, with a trigger and a disposition. Keeping everything forever is not safety; it is unbounded breach exposure, discovery burden, and, for personal data, non-compliance in itself. The periods are chosen by legal, records, and business owners. Your role is the inventory before the decision and the faithful, evidenced enforcement after it. Read schedule rows in four moves — what, how long, since when, then what. Default transfer folders to short conveyance retention, because the system of record carries the long duties. Then make it real: build purges that actually run. Fold the whole thing into a one-page policy for your server that answers auditors before they finish asking. The text file that says "keep for now" can go first.

Frequently Asked Questions

Isn't keeping everything the safest option?
No. Everything you keep can be stolen in a breach and must be searched in a lawsuit. If it is personal data, keeping it may itself violate storage limitation rules. Keeping data past its need does not add safety; it adds surface. The safe position is a written period for each kind of data, followed consistently.
What is a typical retention period for transfer server folders?
There is no universal number. Conveyance copies — files passing through on their way to a system of record — are commonly kept thirty to ninety days. That is enough to cover redelivery and disputes. Folders that serve as the authoritative record of what was sent take the record's full period instead. Your schedule, not a rule of thumb, is the authority.
Who decides how long we keep data?
Legal counsel, records management or compliance, and the business owner of the data — together. The administrator's role is to supply the inventory of what exists and sanity-check that a rule is enforceable. The administrator implements it with automation and keeps evidence that it ran. If you are setting periods alone, escalate; that decision needs owners.
What is a trigger event in a retention schedule?
The moment the retention clock starts. Some rules run from creation ("two years from creation"), others from a business event ("six years after the account closes"). File servers only know file ages, so event-based rules usually need an agreed age-based translation before automation can enforce them.
If I delete a file from the server, is it gone from everywhere?
No. Copies typically remain in backups until the backup cycle expires them, and possibly in other systems the file passed through. A retention schedule should state how backup copies age out, and any statement you make about deletion should be honest about that timing.

From the Sysax team: we build secure file transfer software for Windows — Sysax Multi Server, an FTP, FTPS, SFTP, and HTTPS server, and Sysax FTP Automation for scheduled, scripted transfers. Free trials are on the download page.