Home › Topics › Idempotency › Idempotency

Idempotency in Plain Words: Making File Jobs Safe to Run Twice

Sooner or later, every administrator who runs file transfer jobs meets the rerun question. A nightly job died at step three of five. You fixed the cause, and now you are wondering: is it safe to run this again? For some jobs the answer is obviously yes — run it, shrug, move on. For others, a second run means duplicate invoices, a partner receiving the same batch twice, or a report emailed to two hundred people again. Most of us learn which jobs are which by getting burned.

The property that separates the safe jobs from the dangerous ones has a name: idempotency. A job is idempotent when running it twice — or five times — leaves the world in exactly the same state as running it once. It is one of those ideas that sounds academic until you own a pipeline at two in the morning. At that point, it becomes the most practical idea you know. Idempotent jobs turn failures into non-events: something broke, run it again, done. Non-idempotent jobs turn failures into investigations.

This article is the plain-words foundation for our duplicate detection and idempotency series. By the end you will be able to say precisely what idempotency means for a file job. You will be able to explain why reruns happen constantly no matter how good your automation is. You will recognize the specific operations inside a job that make reruns dangerous. You will also be able to run a short self-test against any job you own to find out where it stands.

Run It Twice, Get the Same Result

Start with the definition, because it is short and worth memorizing. An operation is idempotent when performing it once or performing it many times produces the same final state. Not "does nothing the second time" — produces the same final state. The second run is allowed to do work; it just is not allowed to change the outcome.

Everyday life is full of both kinds of operation, which makes for easy analogies. Setting a thermostat to 68 degrees is idempotent: set it once or set it five times, the room ends up at the same temperature. Turning the thermostat up by one degree is not idempotent: each press changes the state again. Marking an email as read is idempotent; sending an email is not.

Translate that to file work and the pattern holds. Copying report_YYYYMMDD.csv to a destination folder, overwriting whatever is there, is close to idempotent. After the second copy, the destination holds exactly what it held after the first. Appending that file's contents to a master log is not idempotent. Neither is inserting the file's rows into a database table with no duplicate protection. Run either operation twice and every row is there twice. Sending the file to a partner who loads whatever arrives is not idempotent from the partner's point of view. That is true even though your side looks identical after each run.

That last example points at the real-world subtlety: a job is only as idempotent as its most fragile effect. A transfer job is never just "move bytes from A to B." It moves bytes, and then something downstream reacts — a loader ingests, a report generates, a notification fires, a file count increments. When you ask "is this job safe to run twice?", you are really asking that question about every effect in the chain. The chain is longer than the script.

Remember: idempotent does not mean read-only, and it does not mean the second run is skipped. It means the second run cannot change the outcome. A job that re-copies a file over an identical copy did real work — and changed nothing.

Why File Jobs Get Run Twice Whether You Like It or Not

A common first reaction is: "my jobs run once a night, on a schedule — why would they ever run twice?" The honest answer is that duplicate runs are woven into how automated transfer works, and they arrive from more directions than most people expect.

  • Retries. Any competent automation retries failed transfers, because most transfer failures are transient network hiccups. But a retry is, by definition, a second attempt at something that may have partially or even fully happened. Retry logic is the single largest source of duplicate work in file pipelines. That is why our retry and error handling series and this one are close companions.
  • The timeout that actually succeeded. Here is the nastiest special case of a retry. A client uploads a file, and the upload completes on the server. But the confirmation never reaches the client — a dropped connection, a timeout waiting for the final reply. The client, having seen no success, honestly reports failure and retries. The server now receives a second copy of a file it already has. Nobody misbehaved; the duplicate is the natural result of both sides acting correctly on the information they had. We dissect this scenario step by step in where duplicate files actually come from.
  • Sender resends. Partners re-export "just to be safe," re-run their own failed jobs, or resend last week's batch because someone asked a question about it. You do not control the sending side, so you cannot prevent this — you can only survive it.
  • Scheduler behavior. Missed schedules get caught up, a paused job fires immediately on resume, a reboot re-triggers a startup task. The scheduler is doing its job; your job just runs more often than you pictured.
  • Nervous admins. The most human source. The job status is unclear, the log is ambiguous, the business is waiting — so someone runs it again to be sure. In a healthy team this happens weekly; design for it rather than lecturing people out of it.
  • Deliberate reprocessing. Sometimes you want to run yesterday's files again — a bug in the transform, a corrupted output, an audit request. Planned reruns deserve their own machinery, which is the subject of safe reprocessing.

Notice that every one of these is either good practice (retries, catch-up schedules), outside your control (resends), or inevitable human behavior (the nervous rerun). You cannot get rid of duplicate runs. The only durable position is to make duplicate runs boring.

An Anatomy Lesson: Where a Rerun Actually Hurts

To make a job rerun-safe, you need to see it not as one action but as a chain of small operations. Each is either idempotent or not. Here is a typical nightly job, decomposed:

  1. Download every file matching settle_YYYYMMDD.csv from the partner's server.
  2. Move each downloaded file into an incoming folder.
  3. Load each file's rows into a database table.
  4. Move the file to an archive folder.
  5. Email a summary: "3 files, 1,240 rows loaded."

Run that twice and walk the chain. Step 1 downloads the same files again — harmless by itself, if the download overwrites. Step 2 overwrites again — still fine. Step 3 inserts every row a second time — this is the wound. Step 4 archives again — fine. Step 5 emails a second summary claiming another 1,240 rows loaded — a small lie that will confuse someone during the eventual investigation. One step out of five turned a harmless rerun into double-counted settlement data.

The pattern generalizes well enough to put in a table. When you review a job, classify each operation:

Operation Run twice and you get Idempotent?
Copy to a fixed name, overwrite allowed The same file in the same place Yes
Copy with an auto-renaming collision rule (file_(2).csv) Two files where there should be one No
Append rows to a table or log Every row duplicated No
Replace a table's contents from the file The same table contents Yes
Delete a specific file The file gone either way (second run may error) Yes, in effect
Send an email, post a message, call an API that creates something Two emails, two messages, two created things No

Two observations fall out of this table. First, "set the state" operations are idempotent and "change the state" operations are not. Overwrite, replace, and delete describe a destination; append, increment, and create describe a change relative to whatever is there now. Prefer the describing kind whenever you have the choice. Second, the dangerous steps cluster at the end of the chain — the loads, the notifications, the handoffs. That is exactly where the protections in the rest of this series apply.

The Habits of Rerun-Safe Design

Idempotency in a file pipeline is not one feature you switch on. It is a handful of small habits that compound. These are the ones that matter most, roughly in the order you should adopt them.

Give every file one stable identity

A rerun is only detectable if the second copy of a file looks like the first. That starts with naming: a file should get exactly one name at creation — settle_YYYYMMDD.csv, export_YYYYMMDD_HHMMSS_0042.csv. It should keep that name through the pipeline. Names that embed a generation timestamp of the run rather than the data defeat this. The same data resent an hour later arrives under a new name and sails past every name-based check. Our file naming and datestamping series covers conventions that hold up. The short version is that the name should identify the content, not the attempt.

Overwrite, don't accumulate

Wherever a step writes to a fixed location, let the second run overwrite the first — same input, same output, no harm. Be suspicious of any tool setting that "helpfully" renames on collision, because it silently converts an idempotent overwrite into duplicate accumulation. And when a step must append or insert, that step needs a guard. The guard checks against a record of what has already been processed before the step acts.

Keep a record of what has been processed

That record is the processed-files ledger — a durable list of every file the pipeline has already handled. It lives in a flat file or a small database table, keyed by the file's identity. Before the dangerous step runs, the job checks the ledger; if the file is already there, the step is skipped and the skip is logged. The ledger is the workhorse of practical idempotency, and it earns a full article of its own. The article detecting duplicates with names, sizes, hashes, and ledgers covers how to build one honestly. It includes the locking details that keep the ledger trustworthy.

Make each step finish atomically

Reruns interact badly with half-finished work. If a crash can leave a half-written output file, the rerun must be able to overwrite it cleanly. That is why rerun-safe pipelines write to a temporary name and rename into place at the end. That way, any file with a real name is complete. That family of techniques belongs to our partial-file safety series, but note the connection. Atomic writes make reruns safe at the file level, and ledgers make them safe at the processing level. You want both.

Key the side effects, not just the files

The email step in our anatomy example stays dangerous even after the load step is guarded. The fix is the same idea one level up: record that the notification for settle_YYYYMMDD was sent, and check before sending again. Any side effect worth doing once is worth recording that it was done.

The Rerun-Safety Self-Test

Here is the practical core of this article: eight questions to ask about any job you own. Answer them honestly against the job as it is, not as you believe it should be. A "no" is not a crisis — it is a located risk with a known fix.

RERUN-SAFETY SELF-TEST — score one job at a time

 1. If I ran this job again right now, would any data
    be counted, loaded, or billed twice?          [ ] no = pass
 2. Does every file keep one stable name from
    creation to archive?                          [ ] yes = pass
 3. When a destination file already exists, does
    the job overwrite it (not rename, not skip
    silently, not fail)?                          [ ] yes = pass
 4. Is there a durable record (ledger, table,
    archive folder) the job checks before its
    append/insert/send steps?                     [ ] yes = pass
 5. If the job dies halfway, can the next run
    start from the top without manual cleanup?    [ ] yes = pass
 6. Are notifications and downstream triggers
    guarded so a rerun does not re-fire them?     [ ] yes = pass
 7. Can I tell from the logs alone whether a
    given file was processed once or twice?       [ ] yes = pass
 8. Is there a documented, deliberate way to
    reprocess a file on purpose?                  [ ] yes = pass

Score: 8 passes = rerun-safe. 6-7 = safe with care.
5 or fewer = do not rerun without reading the script first.

Questions 1 through 3 cover the transfer and file-handling half of the job, where fixes are usually cheap — a naming cleanup, an overwrite setting. Questions 4 through 6 cover processing and side effects, where the ledger pattern does the heavy lifting. Question 7 is about evidence: when someone asks "did this run twice?", you want the answer to come from logs. Our guide to what to log pairs well here. Question 8 is the maturity marker: rerun-safety is complete when reruns are not just survivable but routine.

Two notes on evidence. On the automation side, your jobs may run through a scheduler with built-in retry. Sysax FTP Automation, for example, retries failed transfers and emails you about errors. With such a scheduler, question 1 is not hypothetical: the tool itself reruns transfers on your behalf. That is exactly the behavior you want, and exactly why the rest of the job must tolerate it. On the server side, Sysax Multi Server logs every transfer to file and database. So "did the partner upload this file twice?" becomes a query rather than a guess — the raw material for question 7.

What Idempotency Does Not Promise

It is worth being precise about the limits, because idempotency gets oversold in the same breath as it gets explained.

It does not prevent failures. Jobs still die: networks drop, disks fill, credentials expire. Idempotency changes the cost of recovery, not the frequency of failure. You still need retry logic, error classification, and alerting — the whole apparatus of retry and error handling. Idempotency is what makes that apparatus safe to use aggressively.

It does not clean bad input. Suppose a partner sends a file with wrong data in it. An idempotent pipeline will faithfully produce the same wrong result every time you run it on that file. Deduplication and validation are different controls; you need both.

It does not guarantee the file arrives exactly once. No transfer mechanism can promise that a file crosses an unreliable network exactly once — duplicates on the wire are a fact of life. What idempotency buys you is the next best thing, and it turns out to be just as good. Files may arrive more than once, but they take effect exactly once. That trade — at-least-once delivery plus idempotent processing — is the central idea of exactly-once thinking, the theory piece of this series.

Gotcha: a job can be idempotent today and lose the property silently tomorrow. The classic regression is a well-meaning edit that adds an append step, a notification, or a collision-rename "improvement" without a guard. When you review changes to a transfer job, ask the rerun question about the diff, not just the job.

Where to Start This Week

You do not need a project plan to act on this article. A realistic sequence for one ordinary week:

  1. Pick your scariest job — the one you would least like to rerun. That fear is information; it means the job fails the self-test somewhere.
  2. Run the eight questions against it with the script open. Ten minutes of work converts dread into a list.
  3. Fix the cheap failures first. Stable names and overwrite-on-collision are usually one-line changes.
  4. Guard the one dangerous step. Most jobs have exactly one append/insert/send step that matters. Put the simplest durable check in front of it — even "skip if the file is already in the archive folder" is a rudimentary ledger.
  5. Rerun the job on purpose. In a quiet window, with a copy of yesterday's input, run it twice and diff the results. Nothing finds the step you missed like a controlled rerun.

Then do the next job next week. Rerun-safety across a whole environment is not a heroic rewrite. It is the same short checklist applied patiently, job by job. Eventually, the two-in-the-morning question — "is it safe to run this again?" — has the same answer everywhere: yes, always, that is how it was built.

The Version to Tell a Colleague

Idempotency means running a job twice leaves things exactly as if it ran once. File jobs get run twice constantly — retries, resends, catch-up schedules, nervous admins. So the property is not optional polish; it is what makes automation safe to operate. The danger concentrates in "change the state" steps: appends, inserts, notifications. Make files keep one stable name, and let overwrites be overwrites. Keep a durable record of what has been processed, and check it before the dangerous steps. Then a rerun is a shrug.

From here, a natural next read is where duplicate files actually come from. It catalogs the duplicate sources in detail with a defense for each. Another is detecting duplicates, which builds the ledger this article kept pointing at. The series ends by assembling all of it in one real pipeline, in a complete worked example.

Frequently Asked Questions

Is copying a file twice idempotent?
Usually, yes — if the copy goes to the same name and overwrites, the destination ends up identical after one copy or ten. It stops being idempotent the moment a collision rule renames the second copy (file_(2).csv) or the destination appends instead of replacing. Check the collision behavior, not just the copy command.
What is the difference between idempotent and read-only?
A read-only operation changes nothing at all. An idempotent operation may change plenty on the first run — it just produces the same final state no matter how many times it runs. Overwriting a file with the same content is idempotent but definitely not read-only.
Do I need a database to make a job idempotent?
No. Many jobs get there with naming and overwrite discipline alone, and a processed-files ledger can be a plain text file if it is updated carefully. A small database table adds convenient locking and querying, but it is an upgrade, not a prerequisite.
Does an idempotent job still need retry and error handling?
Yes — the two work together. Retries determine how often the job recovers on its own; idempotency determines whether those retries are safe. An idempotent job with no retries fails more than it should, and a retrying job with no idempotency duplicates data. You want both halves.
How do I test whether a job is really idempotent?
Run it twice on purpose in a controlled window, with the same input, and compare the results. Compare row counts, file listings, downstream records, and notifications sent. If the second run changed anything a stakeholder could notice, you have found your non-idempotent step. This double-run test belongs in your checklist whenever a job is modified.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.