Home › Topics › Files as Glue › When Glue Fails

Error Handling Across a File Interface

"We sent it at 02:00, same as every night." "Nothing was there at 02:15." When a file interface fails, the failure lands on a boundary owned by two teams. That is what makes it different from every failure inside your own systems. Nobody can see the whole picture alone. The producer's log says the export ran. The consumer's log says nothing usable arrived. In between sits a gap that, in badly run interfaces, gets bridged the expensive way. A meeting brings two teams together to reconstruct the night from mismatched logs and memory. They decide whose side failed by seniority and volume.

Well-run interfaces bridge it with convention instead. Every failure has a layer, and every layer has an owner. Rejects travel in an agreed shape. Reprocessing follows a playbook both teams know, and a shared evidence trail answers "whose side?" in one query. This article builds each of those pieces: a failure taxonomy, whole-file and record-level rejection done properly, and a reject-file convention you can copy. It builds the two-team reprocessing playbook and the evidence design that retires the meeting. This article is part of our Files as Integration Glue series. It assumes the interface has a written contract, because every decision below belongs in one. An interface without a contract handles errors too; it handles them differently every time.

Where File Interfaces Fail: Four Layers

Diagnosis starts with naming the layer, because each layer fails differently, is detected differently, and is owned by a different party. It is the same instinct as the layered troubleshooting method for transfers themselves, applied one level up.

  • Delivery failures. The file never arrives, arrives late, or arrives incomplete. Causes live in infrastructure: network blips, expired credentials, full disks, a scheduler that never fired. Owner: whoever moves the file.
  • File-level failures. The file arrived but is wrong as a unit. It may have unparseable structure, wrong encoding, or a name that matches no expected pattern. Its trailer's count or totals may not match the contents. It may have an unknown format version, or be a duplicate of something already processed. Owner: the producer's export, usually.
  • Record-level failures. The file is structurally fine, but some rows fail validation. Examples are a department code that does not exist, a required field empty, or an employee ID the consumer has never heard of. Owner: the source data, via the producer.
  • Business-level failures. Everything parses and totals, and the content is still wrong. Yesterday's snapshot may have been sent again under today's name, or amounts shifted a decimal place by an upstream change. An extract may have run before the source system finished loading and faithfully totaled the incomplete data. Owner: the producer's upstream world. These are the quietest and most damaging failures, because no structural check can catch them. Only plausibility checks (volume within expected range, totals within historical norms) and downstream reconciliation do.

Half the two-team friction in incidents comes from arguing across layers — one side talking about delivery timestamps while the other is talking about department codes. Put the taxonomy in the contract and make "which layer?" the first question of every incident. It is also a much better first question than "whose fault?"

Delivery Failures Belong to the Transfer Layer

Delivery is the one layer where errors can be handled almost entirely by machinery, and it should be. Transient failures — the connection reset, the momentary DNS wobble — deserve automatic retries with backoff, not a human alarm at first twitch. Permanent failures — authentication rejected (see troubleshooting authentication failures), path not found — deserve immediate escalation. Retrying them just delays the page. Learning to split the two is a discipline of its own, covered in transient vs permanent failures. This is also precisely the layer where automation tooling earns its keep. Sysax FTP Automation retries failed transfers and runs scheduled and folder-triggered jobs unattended. It sends email notifications when a transfer finally does need a human. That automation keeps the 2 a.m. blip from ever becoming a morning incident.

Two delivery-adjacent rules from elsewhere in this series do heavy lifting here. The contract's empty-batch rule (always send a file, even with zero records) makes absence unambiguous. So the consumer's freshness monitoring can alarm on a missing file with confidence. And completion signaling — upload under a temporary name, rename when done — prevents the half-delivered file from ever being read as a whole one. The failure modes it prevents are cataloged in why partial files happen. With those two in place, the delivery layer's error story is short: retry the transient, page the permanent, and never let anyone downstream see a partial. Half a file is not a file; it is a rumor.

Whole-File Rejection, Done Properly

Some failures condemn the entire file. A file may not parse, its trailer may not match, or its version may be unknown. Its business date may be wrong, or it may duplicate a file already applied. The consumer's response should be the same every time, and it has three moves.

Quarantine, untouched. Move the file aside into a quarantine area exactly as it arrived — byte for byte. The bad file is now evidence: the producer will need to see precisely what the consumer saw, encoding quirks and all. This is the same instinct as the dead-letter pattern in job design (see poison files and dead-letter folders). Isolate the poison where it cannot be retried into the import over and over, but never destroy it. The bad file is the only witness that was actually in the room.

Signal, with a reason. Return a negative acknowledgment — the STATUS=REJECTED ack from the acknowledgment patterns article — carrying an agreed reason code. That lets the producer's side route the failure without a phone call. "Trailer count mismatch" and "unknown version" start two very different investigations.

Never repair midstream. The tempting shortcut — the consumer's admin opens the file, fixes the obvious problem, feeds it back in — is how interfaces rot. The repaired file now exists nowhere in the producer's records. The two sides' histories have quietly diverged. The underlying cause survives to fire again next week, and this time the admin who knew the trick is on vacation. Fixes happen at the source, flow through the interface again, and leave a trail. If a genuine emergency ever forces a midstream repair, it happens with both teams' sign-off. In that case, it gets written into the incident record as the exception it is. I have done the midstream repair; it came back the following Tuesday.

Remember: the quarantined original is the single most valuable artifact in any interface dispute. Preserve it unedited, log its checksum, and keep it until the incident is closed and the fix is verified. A consumer that edits or deletes the evidence has forfeited the argument.

Record-Level Rejects and the Reject File

When the file is sound but individual records fail validation, the first question is not how to report them. It is whether partial acceptance is allowed at all. This is the atomicity decision, it belongs in the contract, and both answers are respectable. All-or-nothing (one bad record rejects the file) is right when the records form an accounting whole. Posting half a ledger batch leaves books that do not balance. Partial acceptance (apply the good, reject the bad) is right for roster updates, catalogs, and enrollment feeds. There, holding 8,000 good records hostage to three typos serves nobody. A useful middle setting is a reject threshold. Accept partial failure up to, say, two percent of records, but reject the whole file beyond it. That is because a high reject rate stops meaning "some bad data" and starts meaning "the producer's export is broken." You do not want to half-apply a broken export.

Bluewater Bank's reject threshold earned its clause one spring on the enrollment feed from Kestrel Payroll. The nightly file usually drew two or three rejects, all typos; one night it drew four hundred, every one of them E03 on the department code. The threshold tripped at two percent, and the whole file went to quarantine instead of being half-applied. A single nack with a single reason code went back before anyone had been woken. Kestrel's export had picked up a renamed department table that afternoon. That cause would have been far harder to see spread across four hundred partial updates already posted. The fix ran at the source. The replacement arrived under the next sequence number before the morning cutoff. The quarantined original was archived as the exhibit it was.

Rejected records travel back in a reject file, and its conventions matter enough to spell out. The name mirrors the source file with a reject marker. It flows in the return direction, in its own folder, like an ack. It uses the same format family as the data it echoes. Each reject carries three things: a stable error code from the contract's agreed list, a human-readable reason, and — the crucial part — the original record echoed verbatim. That way, the producer sees exactly the bytes the consumer judged, and can even build the correction mechanically. And the reject file practices what the interface preaches: header, trailer, and its own reject count. A worked example, echoing this series' running feed:

EMPL_ELIG_YYYYMMDD_01.rej

H,EMPL_ELIG_REJ,YYYYMMDD,01,SOURCE=EMPL_ELIG_YYYYMMDD_01.csv
R,E03,UNKNOWN_DEPT,"D,84317,""Okafor, Sam"",QQQ,150.00,YYYYMMDD"
R,E01,MISSING_REQUIRED_premium,"D,84320,Liu Yan,OPS,,YYYYMMDD"
T,2

Agreed codes (from the contract):  E01 missing required field
E02 bad format in field   E03 unknown reference value
E04 duplicate record      E05 out of allowed range

The echoed records are wrapped in quotes with their internal quotes doubled. The reject file obeys the same quoting rules as any other delimited file, a full circle back to flat-file formats. A reject file that mangles the records it reports is a practical joke, not an interface.

The Two-Team Reprocessing Playbook

Rejection is the easy half. The hard half is the correction cycle — two teams, two systems, and a batch that must end up applied exactly once. Agree on this sequence before you need it, and reprocessing becomes choreography instead of negotiation:

  1. Consumer detects and classifies. Name the layer, quarantine what needs quarantining, send the nack or reject file, log the event. Clock starts.
  2. Producer confirms receipt of the reject. A one-line reply or a collected-reject log entry — the tiny loop that prevents the reject itself from vanishing into silence.
  3. Fix at the source. Correct the data in the producing system (or its upstream), so both sides' histories stay consistent and the fix persists into every future run.
  4. Resend in the contract's correction shape. Snapshot feeds resend a full replacement under the next sequence number — same business date, NN incremented. The old file is superseded entirely. Delta feeds send a corrections file containing only the repaired records, explicitly marked as corrections. The contract says which; the producer never improvises the shape.
  5. Consumer applies the replacement safely. The highest-sequence-wins rule discards superseded files deterministically. The import must tolerate seeing records it already applied — the discipline of safe reprocessing patterns. This step is why idempotent imports are not a luxury: every correction cycle reprocesses something.
  6. Verify and close the loop. The replacement's ack — counts and totals matching expectations — closes the incident. No ack, no closure.
  7. Record the pattern. An incident log entry against the interface ID. Recurring rejects with the same code are not bad luck. They are a validation gap on the producer side (worth closing with validation before sending) or a contract clause that needs to evolve.

Two complications deserve their own agreements. When the day's delivery is a set of files and one fails, the contract must say whether the survivors proceed or the set holds together. The tradeoffs are worked through in partial failure in multi-file jobs. And reprocessing takes time the schedule may not have: a replacement that arrives at 09:00 has missed every downstream cutoff the 02:00 original would have made. The playbook therefore includes a communication step — who tells the downstream jobs and the business that tonight's numbers are late. That is really a question about your batch ecosystem's dependency chain, not about this interface alone. The downstream cutoff does not care whose fault the delay was.

The Shared Evidence Trail

Back to the dispute this article opened with. Producer: "We sent it at 02:00, same as every night." Consumer: "Nothing was there at 02:15." Without shared evidence this becomes archaeology — two logs, two clock settings, two formats, assembled in a meeting into a best guess. With designed evidence it is one query. Design in three parts:

Both endpoints log the same facts. For every file, each side records name, size, record count, and checksum. It also records a timestamp for each step it performed (exported, uploaded, downloaded, validated, applied). What makes logs comparable across teams is agreeing on those fields in advance — the general craft is in what to log.

The exchange point is the referee. The server in the middle sees both sides' behavior and belongs to neither team's narrative. Its activity log records every upload and download with account, action, and timestamp. On a Windows exchange point, Sysax Multi Server writes one to file and to a database. In the dispute above, one query settles it. If the log shows the producer's rename completing at 02:03 and no download after it, the consumer's collector never ran. If it shows no upload at all, the producer's job failed before the wire. Either way, the meeting just became a two-line chat message with a log excerpt attached.

The clocks agree and the artifacts survive. Synchronize every participating system to network time, and log in one declared time zone (the contract's). Retain the logs — and quarantined originals — for as long as a dispute could plausibly arise. Evidence with drifting clocks invites the "well, our log says" stalemate the whole design exists to prevent. Two clocks that disagree produce three opinions and no facts.

Telling the Other Team: Error Communication

Finally, the human channel. The contract's error section should define a small severity ladder. Its levels are informational (auto-recovered, no action), warning (late but arriving), error (rejection, same-day action), and critical (deadline blown, downstream impact). Each should have a named contact path and expected response. Severity inflation is real: if everything pages everyone, nothing does.

And standardize the failure notice itself. A useful one fits in six lines. These cover interface ID, file name, failure layer, error code, a pointer to the evidence (quarantine path, log excerpt, reject file), and the action requested with its deadline. Compare "the feed is broken again" with "IF-042, EMPL_ELIG_YYYYMMDD_01.csv, file-level, trailer count mismatch, quarantined at /quarantine/, full replacement requested by 06:00". The second is a work order; the first is an invitation to a meeting.

Boring Failures Are the Goal

Errors on a file interface are not a sign of bad engineering. Interfaces move real data from real systems, and some of it will always be late, malformed, or wrong. The sign of good engineering is that failure is uninteresting. The layer is named on sight, and the reject arrives in the agreed shape with the original record inside it. The replacement follows the sequence rule, and the import survives seeing records twice. The exchange server's log answers the only political question before it is asked. Every piece is small. Together they convert the worst meetings in integration work into log entries. A log entry rarely needs a conference room.

One loose end remains, and it is the biggest one: the errors you cause on purpose. Every format change, field addition, and renaming is a self-inflicted failure waiting to happen on the other side's parser. The closing article in this series is evolving a file interface without breaking the other side. It is the playbook for changing the contract without triggering everything this article just built.

Frequently Asked Questions

Why shouldn't we just fix a bad file by hand and reprocess it?
It is because the repaired file then exists only on your side. The producer's records no longer match what you processed. The root cause survives untouched, and the fix depends on whoever knew the trick. Quarantine, reject with a reason, and let the correction flow from the source through the normal interface. Reserve hand-repair for genuine emergencies with both teams' sign-off.
What is the difference between a nack and a reject file?
A nack is a whole-file verdict: a negative acknowledgment saying the file was rejected, with a reason code. A reject file is record-level detail: each failed record echoed verbatim with its error code and reason. Whole-file failures need only a nack; partial acceptance needs both — an ack reporting the counts and a reject file carrying the casualties.
Should corrections come as a full replacement or a corrections-only file?
Match the feed's model. Snapshot feeds resend a full replacement under the next sequence number, and the consumer applies the highest sequence only. Delta feeds send just the corrected records, marked as corrections. The wrong mix — a partial resend of a snapshot, say — is how consumers end up with half-updated state. So the contract fixes the shape in advance.
How do we keep reprocessing from posting records twice?
Two guards, worn together. One is a deterministic supersede rule — highest sequence number per business date wins, earlier files discarded. The other is an idempotent import that recognizes records it has already applied (by key and content) and skips or safely re-applies them. Correction cycles guarantee some records will be seen twice; design for it rather than hoping.
What evidence actually settles whose side failed?
There are three artifacts, side by side. First is the exchange server's activity log (neutral record of every upload and download, with accounts and timestamps). Second is each team's own job log for the steps it performed. Third is the quarantined original file when content is in question. With synchronized clocks, those three answer nearly every "whose side?" question in minutes.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.