Home › Topics › Files as Glue › Why Files Persist

Why Files Are Still the Universal Integration Interface

"It's a CSV on an FTP server. We should really replace that with an API." The sentence gets said in a design review about once a quarter. It is usually about a feed nobody in the room built and everybody in the building depends on. Somewhere in your estate, probably within one hop of wherever you are sitting, two applications are integrated by a file. A transfer inventory would turn up a dozen more. Payroll exports a fixed-width feed that the benefits provider imports. The warehouse system writes an inventory extract at 2 a.m. and the storefront loads it before dawn. A supplier drops a pipe-delimited price list on your SFTP server every Monday morning. Nobody calls this architecture, but it runs the business. Sooner or later someone looks at one of these flows, says the word "legacy," and proposes a rewrite.

Sometimes they are right. Real-time questions deserve real-time answers, and a file interface cannot give them. Much of the time, though, the file is doing exactly what files are best at, and replacing it would trade real strengths for fashionable ones. I have sat in that review on both sides of the table. The argument is easier to win, in either direction, when it is about shape rather than taste. This article makes the honest case both ways. It covers the four durable strengths of file-based integration and the places where an API genuinely serves better. It provides a decision table for settling the choice on engineering grounds. This article opens our Files as Integration Glue series. The series treats the file interface as a discipline worth doing well rather than a habit to apologize for.

What a File Interface Actually Is

A file interface is three things. A producer — the system that writes the file. A consumer — the system that reads it. And the file itself, which is the message: a batch of records, complete in itself, handed across the boundary between two systems that otherwise share nothing. No shared database, no shared code, no live connection. There is just an agreement: I will put a file shaped like this, in a place like this, on a schedule like this. You will take it from there.

The pattern is older than almost everything else in the machine room, and it is everywhere. Banks settle with each other through batch files. Entire industries run on EDI — electronic data interchange, standardized business documents exchanged largely as files. Our plain-language EDI guide covers that world. Data warehouses are fed by extract files. ERP-style systems — the big business suites that run finance, inventory, and orders — often speak to the outside world primarily through import and export. When people say "integration," they tend to picture something modern and glowing. In practice, an enormous share of it is files, moved on schedules, parsed by jobs. The glowing part is mostly the slide deck.

That longevity is not mere inertia. Files hold four genuine engineering advantages, and being able to name them precisely is what separates defending an architecture from defending a habit.

Strength One: Producer and Consumer Never Have to Meet

An API call requires three things to be true at the same instant. The caller is up, the receiver is up, and the network between them is healthy. A file interface requires none of that simultaneity. The producer writes when it can. The file waits. The consumer reads when it is ready. Either system can reboot, patch, crash, or crawl without the other noticing. The file sits patiently in between, like a parcel in a locker rather than a phone call that must be answered on the third ring.

This is time decoupling, and it is worth more than it sounds. Systems owned by different teams — or different companies — do not share maintenance windows, release calendars, or opinions about downtime. A file interface lets each side run on its own rhythm and meet only at the handoff point. If the consumer is mid-upgrade at the moment the file lands, nothing fails; the import runs when the system comes back. An API integration in the same situation is a pile of failed calls and a retry storm. Two teams page each other about an outage that a file would have absorbed silently. The file, meanwhile, has not noticed there was an outage.

Northgate Retail learned this the slow way. Their nightly inventory feed from the warehouse system to the storefront was rebuilt as a stream of API calls, one per stock line. That was because the file had been labeled legacy. The first month went fine. Then the warehouse system took its regular patch window on a Sunday night. The storefront's calls failed for forty minutes. The retry queue backed up into Monday's trading, and two teams spent the morning paging each other. The old file had crossed that same window every week for years without either side noticing. They kept the API for live stock lookups and put the nightly load back on a file. That was the shape the nightly load had always wanted to be.

The decoupling is organizational as much as technical. The producer team needs no credentials into the consumer's system, no knowledge of its internals, and no seat in its incident calls. Both sides need to understand exactly one thing: the file. That narrowness is a feature. It is the entire surface area of the relationship. That is why this series spends a whole article on writing that surface down as a contract. Unwritten, it is a contract both sides remember differently.

Strength Two: Batch Economics

Every API call pays a toll: connection handling, authentication, request framing, response parsing, logging. Pay the toll once for one record and it is trivial. Pay it two million times for two million records and the toll dwarfs the cargo. Rate limits make it worse — a polite client inserts pauses so the receiving service survives. Pagination turns one logical extract into tens of thousands of round trips. Each one is a fresh opportunity to fail halfway through and leave you wondering which page you reached. Somewhere past page nine thousand, the retry logic develops opinions.

A file amortizes the toll. Two million rows become one artifact. It is written sequentially and compressed to a fraction of its raw size (repetitive tabular data compresses beautifully). It is moved over a single authenticated connection and bulk-loaded on the far side using the import machinery every database is optimized for. The arithmetic is not subtle. For large volumes on a periodic rhythm, batch delivery is routinely orders of magnitude cheaper than the same volume delivered as individual calls. That applies to wall-clock time, compute, and failure opportunities.

There is a reason the heaviest flows in most organizations — ledger feeds, statement runs, catalog syncs, warehouse loads — are files. Where volume is big and timing is periodic, batch economics win, and it is not close. APIs can move bulk data too, and sometimes should; our companion piece on file transfer through REST-style APIs looks at when that shape fits.

Strength Three: The File Is Its Own Audit Trail

When a file interface misbehaves, you hold something no API integration gives you by default: the exact artifact that crossed the boundary. Archive each day's file and you can answer, months later, precisely what was sent. You have the bytes themselves — not what the code should have sent, not what a log line summarizes. You can diff Tuesday's file against Monday's to see what changed. You can replay a file into a test system to reproduce a bug. You can hand an auditor the actual feed behind the ledger entry they are questioning. A log summarizes; the file testifies.

API integrations can reach the same standard, but only through deliberate effort. Capture every request and response payload on both sides, retain it all, and hope the two logs agree when it matters. Few teams do this thoroughly, because it is expensive and nobody demands it until the first dispute. With files, the evidence is the medium. Add a checksum manifest and you can also prove a copy is faithful. Our guides to checksum files and manifests and verifying transfers end to end show how.

The transfer layer adds evidence of its own. When the handoff runs through a server both sides trust, its activity log records who connected, what was uploaded, what was downloaded, and exactly when. A Windows server such as Sysax Multi Server writes that activity log to file and to a database. That turns "did you even send it?" from an argument into a query. Later in this series, the article on error handling across a file interface leans hard on that shared evidence trail.

Strength Four: Everything Speaks File

Every platform you will ever meet can read and write files. The mainframe can. The twenty-year-old ERP-style system can. The cloud service whose only integration option is "export as CSV" can. The lab instrument, the payroll bureau, the government portal, the smallest trading partner with one part-time IT person — all can produce or consume a file. Nearly all of them can manage an SFTP or FTPS connection to move it. Import and export are the lowest common denominator of software. For a meaningful share of the systems in any estate, they are the only integration surface offered at all. For a surprising number of products, "export as CSV" is the whole integration strategy.

This is why files dominate at organizational boundaries. Inside one engineering team you can mandate a common platform. Across two companies you cannot; you negotiate down to what both sides can actually do, and both sides can always do files. A file interface asks nothing of the counterpart except the ability to write records and transfer them. It needs no SDK, no client library, no shared vendor, no framework alignment. When the other end is a partner you cannot re-equip, universality is not a nostalgia argument. It is the argument. Our B2B partner exchange series covers running dozens of such relationships as a managed program.

Where an API Genuinely Serves Better

An honest case has to spend real effort on the other side, because the other side wins often. When it wins, it wins for reasons no file design can answer.

  • Latency. A file interface answers on the next run; an API answers now. If the question is "can this order ship today?" or "is this card stolen?", the answer cannot wait for tonight's batch. Anything a human or a live process is actively waiting on belongs to a request-response interface. No batch schedule short of absurdity changes that.
  • Granularity. Files carry batches; APIs carry single records naturally. When changes are small, frequent, and independent — one order placed, one address updated — wrapping each event in a file of its own is ceremony. In that case, accumulating them into batches reintroduces the very delay you were trying to avoid.
  • Interactivity. An API is a conversation: ask a question with parameters, receive an answer scoped to exactly that question. A file is a broadcast — the consumer takes the whole thing and finds its own answers. Queries, searches, and lookups have no good file equivalent.
  • Typed contracts with immediate validation. A well-built API checks each request as it arrives and rejects bad data to the caller's face, synchronously. That happens while the caller still has the context to fix it. In the file world, validation happens later, and the sender learns about problems after the fact — often the next morning. That gap is why validating files before sending is a discipline of its own. Rejection at the door is a genuinely better failure mode than rejection by voicemail.
  • Per-record feedback. An API returns a status for every item as it is submitted. A file interface reports per-record outcomes only if you build that reporting — acknowledgment files, reject files — yourself.

These are five faces of one theme: APIs shine when interaction is fine-grained, immediate, and two-way. Files shine when the work is bulk, periodic, and one handoff at a time. Most integration questions become easy the moment you say which of those two shapes the flow actually has.

Files or API: The Decision Table

When the debate arrives at your desk, walk it through these questions. No single row decides; the pattern across rows does.

Question to ask Files tend to win when… An API tends to win when…
How fresh must the data be? The next scheduled run is soon enough — hourly, nightly, weekly Consumers need answers in seconds, while someone waits
What shape is the volume? Large batches on a rhythm — thousands to millions of records at once Small independent updates trickling in all day
Must both systems be up together? No — either side may be down, slow, or in maintenance at any time Yes, and both sides are engineered and staffed for that coupling
Who needs feedback, and when? A team, by next morning, via an acknowledgment or report The submitter, immediately, inside the same interaction
How will you prove later what was sent? Archive the files themselves; replay and diff at will You are truly prepared to capture and retain every call payload, on both sides
What can the other side support? Anything that can import and export — including systems nobody can modify Both ends can build and maintain clients against a typed interface
What happens on partial failure? You want a batch you can reject, correct, and resend as one unit You want per-record status handled by the caller in real time

Remember: the choice is made per flow, not per company. "We are an API shop" and "we have always done files" are both slogans. The table above is an engineering decision you can defend in a design review — in either direction.

Most Estates Rightly Run Both

The mature answer to "files or APIs?" is usually "yes." The two patterns are not rivals so much as different tools that pair well, and the strongest architectures use each where it is strong:

  • API in front, file behind. Orders arrive interactively through an API all day; a nightly settlement file reconciles the day's totals between the same two systems. The API handles the moment. The file establishes the durable, auditable record — and catches any drift between the two sides before it compounds.
  • File to seed, API to maintain. The initial load of a million customer records goes as a bulk file, because a million API calls is nobody's idea of a launch plan. Ongoing single-record changes then flow through the API.
  • APIs inside, files at the boundary. Within your own walls, where you control both ends, service calls make sense. At the company boundary — partners, bureaus, banks — the file interface's universality and evidence trail earn their keep.

Whole ecosystems are built on that mixture. The overnight processing world — cutoffs, dependency chains, the choreography of feeds in and feeds out — is its own subject, covered in our nightly batch ecosystems series. The failure state worth avoiding is not "too many files" or "too many APIs." It is dogma: rebuilding a perfectly good nightly feed as twenty thousand paginated calls, or forcing a real-time question to wait for a batch window. That happens because one pattern was declared the modern one. A nightly feed rebuilt as twenty thousand calls is still a nightly feed.

The Bill Files Make You Pay

None of the four strengths comes free. Choosing files means signing up to solve, yourself, several problems that request-response interfaces partially solve for you. Every one of them is tractable — this series and its neighbors exist to solve them. But pretending they do not exist is how file interfaces earn their bad reputation. I have watched every one of them happen to a feed described as "simple."

  • Delivery is silent by default. Nothing in the pattern tells the producer the consumer got the file, read it, or applied it. You add that yourself with acknowledgment patterns — and you define what silence means.
  • A file can be seen half-written. A consumer that reads while the producer is still writing processes a partial batch. Completion signaling — temporary names, atomic renames, marker files — is mandatory hygiene; see temp names and atomic renames.
  • Duplicates will happen. Retries, reruns, and resends eventually deliver the same file twice. Consumers must be safe to feed twice; why duplicates happen explains the mechanics.
  • A missing file makes no noise. An API call that fails throws an error somewhere. A file that never arrives just… doesn't. Monitoring must expect files and alarm on absence — the pattern in freshness checks for expected files.
  • The format is a handshake, not a schema. Nothing stops the producer changing a column and finding out a week later what broke. Format discipline and change management get their own articles later in this series.

There is also the delivery layer itself: the file has to actually move, on schedule, unattended, with retries and alerts when the world misbehaves. That layer is exactly what transfer automation tooling exists for. Sysax FTP Automation runs transfers on schedules and watches folders so a job fires the moment the other side's file lands. It retries transient failures and sends email notifications when a run needs a human. Think of it as the plumbing under every pattern this series describes; the articles concentrate on what flows through the pipes.

Defending the Choice, Either Way

The next time someone calls your file interface legacy, you have a four-part answer. The file interface decouples two systems in time so neither waits on the other's uptime. It moves bulk volume at a cost per record no call-based interface approaches. It leaves an audit trail made of the actual data. And it works with the counterpart you actually have, not the one you wish you had. And the next time you inherit a file interface that is straining, you can argue for the API with the same vocabulary and the same confidence. Users may be refreshing a screen for data that only updates nightly, or single records may be wrapped in ceremonial batches. You will be arguing about shape, not fashion. The quarterly design review will still happen; it will just be shorter.

What files demand in exchange is engineering discipline at the interface. That is the rest of this series. Start with designing a file interface contract, the written agreement both sides sign. Then come the format decisions that keep the interface healthy for years. The pattern is old. Done well, it is not tired.

Frequently Asked Questions

Is file-based integration considered legacy?
The pattern is old, but old is not the same as obsolete. Files remain the best fit for bulk, periodic, loosely coupled exchange — which describes a huge share of business data flows. "Legacy" properly describes a flow whose shape no longer matches its tool. That can be true of an API integration just as easily as a file one.
Can a file interface ever be close to real time?
It can get to minutes. Event-driven designs — the producer drops a file whenever data is ready, and the consumer's folder monitoring reacts immediately — remove the waiting-for-the-schedule delay. Below the level of a minute or so, or for answers someone is actively waiting on, a request-response API is the right tool.
What is the biggest operational risk of a file interface?
Silence. A file that never arrives, or arrives half-written, fails without any error being thrown anywhere. The antidotes are completion signaling (temporary names and atomic renames), acknowledgments from the consumer, and monitoring that alarms when an expected file is absent. Monitoring visible transfer failures alone is not enough.
Why do banks still exchange files instead of using APIs everywhere?
Because settlement is inherently batch: defined cutoffs, large volumes, and a legal need to prove exactly what was exchanged. Files batch economically, archive perfectly, and work identically for every institution in the chain regardless of its technology. Banks also run APIs — for the interactive parts. The two coexist deliberately.
Can an API and a file feed carry the same data without conflict?
Yes, and it is a common, healthy design — for example, an API for live updates plus a nightly file that reconciles totals. The key is declaring one side authoritative when they disagree, so a discrepancy triggers a correction in a known direction instead of a debate.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.