Home › Topics › Build vs Buy › Hidden Costs

The Hidden Costs of Homegrown Transfer Infrastructure

"It'll take five minutes." It took forty. The partner's folder had moved again. The engineer who fixed it went back to the project she was meant to be on, having logged nothing. There was nowhere to log it. Multiply by a quarter. The previous article in this series argued the build case at full strength, and nothing here retracts a word of it. Homegrown transfer scripts really do deliver exact fit, total transparency, and freedom from vendors — and a disciplined team really can hold that position for years. This article is the same ledger, other column: what the position costs to hold, counted fairly, starting with those forty minutes.

"Counted" is the operative word. The genre norm for this topic is scare copy — adjectives about fragility, anecdotes about departed graybeards, and no numbers. It is a genre vendors, ourselves included, are structurally tempted by. You will get none of that here. Instead, we explain why these particular costs evade measurement. We offer a tracking method that produces real figures in one quarter and two specific tests you can run this month. And — because fairness cuts both ways — we include an honest section on the hidden costs of the bought alternative. This is the third article in our Build vs Buy series, and its output is the set of numbers the tipping-point worksheet later consumes.

Why These Costs Hide

Homegrown transfer costs are not hidden because anyone hides them. They hide because of how they are paid.

They are paid in a drip. No single event is big enough to book. There are fifteen minutes rerunning a failed job, half an hour adjusting a script to a partner's new folder. There are ten minutes explaining to a colleague why the Thursday file is special. Each drop rounds to zero. Only the sum is large, and nobody is holding a cup under the drip.

They are paid by the wrong person, invisibly. Script care gravitates to the most capable engineer — the author — whose time is scheduled for other work. The care happens as interruptions, and interruptions are doubly uncounted. The minutes themselves go unlogged, and the broken concentration around them costs more than the minutes. A "quick fix" that takes twenty minutes can take an afternoon's focus with it.

They are paid under other names. Failure triage gets logged, if at all, as "helping accounting." Audit answering is "a meeting." Explaining the estate to a new hire is "onboarding." The transfer infrastructure has no cost center, so its costs are booked to everything else. A bought tool's license quote arrives as one visible number with a signature line. That is precisely why the quote always feels more expensive than an estate that quietly consumes an engineer-week a quarter. Organizations scrutinize line items. The drip has no line item.

And they are protected by survivorship logic: "the scripts work, so they must not be costing much." But "working" describes the files arriving, not the labor making them arrive. The estate that works flawlessly may be consuming heroic effort; the one consuming almost nothing may be one silent failure from disaster. Uptime tells you neither. Only counting does.

The Fair-Accounting Method

The method is one sentence: for one quarter, log every hour anyone spends on the transfer estate, tagged by category and by person. The rigor is in the rules that keep it honest — in both directions.

Count only observed hours, not recalled ones. Memory inflates drama (the weekend outage) and forgets routine (the daily two-minute check that is really twelve minutes). Log at the end of each day, in the tracker below, for a quarter that contains at least one normal month. Do not reconstruct the past; it will lie in whichever direction you already lean. Memory is a loyal witness and a poor one.

Do not count what a bought tool would also cost. Deciding what a flow should do, negotiating formats with partners, cleaning bad source data, owning the flow in audits — this work moves with you. It does not vanish at purchase. Counting it against the scripts rigs the ledger. The categories below are chosen to be build-specific.

Do not count learning at full cost. Hours that taught a junior engineer scripting, protocols, or failure design bought something real beyond the fix, as the build case rightly insists. Log them, but tag them, so the ledger can show both readings.

Do count interruption weight. Note whether each entry was scheduled work or an interruption. An estate that consumes six hours a quarter as planned maintenance is far cheaper than one consuming six hours as pager events, even though the totals match. Six planned hours is a task; six pages is a lifestyle.

HOMEGROWN TRANSFER ESTATE - QUARTERLY HOUR LEDGER
(one line per event: date / person / category / hours /
 scheduled-or-interruption / one-line note)

 A  FAILURE TRIAGE + RERUN     diagnosing, fixing, rerunning,
                               cleaning up downstream effects
 B  CHANGE ABSORPTION          partner or application changes:
                               paths, names, formats, windows
 C  CREDENTIAL + CERT WORK     rotations, expiries, key moves,
                               and failures caused by them
 D  PLATFORM DRIFT             OS/runtime/library updates that
                               broke or required script edits
 E  NEW FLOW BUILD             hours to stand up each new flow,
                               start to trusted
 F  EVIDENCE + AUDIT           answering "prove this transfer"
                               and assembling audit material
 G  KNOWLEDGE TRANSFER         explaining, documenting after the
                               fact, onboarding a new caretaker
 T  TAGGED LEARNING            any hours above that also built
                               durable skill (tag, don't remove)

End of quarter, total by category and by person. Two follow-up
numbers: what fraction was interruption, and what fraction was
paid by one person. Keep the ledger - the worksheet article
turns it into a decision input.

One quarter of this produces the number this whole debate usually lacks. Teams that run the exercise get surprised in both directions. Some discover an engineer-week a quarter nobody had seen. Others discover a genuinely cheap estate and gain the standing to say so with evidence. Both are wins. A low number is a finding, not a failure of the exercise. If the number needs translating into budget language, costing failed jobs and manual work does the conversion.

Meridian Parts ran the ledger for one quarter, mostly to end an argument. The estate was eleven flows, one author, and a shared belief that it cost "an hour or two a week." The ledger came back at a little over an engineer-week for the quarter, which was close to the belief. What nobody had predicted was the shape. Three-quarters of the hours were interruptions, and nine-tenths of those had landed on the author, logged in the timesheet as "helping accounting." The estate was not expensive. It was expensive in one place, at the worst possible moments, and the ledger was the first document to say so.

Reading the Categories

Category A is usually the biggest line, and detection latency is most of it. The expensive part of a failure is rarely the fix; it is the time between failure and discovery. That is because scripted jobs fail silently by default. Every silent day adds downstream cleanup: missed reconciliations, partner escalations, files to re-request. A big category A is also the most fixable line on the ledger. External monitoring and alerting shrink it dramatically without buying anything, which is why counting comes before concluding. (It is no accident, though, that bought automation bundles exactly these defenses. Tools like our Sysax FTP Automation attach retry handling and failure email to every task as configuration, precisely because category A tops most ledgers.)

Category B scales with partners, not with quality. Each change is genuinely cheap — the build case's editor-speed advantage is real. But the count of changes is set by how many counterparties you have and how often they move. Five partners who change something twice a year is a rounding error. Forty partners is a part-time job, arriving as interruptions. Partners move folders the way weather moves: without consulting you.

Categories C and D are the aging costs. An untouched estate still decays, because the world under it moves. Credentials expire, ciphers get deprecated, an OS upgrade changes a default, and a library drops a behavior your loop relied on. These hours surprise teams most, because they arrive without anyone having changed anything — the scripts rot at the speed of their environment. Credential work in particular hides failures of hygiene. If category C is large, read our service account hygiene guide before drawing build-vs-buy conclusions. A bought tool with the same sloppy credential habits pays the same bill.

Category E reveals your estate's architecture. If each new flow costs a day because it reuses a shared library, your estate scales the way the build case promises. If each new flow costs a week because it is a copy-paste fork of the last one, your marginal cost is high and rising. In that case, multiplied by growth, category E becomes the ledger's future. Write down the marginal cost of flow number next; the worksheet will want it.

Category G is the liability account. Hours explaining the estate are the visible interest on an invisible principal: the knowledge concentrated in one head. Which deserves its own test.

The Bus-Factor Test

The bus factor of a system is the number of people whose sudden absence would cripple it (the hit-by-a-bus test runs the same question over your documentation). For homegrown transfer estates the honest number is very often one. The cost of that fact appears nowhere on any quarter's ledger — it is a contingent liability, booked only when it detonates. You can price it in advance with a test that takes one afternoon plus one calendar decision.

THE BUS-FACTOR TEST

 1. List every production flow (if this list does not exist,
    that is finding zero - start the inventory).
 2. For each flow, write the name of a person OTHER than the
    author who has actually diagnosed and fixed a failure of
    that flow, alone. "Could probably figure it out" does not
    count. Blank is blank.
 3. Count the blanks. That is how much of the estate has a
    bus factor of one.
 4. The live test: the author takes two consecutive weeks off,
    genuinely unreachable, announced in advance. Log every
    transfer issue that waits for their return, every workaround
    invented, every task quietly postponed "until they're back."
 5. Price the liability: multiply what two weeks cost in
    delays and risk. A departure is not two weeks - it is
    forever, minus whatever documentation exists.

Two fairness notes. First, a bus factor of one is not negligence; it is the natural resting state of anything one able person built. It becomes negligence only when known and left alone. Second, the test's cure does not require buying anything. Runbooks, rehearsed handoffs, and the conditions checklist from the build-case article restore a passing grade within a quarter or two of honest effort. What the test tells you is whether that effort will actually happen. If the answer has been "next quarter" for six consecutive quarters, the liability is permanent. Permanent liabilities belong on the decision's scale. Anyone who has ever inherited a departed author's estate — the experience documented in inheriting undocumented FTP automation — can testify to what the detonated version costs: months of archaeology, conducted during incidents.

Remember: the ledger prices the estate you run; the bus-factor test prices the estate you would suddenly own on the author's last day. Decisions that look only at the first number are pricing the ship and ignoring the iceberg contract.

The Audit-Labor Line

Category F deserves its own section because it is binary across organizations. It is near zero for teams nobody examines, and a dominating cost for teams under recurring scrutiny. Most teams migrate from the first group to the second as they grow, usually without noticing the crossing.

Run the drill once and you will know which group you are in. Pick a real transfer from last month and produce evidence as if an auditor asked. Show the file was sent, received intact, on time, by an authorized identity. Show evidence you would have known if it had failed. Time the exercise. I have timed it at four hours for a single delivery, most of them spent finding which machine's log held the answer. In a scripted estate the answer is assembled by hand from per-script logs in per-script formats across several machines. The one person who knows where everything is assembles it; an hour to half a day per question is typical. The guide to what auditors actually accept raises the bar further, because prose reconstructions convince nobody. Multiply by questions per year, note who pays the hours (almost always the author, interrupted), and enter it on the ledger.

Uniform, queryable transfer records are much of what the buy side actually sells, beneath the feature lists. A server such as Sysax Multi Server writes activity logs to file and to a database so that one query spans every account and flow. A script drawer requires one investigation per question. If category F is your dominant line, that is worth knowing precisely. It is equally worth knowing that a small evidence discipline — consistent log formats and retention, as our logging guide describes — can shrink the line substantially without a purchase. Counting first keeps both options honest.

The Bought Side's Hidden Costs — Counted the Same Way

A ledger that only audits one side is advocacy wearing a green eyeshade. Bought transfer infrastructure has its own drip of quiet costs. And — full disclosure — we sell on that side of the street. So weigh this section as the interested testimony it is, and verify it through the vendor assessment series, which exists to arm you against vendors including us.

License growth. Products are priced on axes — users, connections, flows, servers, features — and your usage moves along those axes as you grow. The step function you accepted at purchase can step again at renewal, and the tiers you outgrow are rarely the ones you predicted. Ask any vendor how spend evolves as each axis doubles, in writing, before signing. Watch how long the answer takes to arrive.

Upgrade projects. Self-hosted products need updates applied, tested, and scheduled — real recurring hours, non-negotiable for an internet-facing transfer tool, as the industry's breach history brutally established. The trade-offs run exactly parallel to the ones in self-hosted versus SaaS: someone always patches; the question is who, and on whose schedule.

Lock-in, compounding quietly. Every flow configured in a product's proprietary format, every partner pointed at its endpoints, every retention year of logs in its store deepens the cost of ever leaving. This is the bought side's bus-factor equivalent: a contingent liability, booked only at exit time, and steadily growing while unexamined. The final article in this series covers the contract with yourself that keeps it bounded.

Vendor-watching labor. A vendor is a dependency with a trajectory — advisories to track, an acquisition to react to, a support quality that can sag. Managing that is the recurring work described in ongoing vendor monitoring, and it belongs on the bought side's ledger at a few hours a quarter, forever. Builders pay zero here; buyers should never pretend otherwise.

The glue and the owner. The tool automates flows; it does not own itself. Someone administers it, and the odd corners of your estate still get scripts around the edges. Buying shrinks the homegrown ledger; it does not zero it. Every estate keeps a junk drawer; buying just makes it smaller.

What the Numbers Mean

After a quarter you hold something rare in this debate: evidence. Read it without prejudice in either direction. Evidence is rude to everyone equally.

If the total is small, the interruption share is low, the bus-factor test produced names instead of blanks, and category F barely exists — your hidden costs are not hidden. In that case, they are merely modest. Keep building, on the merits, and re-run the ledger when growth or turnover changes the facts. This outcome is common in small stable estates and deserves to be called what it is: the build case, empirically confirmed.

If the total is large, resist the reflex to conclude "buy" in the same breath. First separate the practice-shaped costs from the structure-shaped ones. Categories A, C, and G respond to disciplines you can adopt in a quarter — monitoring, credential hygiene, runbooks — without spending anything but the hours. An estate cured by practices may re-run the ledger into the healthy column. What does not respond to practice is structure. That includes category B scaling with partner count, category E scaling with growth, category F scaling with scrutiny. It includes a bus factor that six quarters of good intentions have not moved. Structural costs rise with the curve you are on, not with the diligence you apply. When they dominate the ledger, you have located your position on the cost curves from the opening article. At that point, it is time to score the decision properly. Take the ledger, the blanks, and the drill timings to the tipping-point worksheet; they are its first three inputs.

Frequently Asked Questions

Isn't "hidden costs" just vendor fear-mongering?
The genre often is, which is why this article refuses adjectives and asks you to log hours for a quarter instead. Measured costs cannot be fear-mongered — they are whatever they are, and a low total is explicitly called out here as a legitimate verdict for keeping your scripts. If a vendor's cost argument cannot survive being replaced by your own measurements, it deserved to die.
How long do we really need to track hours?
One quarter that includes at least one ordinary month. Shorter samples get distorted by whether a bad week happened to land inside them; recalled estimates are worse, inflating drama and forgetting routine. If a full quarter is politically impossible, six honest weeks beats any amount of remembering.
Our tracked hours came out low. Did we do it wrong?
Possibly right. Small, stable, well-tended estates genuinely are cheap — that is the build case working. Before celebrating, check the two blind spots: did everyone log (including the author's interruptions), and did the quarter include any partner changes or failures at all? If yes and yes, your number is real. Keep building, and keep the ledger habit at a lighter cadence.
Do bought tools eliminate these cost categories?
No — they reshape them. Triage shrinks where retry and alerting are built in; change absorption gets cheaper per change; evidence becomes a query instead of an investigation. But license spend grows with usage axes, upgrades become scheduled projects, lock-in compounds, and the vendor itself needs watching. The honest comparison is ledger versus ledger, not ledger versus brochure.
Which category is usually the biggest surprise?
Knowledge concentration — not as a line of hours, but as the liability the bus-factor test exposes. Teams expect triage to dominate and often it does. Yet the finding that changes decisions is usually the column of blanks where second names should be. That column prices a risk the quarterly ledger structurally cannot show.
How do we put a number on the author leaving?
Run the two-week unreachable test and extrapolate. Everything that waited, was worked around, or was postponed becomes permanent on departure day, minus whatever documentation exists by then. Teams that inherit undocumented estates typically spend months reconstructing them, mostly during incidents. That figure — months of senior time, spent at the worst moments — is the honest size of the liability.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.