The Transfer Inventory: Every Flow on One Page
"How many file transfers do we run?" "About thirty. Maybe fifty." "Which ones matter if the SFTP server dies tonight?" A longer pause. "I'd have to check." The check involves a scheduler on one box and a folder of scripts on another. It also involves a partner spreadsheet somebody built two jobs ago, and the memory of the administrator who set most of it up. The knowledge exists. It is just not anywhere you can point at, hand over, or trust under pressure, and the pause is the sound of that.
A transfer inventory fixes that. It is a single list with exactly one row per flow. A flow is one recurring movement of files from a source to a destination for a business purpose. That means "the nightly orders file to our freight partner," not "the SFTP server" and not "the script on the batch box." The inventory records what each flow is, where it goes, when it runs, and who owns it. It also records what it depends on and where the instructions live for when it breaks. Every other piece of documentation hangs off this page, which is why this page has to exist first.
This article shows you how to build one. It gives you the record schema column by column and a worked example from a fictional company running about forty flows. It includes a Markdown template you can copy and a short discovery script that seeds the list from what is actually scheduled. This is the foundation article in our Flow Documentation series. The articles on ownership, dependencies, runbooks, and currency all assume this list exists. They are optimists. For the next twenty minutes, so are you.
What a Flow Is, and Why One Row Each
Think in flows, not in servers and scripts. Servers and scripts are what you log into, so they are what a new administrator naturally counts; the inventory asks for a different unit. A server hosts many flows. One flow may be spread across several scripts. These could be an export on the application server, a push job on the transfer box, and a confirmation pull the next morning. The business does not care about those pieces individually; it cares that the orders reached the freight partner before dispatch started. That outcome is the flow, and it is the unit you document.
Take our worked example. Meridian Parts is a fictional distributor with about forty flows. Some are outbound to partners: a nightly orders export to Acme Freight, a daily payment file to Bluewater Bank, a weekly payroll extract to Kestrel Payroll. Some are inbound: a price list pulled from Northgate Retail every morning, invoices dropped by suppliers into a hosted SFTP folder. Many are internal: warehouse scanner uploads, a report bundle for finance, a log archive shipped to a retention server. Forty flows across five servers, three scheduling tools, and two administrators — one of whom built most of it.
Three terms you will meet throughout this series, defined once here. The owner of a flow is the person accountable for it. They can say whether it still matters, approve a change, and answer for it if it is wrong. The custodian (or technical owner) operates it day to day. They are often different people — the logistics manager owns the orders flow; the transfer administrator is its custodian. And a source of truth is the one copy of a fact that wins when copies disagree. The inventory is the source of truth for "this flow exists and here is what it is". A wiki page, a scheduler comment, or a partner's onboarding form should point to the row, not contradict it. The wiki page will try.
One row per flow forces a decision for every ambiguous case. Is the acknowledgement pull part of the orders flow or its own? Its own, if it can fail independently and someone would need to act. Is one report sent to three offices one flow or three? One if they succeed or fail together; three if each has a different owner or schedule. Making those calls once, on paper, saves making them at three in the morning.
The Lists You Probably Already Have
Almost nobody starts from nothing, and almost nobody starts from one thing either. Most estates already hold several partial lists, each built for a different purpose and each quietly certain it is the real one. The inventory's first job is to merge them:
- A flow census. If you ever walked the business asking "what files move where," you produced one — see taking a file flow census. A census is a snapshot; the inventory is that snapshot kept alive.
- A job list. An export of the scheduled tasks on the batch box. It lists jobs, not flows, but every job belongs to a flow.
- A partner register. Trading partners with contacts and connection details, per partner exchange basics. Each partner maps to one or more flows.
- A sprawl register. The endpoints a discovery sweep found, per discovering every server and script. The inventory groups them into flows.
- A migration workload inventory. A project list with cutover sequencing and a move-or-retire verdict — see the workload inventory for migration. That snapshot has an end date; yours is the operational record that outlives it. When the migration closes, fold its list back into yours.
The Flow Record Schema
A schema is the agreed set of columns and what each one means. The one below has twenty, which sounds like a lot until you try to run a flow from a row that lacks one. Each exists because someone was paged at night without that fact. None of them were added in daylight. The examples are from Meridian's FLOW-0007, the nightly orders export.
| Column | What goes in it | Example (FLOW-0007) |
|---|---|---|
flow_id | Permanent identifier, zero-padded so it sorts. Never changes, never reused. | FLOW-0007 |
flow_name | Short human name: what moves, to or from whom. | Nightly orders export to Acme Freight |
purpose | One sentence on why the business needs it. | Confirmed orders reach the carrier for next-day dispatch |
direction | outbound (we send), inbound (we receive), or internal. | outbound |
source | Host and path where the file starts. | erp01.corp.example.com, D:\Exports\Orders\ |
destination | Host, port, and path where it ends. | sftp.acmefreight.example.com:22, /inbound/meridian/ |
protocol | SFTP, FTPS, FTP, HTTPS, SMB — plus the authentication method. | SFTP, SSH key |
credential_ref | Where the secret is kept. Never the secret itself. | vault: acme-freight-sftp-key; account meridian_orders |
schedule | When it runs, in words, plus its deadline. Watch-folder flows say "on arrival." | 02:00 daily; must complete by 04:00 |
file_pattern | Naming pattern, expected count, typical size. | ORDERS_YYYYMMDD.csv, one file, two to forty megabytes |
job_ref | Host and the exact scheduler entry, script, or automation task that executes it. | mft01: Task Scheduler \Meridian\FLOW-0007-orders-acme |
business_owner | The accountable role and name. | Logistics manager (P. Okafor) |
technical_owner | The custodian who operates it. | Transfer admin (D. Reyes) |
backup_owner | Who steps in when the custodian is away. | S. Lindqvist |
external_contact | Partner or vendor contact, as a reference into the contact sheet. | contacts: ACME-EDI |
dependencies | What must happen first (upstream) and what waits on this (downstream). | up: ERP day-end close; down: FLOW-0008 ack pull, Acme import 04:30 |
criticality | A tier, and what a miss costs in plain words. | Tier 1 — a miss delays customer deliveries by a day |
runbook | Link to the per-flow runbook. | runbooks/FLOW-0007.md |
status | planned, active, suspended, or retired. Rows are never deleted. | active |
last_reviewed | Date and initials of whoever last confirmed the row is true. | YYYY-MM-DD, DR |
Three columns deserve a second look. credential_ref is the one people get wrong first. The inventory is widely shared, so it must never contain a password or key. The credential reference holds only a name that a person with the right access can look up. job_ref is the bridge between the flow (a business idea) and the mechanism (a task on a specific host). It is what lets a stranger find the thing to restart. And status with last_reviewed turns a list into a living record. Without them you cannot tell a row checked last month from one that was true three administrators ago. Keep the column names lowercase with underscores exactly as shown. The article keeping documentation current compares this list to the live scheduler with a script. Scripts prefer stable headers, because unlike people they refuse to guess.
A Second Row: The Inbound Price List
Outbound pushes like FLOW-0007 are the easy case; inbound pulls have their own wrinkles. FLOW-0023 is "Daily price list pull from Northgate Retail": source sftp.northgate.example.com:22, path /outbound/pricing/; destination mft01.corp.example.com, D:\Inbound\Northgate\; SFTP with a password under vault: northgate-sftp. Because Meridian cannot control when Northgate publishes, the schedule column reads "poll at 06:00, 06:30, and 07:00; file expected by 06:00; escalate if absent at 07:00." The file pattern is PRICES_YYYYMMDD.csv, exactly one file. A row like that is what lets a monitor raise an alarm for a missing file, the subject of freshness checks for expected files. Upstream is Northgate's publishing job, which Meridian cannot see; downstream is the ERP price import at 07:15. Criticality is Tier 2: a miss means the sales team quotes yesterday's prices for a day.
Notice what the row does not contain: no password, no explanation of SFTP, no incident history. The row is deliberately thin so that forty of them fit on one screen. Fat rows grow a column called notes, then one called ??, then dust.
The Markdown Flow Record
A spreadsheet row is the right shape for scanning forty flows at once and the wrong shape for reading one carefully or for version control. So most teams keep two views: the inventory table for the overview, and one flow record per flow in Markdown. Markdown is a plain-text format that renders in a wiki or a code repository and can be diffed line by line. The record holds the same twenty fields plus notes. Copy this template, one per flow:
# FLOW-0007 — Nightly orders export to Acme Freight - **Status:** active - **Purpose:** confirmed orders reach the carrier for next-day dispatch - **Direction:** outbound - **Source:** erp01.corp.example.com D:\Exports\Orders\ - **Destination:** sftp.acmefreight.example.com:22 /inbound/meridian/ - **Protocol:** SFTP, SSH key authentication - **Credential ref:** vault: acme-freight-sftp-key (account meridian_orders) - **Schedule:** 02:00 daily incl. weekends; must complete by 04:00 - **File pattern:** ORDERS_YYYYMMDD.csv, one file, 2-40 MB - **Job ref:** mft01: Task Scheduler \Meridian\FLOW-0007-orders-acme - **Business owner:** Logistics manager — P. Okafor - **Technical owner:** Transfer admin — D. Reyes - **Backup owner:** S. Lindqvist - **External contact:** contacts: ACME-EDI - **Upstream:** ERP day-end close (normally done by 01:30) - **Downstream:** FLOW-0008 acknowledgement pull; Acme import at 04:30 - **Criticality:** Tier 1 — a miss delays customer deliveries by a day - **Runbook:** runbooks/FLOW-0007.md - **Last reviewed:** YYYY-MM-DD by DR ## Notes - Weekend files are small; a zero-byte file on Sunday is NOT normal.
The table and the records must agree, and the way to guarantee that is to generate one from the other with a small script. Editing both by hand and hoping is not a method.
Seeding the List From What Is Actually Running
The fastest way to a first draft is not to interview people; it is to ask the machines. Every unattended flow is triggered by something — a scheduled task, a cron entry, a systemd timer, an automation tool's task list, or a watch folder. Every inbound flow lands on an account. Listing those gives you raw jobs and accounts to group into flows. On the Windows side, run this as an administrator on each transfer or batch host:
# seed-inventory.ps1
$stamp = Get-Date -Format 'yyyyMMdd'
# 1. Scheduled tasks that are not Windows' own, with what they execute
Get-ScheduledTask | Where-Object { $_.TaskPath -notlike '\Microsoft\*' } |
Select-Object TaskName, TaskPath, State,
@{n='Action'; e={ ($_.Actions | ForEach-Object { "$($_.Execute) $($_.Arguments)" }) -join ' ; ' }} |
Export-Csv -NoTypeInformation "seed-tasks-$env:COMPUTERNAME-$stamp.csv"
# 2. Local accounts (a hosted SFTP/FTP server often has one per partner)
Get-LocalUser | Select-Object Name, Enabled, LastLogon, Description |
Export-Csv -NoTypeInformation "seed-accounts-$env:COMPUTERNAME-$stamp.csv"
The Linux equivalent, run as root, is shorter:
#!/bin/sh
# seed-inventory.sh — every cron entry, systemd timers, transfer-group members
{ cat /etc/crontab /etc/cron.d/* 2>/dev/null
for u in $(cut -d: -f1 /etc/passwd); do
crontab -l -u "$u" 2>/dev/null | grep -v '^#' | sed "s/^/$u: /"
done
} > "seed-cron-$(hostname -s).txt"
systemctl list-timers --all --no-pager > "seed-timers-$(hostname -s).txt" 2>/dev/null
getent group sftpusers > "seed-accounts-$(hostname -s).txt"
Collect the outputs from every host that might run or receive a transfer, and start grouping. A task named orders-acme and a script that mentions sftp.acmefreight.example.com are the same flow. An account named northgate on the hosted server and a cron line pulling from Northgate are two halves of one relationship. The grouping will raise questions the machines cannot answer. That is when you ask the humans, with a list in hand rather than a blank page. I have never had a useful answer to "what runs on this box?"; I have had dozens to "is this task still yours?"
This is a seeding pass, not a discovery sweep. A sweep also scans the network for listeners you did not know about and mines DNS and firewall rules. The funnel is in discovering every server and script, with the plain-FTP variant in finding all your FTP. Seed first; sweep when you suspect flows on machines you have never logged into.
If your automation tool keeps its own task list, export that too. A scheduled-transfer product such as Sysax FTP Automation holds each scheduled task and script in one place. That makes a good habit easy: name each task with its flow ID and record the task name in job_ref. That way, the inventory row and the task that executes it point at each other.
Where the Inventory Lives
The format matters less than three properties: everyone who might need it can read it, few can change it, and every change is recorded. A shared spreadsheet meets the first two and, with discipline, the third. A CSV in version control beside the Markdown records meets all three naturally. A wiki table works if the wiki keeps page history. A CMDB — a configuration management database, the asset register large IT departments keep — is the most formal option. It is right when the rest of the organization already lives there and wrong when it does not. Nobody corrects a record they never visit.
Whatever you choose, apply these rules from day one:
- IDs are permanent. FLOW-0007 stays FLOW-0007 when renamed, moved, or retired. Never reuse a number.
- Rows are never deleted. A retired flow gets
status: retiredand a note saying when and why. Whoever investigates a mystery file six months later will thank you. - One column, one fact. Do not pack host, path, and port into a paragraph.
- The inventory has an owner. Somebody is custodian of the document itself — responsible for its accuracy, not for every flow in it. Without that, it decays.
- Secrets never enter it. References only, always.
Remember: the inventory is a list of flows, not of servers or scripts. If a row cannot answer "what business outcome does this produce, and who cares if it stops," it is not a flow. It is a mechanism that belongs in some flow's job_ref column.
The on-call administrator needs it from a laptop at home during an outage, so it cannot live only on the server that is down. For the same reason, the DR runbook has the same requirement. Documentation has a gift for living on exactly the server that is down.
What the Inventory Unlocks
Once the list exists, many other jobs turn out to be filters over it.
- Disaster recovery scope. "Which flows must come back first?" is the Tier 1 rows sorted by deadline — the scope list the disaster recovery series starts from.
- Audit evidence. Auditors ask for a list of data movements, who owns each, and how you know it is complete. The inventory with its
last_reviewedcolumn answers all three. That is why it appears among the evidence auditors accept. - Policy compliance. A policy that says "every automated transfer must be documented and owned" needs a place where that lives; see why a transfer policy.
- Monitoring. Every row with a file pattern and a deadline is a monitor waiting to be configured.
- Ownership, runbooks, dependencies. The owner columns point into the contact sheet from flow ownership and contacts. The runbook column points at runbooks per flow. The dependencies column compresses the map from mapping flow dependencies.
The inventory also shapes server configuration. A hosted server with per-account settings — Sysax Multi Server is one — lets you give each inbound flow its own account. So the account name on the server matches the credential_ref in the row. The server's activity log becomes evidence of the flow running.
Building It Without Stalling
The most common way an inventory fails is by never being finished. The team aims for forty perfect rows, gets to twelve, and stops when something urgent comes up. I have been on that team, and the twelve rows were beautiful. Aim instead for forty rough rows in the first week. Fill flow_id, flow_name, job_ref, and technical_owner for everything the seed pass found; mark the rest TBC. Forty flows with visible gaps beat twelve complete rows, because gaps can be assigned. Then walk the list with each owner, ten minutes per flow, aiming for no TBC in a Tier 1 row within a month. Finally, publish it: an inventory people use is an inventory people correct.
Bluewater Bank's first seed pass turned up a scheduled task called push_old on a batch host. It sent one file every night to an SFTP folder nobody on the team could name. The task got a row with the owner marked unknown, status set to suspended, and a switch-off date announced to the whole department. Eleven days later a treasury analyst asked, politely, why the morning reconciliation had stopped balancing. The flow had an owner, a proper name, and a runbook by that afternoon. The task had outlived the three people who knew what it was for, and the row was the first thing in years that had asked.
Wrapping Up
The transfer inventory is one row per flow, twenty columns, no secrets, permanent IDs, never deleted, always reviewed. It merges the census, the job list, the partner register, and any sprawl or migration lists into one operational record that outlives them all. Seed it from the schedulers, fill it with the owners, keep it reachable during an outage, and give it a custodian. The next time someone asks how many transfers you run, there is no pause.
Next, flow ownership and contacts fills the three owner columns, and mapping flow dependencies expands the dependencies column into a map. When the list is complete, keeping transfer documentation current shows how to stop it drifting.
Frequently Asked Questions
What counts as one flow?
Should the inventory be a spreadsheet, a wiki page, or a CMDB?
Can I put the account password in the credential column?
How is this different from the migration workload inventory?
What do I do with a flow nobody claims?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
