Runbooks Per Flow: What to Do When It Breaks
Ten past two. The alert for FLOW-0007 has fired, and the on-call administrator has the inventory open in one window and the dependency map in another. The inventory says what the flow is. The map says what it needs. Neither says what to do. They do not say which log to open, which command re-runs the job, or whether re-running is safe. They do not say who to call if it is still broken at three. That knowledge is on a sticky note on the custodian's monitor, in an office forty minutes away, and the sticky note is not on call. What should be there instead is a runbook: a short, per-flow set of instructions written for someone who has never touched the flow. It is to be read under pressure, in the dark, on a laptop at the kitchen table.
Most transfer estates have runbooks for the platform ("how to restart the SFTP service") and none for the flows. But flows are what fail. The service is up; the orders file still did not arrive, because a partner rotated a host key or the ERP export finished late. A flow runbook covers that, if it was written by the person who knows the flow for the person who does not.
This article gives you the runbook template and the test that tells you whether a runbook is good enough. It includes a worked example for Meridian Parts' nightly orders export and a method for building the common-failures section. It also gives you the rules that keep runbooks short enough to be read at three in the morning. This article is part of our Flow Documentation series and fills the runbook column of the transfer inventory.
What a Runbook Is, and What It Is Not
A flow runbook answers one question: this flow has failed, what now? It is specific to one flow, organized by symptom rather than by topic, and short — one page, two at most. It assumes the reader can log in and run commands but knows nothing about this flow, this partner, or this file.
Three neighbors are easy to confuse with it. The troubleshooting method is the generic, layer-by-layer way of finding out why any transfer failed — connectivity, then authentication, then permissions, then protocol, then content. It is the same for every flow, so it is written once, in the troubleshooting series. Every runbook points to it for the cases the runbook did not anticipate. The disaster recovery runbook is for the day the whole estate is gone and must be rebuilt somewhere else. It is one document for the platform, covered in the disaster recovery series and written up as the DR runbook. And the flow record is the reference sheet — hosts, paths, owners, schedule — which the runbook links to rather than repeating.
The runbook is the thin layer between all of those and the operator's hands. It covers how you know it is this flow and what to check before touching anything. It covers the five ways it usually breaks and the fix for each. It also covers how to re-run it safely, and who to call and when.
The Three-in-the-Morning Test
A runbook is good enough when a competent colleague who has never seen the flow can follow it to a fix or to a correct escalation. They must be able to do this at three in the morning, with no one to ask. Everything about how runbooks are written follows from that test.
- Readable in two minutes. Headings, short lines, no paragraphs of history. The reader is scanning, not studying.
- Symptom first. The reader has an alert text and a log line. The runbook's failure section is organized by what they see, not by underlying cause, because they do not know the cause yet.
- Every command copy-pasteable. Exact task names, exact paths, exact hosts. "Re-run the job" is not an instruction;
schtasks /Run /TN "\Meridian\FLOW-0007-orders-acme"is. - Every decision has a default. "Consider whether to resend" leaves the reader guessing. "Resend if the partner confirms nothing was received; otherwise do not, and escalate" does not.
- Nothing assumed. Where the log is, how to read the success line, what the file normally looks like. A reader who has to guess will guess wrong at that hour.
- A stopping rule. The runbook says when to stop fixing and start escalating — in minutes, relative to the deadline. That way, nobody spends two hours on a Tier 1 flow that needed a phone call at twenty past two.
The way to know whether a runbook passes is to make someone try it, which is the drill described in the hit-by-a-bus test. Until then, write as if that drill is tomorrow. The unscheduled version might be.
The Runbook Template
Nine sections, in the order the reader needs them. The template is Markdown so it can live beside the flow record and be diffed. Copy it for each flow and delete nothing. An empty section that says "none known" is information.
# Runbook: FLOW-XXXX — <flow name> Tier: <1/2/3> Deadline: <time> Custodian: <contact_id> Backup: <contact_id> Flow record: <link> Last tested by a non-owner: YYYY-MM-DD (<initials>) ## 1. What this flow does (two sentences) <What moves, from where, to whom, why it matters.> ## 2. What normal looks like - Runs at <time> on <host> as <job_ref> - Produces/collects <file pattern>, <typical size>, <count> - Success line in <log path> looks like: <exact example> ## 3. How you will know it failed - Alert name(s): <...> - Symptoms without an alert: <file absent at X, partner call, etc.> ## 4. Before you touch anything (two minutes) - [ ] Is the upstream dependency done? <what to check> - [ ] Is a shared component down? <host, key, firewall — check for sibling alerts> - [ ] Has it already recovered on retry? <where to see retries> - [ ] What has already been sent? <how to see partial success> ## 5. Common failures and fixes | Symptom (what you see) | Likely cause | Fix | Safe to re-run? | |---|---|---|---| | <log text> | <cause> | <exact steps> | yes / no / ask <who> | ## 6. How to re-run safely - Command: <exact command> - Before: <what must be true — e.g. remove partial file, check nothing sent> - After: <what to check — log line, file at destination, size> ## 7. Escalation - Stop fixing and escalate if: <not fixed within N minutes AND deadline within M> - Order: <contact_id> (hours) → <contact_id> (hours) → business owner (wake for Tier 1) ## 8. After the fix - Tell downstream: <flow IDs and owners> - Log the incident: <where> - If the failure was new, add a row to section 5 ## 9. Do not - <the tempting wrong moves for this flow>
Section 4 is the one people leave out and the one that saves the most time. Half of all night-time flow failures are not the flow's fault: the upstream export was late, or the shared host is down and six sibling alerts prove it. Two minutes of checking prevents an hour of fixing the wrong thing. I have spent that hour, more than once, with great confidence. The dependency map from mapping flow dependencies is where the items for this section come from.
Section 9 sounds odd until you have written one. Every flow has a tempting wrong move. Examples are re-running a payment file that was half-received, or deleting a "stuck" file that was still being written. Another is accepting a changed host key without checking. Naming them in the runbook is how the last person's scar becomes the next person's warning.
A Worked Example: FLOW-0007
Here are the sections that carry the most weight, filled in for Meridian's nightly orders export. Section 2 first, because "what normal looks like" is the reference the operator compares everything against:
## 2. What normal looks like
- Runs 02:00 daily on mft01 as Task Scheduler \Meridian\FLOW-0007-orders-acme
- Sends D:\Exports\Orders\ORDERS_YYYYMMDD.csv (2-40 MB weekdays, under 1 MB weekends)
to sftp.acmefreight.example.com:22 /inbound/meridian/ as meridian_orders
- Success line in D:\Jobs\logs\FLOW-0007.log:
Mar 14 02:03:41 FLOW-0007 OK ORDERS_YYYYMMDD.csv 18442310 bytes uploaded, renamed from .tmp
- Retries: the job retries the connection three times, five minutes apart, before alerting.
If you see the alert, those retries have ALREADY happened.
## 6. How to re-run safely
- Command (on mft01, elevated): schtasks /Run /TN "\Meridian\FLOW-0007-orders-acme"
- Before: confirm the source file is dated today (dir D:\Exports\Orders\). If it is
yesterday's, STOP — the ERP close has not finished; see section 4.
- The job uploads to a .tmp name and renames on completion, so a re-run after a failed
upload is safe: Acme never sees a partial file. Re-running after a SUCCESSFUL run
sends a duplicate — check the log first.
- After: watch for the OK line; then confirm the size matches the source file.
And the common-failures section, as a table. Every row began life as a real incident:
| Symptom (in the log) | Likely cause | Fix | Safe to re-run? |
|---|---|---|---|
no files matched ORDERS_*.csv | ERP day-end close ran late or failed (upstream) | Check the ERP batch status page. If close is still running, wait; if it failed, escalate to MER-ERP. Do not send yesterday's file. | Yes, once today's file exists |
host key verification failed | Acme rotated their SSH host key (they do this on a fixed cycle) — or, rarely, something is wrong on the path | Confirm the new fingerprint with ACME-EDI before accepting it. Never accept blind. Update the known-hosts entry, then re-run. | Yes, after the key is confirmed |
connection timed out after three retries | Acme endpoint down, or our egress firewall rule changed | Test port 22 from mft01. If it fails from mft01 but not from the jump host, it is our firewall — escalate to MER-NET. If it fails everywhere, call ACME-EDI. | Yes, once reachable |
permission denied on /inbound/meridian/ | Acme changed folder rights, or the account was disabled | Nothing to fix on our side. Call ACME-EDI with the exact log line and the account name. | Yes, after Acme confirms |
OK line present but 0 bytes on a weekday | ERP wrote an empty file (its export failed silently) | Treat as a failure. Escalate to MER-ERP; when a real file exists, re-run — Acme's import ignores the empty one. | Yes, with the real file |
Notice the shape of each row. The symptom is literal log text, because that is what the operator will search for. The fix names contacts by their contact-sheet ID, so the reader goes to one place for the phone number. And the last column answers the question that matters most and is hardest to answer at night: whether running it again could make things worse. For this flow the answer is almost always yes, because the job uploads to a temporary name and renames on completion. For the bank payments flow the same column reads "ask MER-FINANCE before any resend," because a duplicate file means duplicate payments. The reasoning behind those two answers is in safe reprocessing patterns.
Building the Common-Failures Section
The template gives you the columns; the rows come from three sources. The first is history. Go through the last year of tickets and chat threads about this flow. For every time it broke, write down what the operator saw and what fixed it. Most flows have three to six recurring failures, and a runbook that covers those covers most nights. The chat thread was the runbook all along; this just moves it somewhere findable.
The second source is the dependency table. Each dependency produces a row: "upstream not done" for the data dependency, "shared host down" for the component, "partner cutoff passed" for the timing. These rows often have no history yet, and that is exactly why they belong in the runbook. The first time they happen should not be the first time anyone thinks about them.
The third source is the layered troubleshooting method. For each layer — connectivity, authentication, permissions, protocol, content — ask "what would that look like for this flow, and what is the fix here?" That gives you generic rows in flow-specific words: not "check connectivity" but "test port 22 to sftp.acmefreight.example.com from mft01." The table's last row then says "anything else — follow the layered method," with a link. The runbook handles the known; the method handles the unknown.
The "safe to re-run" column deserves its own pass. Distinguish transient failures — a timeout, a busy server — which a retry fixes, from permanent ones — wrong password, missing folder — which a retry only repeats. The distinction is worked through in transient versus permanent failures. And record what the job already did on its own. If the automation tool retried three times before alerting, as the retry handling in Sysax FTP Automation can be configured to do, the runbook must say so. Otherwise, the operator will spend twenty minutes repeating retries that already failed.
Escalation Paths
Escalation is not a sign of failure; it is a step in the runbook, with a trigger and an order. The trigger is a stopping rule expressed in time: "if not fixed within thirty minutes, or if the deadline is within ninety minutes and the cause is not yet known, escalate." Fix the numbers per flow from the deadline and the slack in the dependency map. A Tier 1 flow with a four o'clock deadline and a two o'clock failure has two hours. A rule that escalates at half past two leaves ninety minutes for the escalated person to act, which is about right. I have never heard anyone regret escalating at half past two.
The order is a list of contact-sheet IDs, each with the hours they can be reached, taken from the sheet built in flow ownership and contacts. For FLOW-0007 it starts with ACME-EDI (staffed from six in the morning — so before six, skip to the next step). Next is MER-ONCALL escalation (always), then the business owner MER-LOGISTICS, who has agreed to be woken for Tier 1. That last clause is a decision the owner made in daylight so that nobody has to make it at night. Partners' hours and promised response times come from the partner agreement; the runbook only repeats the ones the operator needs.
Escalation also means telling people who are not fixing anything. Section 8's "tell downstream" line exists for the owners of FLOW-0008 and FLOW-0009. They would rather hear "orders were sent at 03:40, your flows will run late" than discover it from their own alerts. Nobody enjoys learning about their own flow from a monitor.
Remember: the runbook's stopping rule is the most valuable sentence in it. An operator who knows when to stop fixing and pick up the phone will beat a more skilled one who does not, every single night.
Keeping Runbooks Short Enough to Read
Runbooks grow. Every incident adds a row, every clever administrator adds background, and within a year the one-page runbook is six pages nobody reads under pressure. Five rules hold the line.
- One page for the path; links for the depth. Background on how SFTP host keys work belongs in the library, linked from the row — see host keys and known hosts — not in the runbook.
- No history. "This used to be FTPS until Acme changed" is a change-log line in the flow record, not a runbook sentence.
- Retire rows. A failure that has not happened in two years and whose cause has been removed comes out of the table. Keep it in the record's change history if you must.
- Same structure everywhere. Nine sections, same order, same headings, in every runbook. An operator who has used one has used them all, and can find section 6 without reading sections 1 to 5.
- Test it, then trust the test. The "last tested by a non-owner" date in the header is the runbook's expiry date. A runbook not tested in a year is a rumor.
Acme's runbook for its payroll extract had been written carefully and then left alone. When the flow failed for real one night, the operator followed the runbook faithfully to a host that had been decommissioned. Then the operator followed the runbook to the host's replacement, which had also been replaced. The runbook was three servers out of date, and every step in it had been true once. The flow was fixed from memory by the custodian, at home, on the phone, which is the exact situation the runbook existed to prevent. The fix was a quarterly review in which the backup walks the runbook against the live estate and every wrong hostname becomes a ticket. The next failure was fixed by the runbook alone, and the custodian slept through it.
Where runbooks live follows the inventory's rule: somewhere reachable when the transfer estate is not, with page history. For Tier 1 flows, keep a printed copy in the on-call folder. Version control beside the flow records is the natural home. It makes "add a row after each incident" a small, reviewable change rather than a wiki edit nobody sees.
For inbound flows, one line in section 2 is worth adding for every hosted server. It says where the server's own activity log is and what a normal partner login looks like in it. A server that logs each session to a file or a database, as Sysax Multi Server does, lets the operator answer "did the partner even connect?" in one query. This is the first question for any inbound failure and the one the partner will ask you. Reading those logs fluently is covered in reading transfer logs.
Wrapping Up
A runbook per flow is a one-page, symptom-first, copy-pasteable set of instructions with a stopping rule. It is written by the person who knows the flow for the person who does not. There are nine sections, the same in every runbook. There is a failure table built from history, dependencies, and the layered method. There is a re-run section that says whether running again is safe, and an escalation order in contact-sheet IDs with hours. Keep it short by linking out, retiring rows, and testing it with a stranger. The sticky note can retire.
The stranger test is passing the hit-by-a-bus test. The reasons flows fail without anyone noticing — the failures your section 3 must anticipate — are in why jobs fail silently. The routine for recovering a whole night's chain after a fix is in catch-up after failures. When an incident reveals a row the runbook lacked, the discipline for turning it into a fix is running your own postmortem.
Frequently Asked Questions
Do I really need a separate runbook for every flow?
How is a flow runbook different from the troubleshooting method?
What if I do not know whether re-running is safe?
How long should a runbook be?
Who writes the runbook, and who keeps it current?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
