Home › Topics › Flow Documentation › Bus Test

Passing the Hit-by-a-Bus Test

"What happens if Dana is on a plane with no internet?" The question comes up in every transfer team, usually as a joke, and usually about one person. Dana set up most of the flows and knows which partner's server needs the old cipher. She remembers that the payroll file must never be resent without a phone call. She keeps the one key three partners trust on her laptop. Dana is an asset, and none of this is her fault. The estate depends on her being present, awake, and employed. Nobody can promise all three, least of all Dana, who would like a holiday.

A single point of knowledge is a person without whom a flow cannot be run, fixed, or safely changed. The number of people who would have to be unavailable before that happens is sometimes called the bus factor. The bus is traditional, but a month's leave, a lottery win, or a better job does the same thing more politely. A flow with a bus factor of one is one long-haul flight away from being unrunnable. The hit-by-a-bus test is the drill that measures it. A colleague who has never touched the flow sits down with the documentation and nothing else. They try to run it, break it, and fix it. Either they can, or you now know exactly what is missing.

This article is the last in our Flow Documentation series and is where the earlier articles are proven. You will get the drill step by step, a scorecard, and a method for hunting the knowledge that never made it onto paper. You will also get a shadowing plan and a handover checklist. The example is Meridian Parts, our fictional distributor with about forty flows and, until recently, one Dana.

What the Test Measures

"Could someone else run this?" hides three different questions, and a flow can pass one while failing the others. Can they find it? Given only the flow's name, can a stranger locate the inventory row, the runbook, the host, the job, and the credentials — through references, not by asking? Can they run it? Can they read what normal looks like, trigger a run or a re-run, and confirm success from the log? Can they fix it? Given a realistic failure, can they diagnose it from the runbook's failure table, apply the fix, and escalate correctly when the fix is beyond them?

The three questions are also what a new administrator needs on day one. A day-one pack for a transfer role is short. It covers where the inventory is and how to read a row, and where the runbooks are and how they are structured. It says where the contact sheet and on-call rota are, and how credentials are referenced and how to get vault access. It also covers which hosts run flows and how to log in, and the review calendar they will join. A new starter with those six things can be useful in their first week. One who has to ask for each is learning the estate from Dana, which recreates the problem.

Notice what the test does not measure: whether the documentation is complete, elegant, or long. A two-page runbook that gets a stranger to the fix beats a twenty-page one they cannot navigate at speed. A twenty-page runbook is thorough the way a haystack is.

The Tabletop Exercise

A tabletop exercise is a rehearsal done at a desk rather than during a real incident: a scenario, a person responding to it, and a facilitator observing. For a flow, it takes about an hour and follows a fixed shape.

  1. Pick a flow and a stranger. Start with a Tier 1 flow. The responder must never have operated it — the backup owner is ideal, because they will actually need to. The custodian facilitates and does not touch the keyboard.
  2. Give them only what they would have. The documentation, their normal access, and a laptop. No verbal briefing. If the responder asks a question, the only permitted answer is "the documentation does not say that," and the facilitator writes the question down. That list is the exercise's most valuable output.
  3. Scenario one: run it. "Show me the flow ran successfully last night, and how you would run it again if asked." The responder must find the row, the job, the log, and the success line, and produce the exact re-run command. They must not execute it against a partner unless the runbook says that is safe today.
  4. Scenario two: it broke. The facilitator hands over a failure card. This is a real log excerpt from a past incident, such as host key verification failed. Or it is a message like "Acme called to say nothing arrived." The responder works the runbook to a fix or to the point where they would escalate. They say out loud who they would call and when.
  5. Scenario three: a decision. "The job shows the file was sent, but the partner says they have half of it. What do you do?" This tests whether the runbook's "safe to re-run" answer and the escalation order are clear enough that the responder does not guess.
  6. Stop at the time limit. Forty-five minutes for the three scenarios. Running out of time is a finding, not a failure of the responder.

Do the parts that change anything in a safe place. A staging copy of the flow, built as the testing and staging series describes, lets the responder actually execute the re-run and see the log line. Where no staging exists, the responder performs every read-only step for real and describes the write steps precisely. "I would run schtasks /Run /TN "\Meridian\FLOW-0007-orders-acme"" is enough to show they found it. Never let a drill resend a payments file to a real bank to prove a point.

If the flows run inside an automation tool rather than the operating system's scheduler, the drill includes opening the tool and finding the task by its flow ID. A team using Sysax FTP Automation, for example, should expect the responder to locate the scheduled task named for the flow and read its last result there. If the naming convention holds, that takes seconds; if not, it is a finding.

Scoring the Drill

Scoring turns "that went okay, I think" into a number that can be compared between flows and over time. Score each item zero (could not), one (did it with difficulty or a logged question), or two (did it cleanly from the documentation).

Item What "two" looks like Where a gap gets fixed
Found the inventory row and runbookFrom the flow name alone, in under five minutesWhere the inventory lives; the day-one pack
Identified the host and jobLogged in to the right host and opened the right task or cron entryThe job_ref column
Read what normal looks likeFound last night's success line and the file at the destinationRunbook section 2
Obtained credentials via the referenceFollowed credential_ref to the vault with their own accessVault access for the backup role
Produced the re-run command and its safety conditionsExact command, plus "check X before, Y after"Runbook section 6
Diagnosed the failure cardMatched the symptom to a runbook row and applied its fixRunbook section 5
Checked upstream and shared components firstAsked "is this the flow, or something it depends on?" before fixingRunbook section 4; dependency table
Escalated correctlyNamed the right contact, in order, with hours, and the stopping ruleRunbook section 7; contact sheet
Made the decision scenario safelyDid not guess; found the answer or the person who owns itRunbook sections 6 and 9; RACI
Finished within forty-five minutesAll three scenarios, no facilitator answers neededRunbook length; structure

Twenty points are available. Seventeen or more is a pass: the flow survives Dana's absence. Eleven to sixteen means the documentation is close and the question log says what to fix before re-testing. Ten or less means the flow has a bus factor of one, whatever the inventory says about its backup owner. Record the score and date in the runbook header — the "last tested by a non-owner" line. Also record them in a simple tracker of flow, date, responder, and score. That lets you watch the estate's scores rise over a year. The first number is usually humbling; that is the number doing its job.

Meridian's first drill was on FLOW-0012, the bank payments file, with the transfer administrator as responder and the finance analyst facilitating. Score: nine. He found the row and runbook quickly, but the credential reference pointed at a vault entry he had no rights to open. The "safe to re-run" cell said "ask." The failure card — an expired certificate on the bank's FTPS endpoint — matched no row. Nobody had done anything wrong; the documentation had never been asked a question by a stranger before. Three findings, three tickets, and a re-test a month later scored seventeen.

The Tribal-Knowledge Hunt

Tribal knowledge is what the team knows and the documentation does not. Examples are the partner who needs a phone call after every key change, or the export that finishes ten minutes late on month-end. Another is the folder that must not be emptied because a second job reads it. The drill's question log is the first place it surfaces — "which log is current?", "why wait until 02:30?", "who is R. in the notes?" Every logged question is tribal knowledge caught in the act of not being written down.

The second technique is the interview, done in daylight with the custodian. Walk through the flow record line by line. After every fact ask "how do you know that?" and "what would go wrong if it were different?" Listen for the phrases that mark undocumented knowledge — "everyone knows," "just," "usually," "it should be fine," and any sentence ending in "ask Dana." Each one becomes a note in the record, a row in the runbook table, or a line in the do-not section. The word "just" is usually carrying an entire runbook.

The third technique is a sweep for knowledge that has hardened into personal artifacts. These are things that only work because a particular person's account, laptop, or home folder exists. These break not when the person is absent but when their account is disabled on their last day. Look for scheduled tasks running as a person rather than a service account, and scripts under a user profile. Also look for keys in a personal folder and saved passwords in one person's client. On a Windows host, the first two are quick questions:

# Tasks that run as a named person instead of a service account (adjust the pattern to your naming)
Get-ScheduledTask | Where-Object { $_.TaskPath -notlike '\Microsoft\*' -and $_.Principal.UserId -notlike 'svc_*' } |
  Select-Object TaskName, TaskPath, @{n='RunsAs'; e={ $_.Principal.UserId }}

# Transfer scripts living in someone's profile rather than the jobs folder
Get-ChildItem -Path C:\Users -Recurse -Include *.ps1,*.bat,*.cmd -ErrorAction SilentlyContinue |
  Select-String -Pattern 'sftp|ftps?://|scp |curl|lftp' -List | Select-Object Path

Every hit is a handover risk with a name on it. The fixes are ordinary. Move jobs to service accounts per service accounts for jobs. Move credentials into a store the role can reach per job credentials storage. Move scripts into the repository. But the sweep finds these risks before the leaving date does. What happens when a personal credential outlives its owner is told in the leaked credential.

Northgate Retail ran that sweep two weeks before an administrator's last day. It found the supplier invoice pull running as her personal account, from a script in her profile. The SFTP password was saved in her desktop client. None of it was a secret and none of it was careless. It was how the flow had been built on a busy afternoon years earlier and never had a reason to be revisited. Moving the task to a service account, the script to the jobs folder, and the password to the vault took one day. On her last day her account was disabled at five, and the pull ran at six as if nothing had happened, which was the whole idea.

Remember: the goal is not to make Dana replaceable. It is to make the flows not depend on her. That way, Dana can take a holiday, get promoted, or work on something interesting without the estate holding its breath.

Shadowing and Handover

The drill is for flows that are staying. Handover is for people who are leaving — or arriving. Both use shadowing, in two directions. In forward shadowing, the newcomer watches the custodian operate the flow and asks questions; this transfers context quickly but tests nothing. In reverse shadowing, the newcomer operates and the custodian watches in silence, allowed only to prevent harm. This is the drill in slow motion, and it is where handovers succeed or fail. Do a week of forward, then at least a week of reverse, and treat every intervention during reverse shadowing as a documentation finding. The silence is the hard part; most custodians need to sit on their hands.

When someone gives notice, the time available is fixed and short, so run the handover from a checklist rather than from memory. This is Meridian's, written after Dana's move to another team gave them four weeks and a strong incentive:

# Handover checklist — <person> → <successor>, target date YYYY-MM-DD

## Week 1 — find everything
- [ ] List every inventory row naming <person> as custodian, backup, or business owner
- [ ] Run the tribal-knowledge sweep: personal-account tasks, profile scripts, personal keys
- [ ] Interview: "how do you know that?" through every Tier 1 record and runbook
- [ ] Vault: every credential <person> can reach is owned by a role, not a person

## Week 2 — forward shadowing
- [ ] Successor watches each Tier 1 flow run and each runbook walked through
- [ ] Every question asked becomes a record or runbook edit the same day

## Week 3 — reverse shadowing and drills
- [ ] Successor operates; <person> silent unless preventing harm
- [ ] Tabletop drill on each Tier 1 flow, successor as responder; score >= 17 or re-test
- [ ] Partners told who the new contact is; contact sheet updated

## Week 4 — cut over
- [ ] Inventory rows reassigned; leavers script confirms no row names <person>
- [ ] Keys and accounts tied to <person> reissued to roles and the old ones retired
- [ ] Review calendar duties reassigned
- [ ] "Questions I still have" list written by the successor and answered
- [ ] Sign-off by both, dated; findings still open become tickets with owners

The credentials and keys items carry the most risk and have their own procedures. These are retiring keys at offboarding, service account hygiene, and the account side in offboarding that closes the account. The checklist points to these rather than repeating them. The "questions I still have" item is the one teams skip and should not. What the successor still does not understand on the last day is the first week's documentation backlog, in priority order.

Handover without a leaver is the same checklist run slowly. When a new administrator joins a stable team, weeks one to three happen over their first two months, with every Tier 1 flow drilled at least once. By the end they are a real backup owner, and the inventory's backup column has become true rather than aspirational. That is the point made in flow ownership and contacts.

Turning Findings Into Fixes

A drill that produces a question log and no changes was a pleasant hour. The value is in the fixes, which are easier if each finding is sorted into one of four kinds. A missing fact — a host name, a path, who R. is — goes into the flow record. A missing step — check this before that, the command is actually this — goes into the runbook. Usually that step becomes a new row in the failure table or a line in section 4. Missing access — the responder could not open the vault entry or log in to the host — is a role fix, not a documentation fix. Missing access is the most common cause of a low score. And knowledge that resists documentation includes judgment about a partner's habits or a feel for what a normal file size is on month-end. It is handled by shadowing and by writing down as much of it as can be written, in the record's notes.

Prioritize by tier and score: the lowest-scoring Tier 1 flow first. Raise one ticket per finding, small enough to close in an hour, owned by the custodian, with the re-test date set when the ticket is raised. Then re-run the drill — with a different responder if you can, because the second stranger finds what the first happened to know. I have been the first stranger, and I happened to know a great deal. Two passes with two different people is Meridian's standard for calling a flow bus-proof.

Feed the findings back into the wider system. A missing-access finding often reveals that the backup role was never granted the rights the custodian has — a gap the quarterly access review should have caught. A pattern across several drills — every responder struggled to find the current log — is not a per-flow fix but a convention to change everywhere. That is the kind of conclusion running your own postmortem is designed to reach. Keep the findings in the loop that keeping transfer documentation current describes: finding, ticket, fix, re-test, and the calendar that schedules the next drill.

Finally, schedule the drills so they happen: one Tier 1 flow a month, rotating, so every critical flow is tested by a stranger once a year. The drill is also the closest thing you have to a rehearsal for the day the whole estate must be rebuilt. A team whose members can each run any flow from the documents can follow the restore order in the disaster recovery series without the one person who wrote it. And restore drills are the same idea at estate scale.

Wrapping Up

The hit-by-a-bus test asks whether a flow depends on a person or on its documentation, and answers by making a stranger try. Run it as a forty-five-minute tabletop with three scenarios and score it out of twenty. Log every question and turn each into a record edit, a runbook row, an access fix, or a shadowing session. Hunt tribal knowledge in interviews and personal-account sweeps. When someone leaves, run the four-week handover checklist; when someone joins, run it slowly. Re-test until two different strangers pass. Then Dana can get on the plane.

Everything the drill checks was built earlier in this series: the rows in the transfer inventory, the instructions in runbooks per flow, and the loop in keeping transfer documentation current. The drill is how you find out whether they work, before a long-haul flight does.

Frequently Asked Questions

What is the bus factor of a flow?
The number of people who would have to be unavailable before nobody could run, fix, or safely change the flow. A bus factor of one means the flow is one absence away from being unrunnable. The drill measures it; the documentation raises it.
Is it safe to run the drill on production flows?
Run the read-only parts for real — finding rows, logs, tasks, and credentials by reference. Do anything that sends or changes data in a staging copy, or have the responder describe the exact command instead. Never resend a real partner file to prove the responder could.
Who should be the responder?
Someone who has never operated the flow, ideally the person named as its backup owner, because they are the one who will genuinely need to do it. The custodian facilitates, stays silent, and writes down every question — that log is the main output.
What score counts as a pass?
Out of twenty, seventeen or more with no facilitator answers needed. Ten or less means the flow has a bus factor of one regardless of what the inventory says. Record the score and date in the runbook header and re-test after the fixes with a different responder.
How often should we run it?
Test one Tier 1 flow a month, rotating, so each critical flow is tested by a stranger about once a year. Also drill every flow a new backup owner inherits, within that owner's first quarter. Any handover triggers drills on all of the leaver's Tier 1 flows.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.