Home › Topics › Chain of Custody › Documenting the Journey

Documenting a File's Journey from Origin to Destination

"Walk us through exactly what happened to the June benefits file." The email is four lines long and copied to three people you have never met. Maybe it comes from a partner disputing totals, maybe from an auditor sampling one transfer, maybe from your own compliance team. Either way, the person asking does not want reassurance. They want a document: every hop, every actor, every timestamp, with evidence behind each line.

You almost never need new tooling to produce one. A usable custody record is assembled from materials your systems already generate — transfer logs, job histories, hash values, receipts. The other ingredient is the discipline of connecting them in time order. The skill is knowing what to pull from where, and what the finished document should look like. The four-line email is the easy part.

This article, part of our Chain of Custody series, does the assembly once, end to end, on a realistic multi-hop flow. We follow one file from the application that created it to its final archived copy. We collect the evidence at each checkpoint and finish with a custody record you can copy as a template. If the concept of custody itself is new, start with the foundation article and come back.

The Flow We'll Document

Our synthetic example: Northfield, a mid-sized employer, sends a monthly benefits contribution file to its benefits administrator, Meridian Bank. The flow has more hops than a beginner would guess, which is exactly why it makes a good worked example. Most real flows look like this. (The ones that look simpler usually have a hop nobody has drawn yet.)

  1. The finance application on ledger01.northfield.example exports benefits_jun30.csv at month end.
  2. An export post-step uploads it over SFTP to the internal transfer server, transfer01.northfield.example.
  3. A scheduled job on transfer01 verifies the file, encrypts it with OpenPGP, and uploads it to the partner's endpoint, sftp.meridianbank.example.
  4. The partner's system processes it and drops an acknowledgment file, which our side retrieves.
  5. The original is moved to an archive folder under a retention rule.

The diagram below shows the journey with its custody checkpoints — the five points where evidence gets created. Each checkpoint answers, for its own hop, the three custody questions: who, when, and what changed.

A benefits file's journey from the origin system ledger01 to the transfer server transfer01, then to the partner endpoint at Meridian Bank, with the original archived. Five custody checkpoints mark where evidence is recorded: origin hash, arrival verification, encrypt and send, partner receipt, and archive.

The Raw Materials, and Where They Live

Before touching the specific file, know your sources. A custody record draws on four of them. Writing down where each lives turns a stressful request into an unhurried one. The flow's runbook is the natural place.

  • Application and job logs at the origin — the export job's own history: when it ran, as which account, what it wrote. Often the most overlooked source, and the only one that covers the file's birth.
  • Transfer server activity logs — every session, login, upload, and download, with account, address, byte count, and result. What belongs in these logs is covered in what to log. How to read them fluently is covered in reading transfer logs.
  • Hash logs — output of the small scripts that compute and compare fingerprints at each checkpoint. These rarely exist by default; they exist because someone decided custody mattered and scripted the checks around the transfers.
  • Receipts — the far side's confirmation, in whatever form the partnership produces: an acknowledgment file, a signed receipt, a log extract they share on request.

One property matters more than any formatting concern: all four sources must be contemporaneous — written by machines at the moment of the event. A record assembled later from these sources inherits their credibility. A record typed from memory after a dispute begins has almost none, which is a distinction legal teams typically care about a great deal.

Walking the Journey, Checkpoint by Checkpoint

The walk goes one checkpoint at a time, pulling the actual evidence each one produced — the same walk you would do with your own logs open in another window.

Checkpoint one: origin

Custody begins when the file does. The export job's log establishes creation — when, by what process, running as which account:

Jun 30 20:58:12 UTC  ledger01  job benefits-export (svc_ledger): run started
Jun 30 20:58:13 UTC  ledger01  wrote D:\exports\benefits_jun30.csv (61,448 bytes, 1,207 rows)
Jun 30 20:58:14 UTC  ledger01  post-step hash-and-log completed
Jun 30 20:58:16 UTC  ledger01  job benefits-export: success

That hash-and-log post-step is a two-line script that computes the file's SHA-256 fingerprint and appends it to a hash log. On a Windows origin it is typically PowerShell:

Get-FileHash -Algorithm SHA256 D:\exports\benefits_jun30.csv

Algorithm  Hash                                                              Path
---------  ----                                                              ----
SHA256     7B42E09C1D5A88F0342AC6B91E0D4F773A5518C2B6E9D00F14A7C83B52961EDA  D:\exports\benefits_jun30.csv

This value — recorded before the file moves anywhere — is the anchor for the entire journey. Every later "did it change?" question is answered by comparing against it. The next article covers the reasons SHA-256 is the right choice here, and exactly what a match does and does not prove. It is hashes as custody evidence. For now, treat the origin hash as the file's recorded fingerprint at birth.

Checkpoint two: arrival on the transfer server

The export's final step uploads the file over SFTP to transfer01. Here the evidence changes hands: the server's activity log takes over as the primary source. Log layouts vary by product, but the events are universal — connection, authentication, transfer, disconnect:

Jun 30 21:02:44 UTC  session 5417 opened from 10.20.4.15 (SFTP)
Jun 30 21:02:45 UTC  session 5417 user svc_ledger authenticated (public key)
Jun 30 21:02:47 UTC  session 5417 upload /inbound/benefits/benefits_jun30.csv 61,448 bytes OK
Jun 30 21:02:48 UTC  session 5417 disconnected

Three details in those four lines quietly do custody work. The account is svc_ledger — a named service identity, not a shared login, so "who" has an answer. The byte count matches the origin log. And the session number lets you pull the complete session if anyone asks what else that connection did. This is why server-side logging is the backbone of custody. A server like Sysax Multi Server records every session, login, upload, and download to its activity log. The log is written to file and optionally a database, with rollover. So this checkpoint exists whether or not anyone was thinking about custody that night.

Arrival is also the first re-verification point. A scheduled task on transfer01 hashes the received file and compares it to the origin value. The technique is the standard end-to-end verification ritual, with one addition — the result is logged, not just glanced at:

Jun 30 21:03:05 UTC  transfer01  verify benefits_jun30.csv sha256=7B42E09C...52961EDA  MATCH origin

Checkpoint three: the outbound leg

At ten past the hour, the outbound job wakes up. In our flow it is a scheduled task in Sysax FTP Automation. It watches the /inbound/benefits/ folder. When a new file lands, its pre-processing step runs the hash check. Its OpenPGP step then encrypts the payload for the partner's key, and it uploads the result over SFTP. The job's run log is your evidence for the entire hop:

Jun 30 21:10:00 UTC  job benefits-outbound started (runs as NORTHFIELD\xferbot)
Jun 30 21:10:01 UTC  pre-process verify_hash.cmd: MATCH against recorded origin value
Jun 30 21:10:03 UTC  pgp encrypt: benefits_jun30.csv -> benefits_jun30.csv.pgp (62,112 bytes)
                     sha256(benefits_jun30.csv.pgp)=C08F44D2...19AB recorded
Jun 30 21:10:09 UTC  upload to sftp.meridianbank.example /drop/ as northfield-payroll: OK
Jun 30 21:10:10 UTC  notify: run summary emailed to transfers@northfield.example
Jun 30 21:10:10 UTC  job benefits-outbound finished: success

The upload happened under northfield-payroll — the named account the partner issued to your organization. Their logs will show that name, which means your record and theirs can be joined later. Then look at line three.

Gotcha: encryption changes the bytes, so it changes the hash. The .pgp file's fingerprint will never match the origin value — and that is fine, because the transformation was intentional and recorded. The custody rule applies to any deliberate change (encryption, compression, renaming with content edits). Log the event, the actor, and both the before and after hashes. An unexplained hash change is evidence of a problem; an explained one is just another link in the chain.

Checkpoint four: the far side

Your own evidence has an honest limit: your logs prove what you sent and that the far end's server accepted it. They cannot prove what happened inside the partner's systems afterward. For that, you need the partner to hand you evidence — a receipt.

In our flow, Meridian Bank's intake process decrypts the payload, verifies it, and writes an acknowledgment file to a pickup folder. Our automation retrieves that file on its next poll:

Jun 30 21:41:22 UTC  retrieved /ack/benefits_jun30.ack from sftp.meridianbank.example

contents of benefits_jun30.ack:
  RECEIVED   benefits_jun30.csv.pgp  62,112 bytes
  DECRYPTED  benefits_jun30.csv
  SHA256     7B42E09C1D5A88F0342AC6B91E0D4F773A5518C2B6E9D00F14A7C83B52961EDA
  STATUS     ACCEPTED

Read that SHA256 line again: the partner decrypted and computed the fingerprint of the plaintext, and it equals your origin value. That single line closes the loop — the file Meridian accepted is byte-for-byte the file the finance system exported, on the partner's own testimony. Not every partnership produces receipts this good. The patterns for getting them — confirmation files, signed acknowledgments, correlated logs — are covered in proof-of-delivery patterns. B2B setups that use AS2 get the formalized version, the MDN, explained in MDNs and proof of delivery. Whatever form the receipt takes, the custody rule is the same: retain it with the record, verbatim, as received. Verbatim. Not summarized, not reformatted, not "the important bit."

And when a partner offers no receipt at all? Your fallback is the transfer transcript from your own client side. The job log's upload ... OK line proves the remote server accepted every byte. That is real evidence, just weaker than an application-level acknowledgment. It proves arrival at their doorstep, not processing inside the house. If the flow matters enough to document, it matters enough to raise receipts with the partner at the next review. The middle of a dispute is the worst possible time to discover the loop was never closed. I have made that discovery mid-dispute twice, and do not recommend it.

Checkpoint five: archive and disposition

Custody does not end at delivery — it ends at documented disposition. The post-processing step moves the original into an archive folder, verifies the hash one last time, and logs it:

Jun 30 21:41:25 UTC  archive: benefits_jun30.csv -> /archive/benefits/ sha256 MATCH origin

The archive copy lives under a retention rule — in this flow, kept for seven years, then purged. The custody record should say so. "Where is it now, and when will it stop existing?" is a question the record must answer for its whole retention life. Retention decisions and the machinery that enforces them are their own subject; our retention and deletion series covers it.

The Copies You Almost Forgot

A careful reader of your record will eventually ask: the archive holds one copy — how many others exist? Walk the journey again counting copies instead of hops. The export wrote one in D:\exports\ on ledger01. The upload created a second in /inbound/benefits/ on transfer01. Encryption produced a third — the .pgp artifact. Delivery placed a fourth on the partner's server, and archiving made a fifth. Five copies of a sensitive payroll file, and the custody record so far only narrates the fate of the one that traveled. Nobody draws the ones that stay home.

A complete record accounts for the stragglers. In our flow, the job's post-processing cleans up as it goes. The staging copy and the .pgp artifact are deleted after successful delivery. Each deletion is a logged event with an actor and a timestamp, exactly like every other custody event. The origin copy on ledger01 is overwritten by the next month's export, and the record says so. Copies that linger unmanaged in staging folders are both a custody loose end and a data-exposure problem. If that sentence describes your servers, the copy-mapping exercise in our retention series is the place to start.

Assembling the Record

The assembly step, done with sources like these, is mostly transcription. Order every event by UTC time; give each line an actor and an evidence source; note every hash result. Format matters less than people fear: plain text or a simple table both work, and plain text ages best. It opens on anything, diffs cleanly, and cannot hide a formula or a stale cached value the way a spreadsheet can. What matters is that the record cites its sources line by line, so any reader can trace every claim back to a log they can inspect. Plain text has also never once auto-corrected a hash.

Here is the finished custody record for our file, in a plain-text format you can copy as a template:

CUSTODY RECORD
File:            benefits_jun30.csv (61,448 bytes, 1,207 rows)
Origin sha256:   7B42E09C1D5A88F0342AC6B91E0D4F773A5518C2B6E9D00F14A7C83B52961EDA
Flow:            ledger01 -> transfer01 -> Meridian Bank; original archived
Prepared by:     P. Rao, infrastructure team, from sources listed per line

#  When (UTC)       Event                                Actor              Evidence source
1  Jun 30 20:58:13  created by benefits-export job       svc_ledger         ledger01 job log
2  Jun 30 20:58:14  origin sha256 recorded               svc_ledger         ledger01 hash log
3  Jun 30 21:02:47  uploaded to transfer01 (SFTP)        svc_ledger         transfer01 activity log, session 5417
4  Jun 30 21:03:05  arrival hash verified: MATCH         xferbot            transfer01 hash log
5  Jun 30 21:10:03  encrypted to .pgp; new hash recorded xferbot            benefits-outbound job log
6  Jun 30 21:10:09  delivered to partner endpoint        northfield-payroll benefits-outbound job log
7  Jun 30 21:41:22  partner receipt: decrypted sha256    Meridian Bank      ack file, retained verbatim
                    MATCHES origin; status ACCEPTED
8  Jun 30 21:41:25  original archived; hash MATCH        xferbot            job log + archive hash log
Disposition:     /archive/benefits/, retained seven years per retention schedule
Gaps:            none identified between events 1 and 8

Notice the last line. Stating "no gaps identified" — or honestly listing the gaps you found — is part of the record. It tells the reader you checked continuity rather than assuming it, and it is the difference between a timeline and a chain.

The Checks Before You Call It Complete

Before the record leaves your hands, walk it against four failure points. These are the places where custody records fall apart under a skeptical reader.

  • Clock agreement. Events 3 and 4 came from different machines. If their clocks disagree, your timeline can show effects before causes, and a challenger will use that. All sources here log UTC from synchronized clocks. If yours don't, fix that before the next dispute, not during it. The article on timestamps and evidence explains how much skew matters and why.
  • Attribution. Every actor in the record is a singular, named identity. If any line says ftpuser or "someone in finance," you have an attribution hole where a custodian should be.
  • Source survival. The record cites logs — so the logs must still exist when questions arrive, which can be months later. Check that log retention outlives dispute windows, and that the cited extracts are preserved with the record itself.
  • Coverage. Scan the intervals between events. Between 21:03 and 21:10 the file sat on transfer01 — who could touch it there? If the folder's permissions restrict it to the two service accounts, say so; that closes the interval. If thirty people can write to it, the record has a soft spot even though every event line is true.

Bluewater Bank assembled a record like this one for a client dispute, and it read beautifully until event 3. That event's evidence source was an activity log that rolled over at ninety days. The dispute arrived on day one hundred and twelve. The origin hash and the client's receipt bracketed the missing hop and settled the content question. But the record went out with "session log no longer available" where a session number should have been. Log retention was changed to two years the same afternoon, a four-minute ticket. The sentence it replaced took a good deal longer to live down. How rollover and log growth behave, and how to keep them from eating evidence, is the subject of log growth and service logs health.

If a check fails, resist the urge to quietly smooth it over. A documented weakness is survivable; a discovered concealment is fatal. What breaks in real custody chains, and how to handle a genuine gap honestly, is the subject of breaks in custody.

Putting It Together

Documenting a file's journey is an assembly job. Origin log and hash anchor the start. The transfer server's activity log carries each hop. The automation job's history covers transformations and delivery. The partner's receipt closes the loop. The archive entry ends the story. Ordered by UTC time with named actors and cited sources, those materials become a custody record that answers "walk us through exactly what happened" in a page.

The deeper skill is noticing that every checkpoint in this article existed because someone configured it in advance. That includes the hash steps, the named accounts, the logging, and the receipt handling. None of it can be conjured after the fact. That is the case for building flows where the record assembles itself, which is exactly where this series ends up: designing flows that keep custody automatically. The four-line email, it turns out, was answered months before it was sent.

Frequently Asked Questions

Do I need to document every file this thoroughly?
No — you need every file on sensitive flows to be documentable, which is different. Configure the checkpoints (logging, hash steps, receipts) once per flow, and the evidence accumulates for every file automatically. You only assemble a formal record when someone asks.
What if one hop has no log at all?
Record the gap honestly in the custody record, and bracket it with the evidence on either side. Use hashes to show whether content survived the unlogged hop intact. Then fix the flow so the gap doesn't recur. Never fill a gap with reconstruction from memory presented as if it were a log.
Does the partner have to cooperate for the record to work?
Your own logs and hashes stand on their own for everything up to delivery. The partner's cooperation — a receipt, an acknowledgment file, a log extract — is what extends the record past the handoff. Agree on receipt format during onboarding; retrofitting one mid-dispute is much harder.
Why record hashes in a log instead of just checking them?
A check that isn't recorded proves nothing later — you can't cite a glance. A logged verification line carries the time, the actor, the value, and the result, which turns a passing check into reusable evidence. The recording habit is the whole difference between verifying transfers and keeping custody.
Who should assemble the custody record when a request comes in?
Whoever administers the transfer systems, because they know the sources. The record should clearly name who prepared it and where each line came from. If the matter is legal, hand your assembled record and preserved sources to counsel. Interpreting and arguing from it is their job, not yours.

From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.