Using Hashes as Chain-of-Custody Evidence
Everything else in a custody record is testimony. This account logged in, this file was uploaded, this job ran — true, but the reader is being asked to believe your machines. A hash — the short fingerprint computed from a file's bytes — is the one line a skeptic can recompute for themselves and get the same answer. That makes hashing the most objective link in the chain, the part that turns "we believe nothing changed" into "check for yourself."
A hash carries that weight only if it was recorded with discipline. It must be computed at the right moments, written down with its context, and stored where later tampering can be ruled out. A value checked once on a screen and written nowhere proves exactly nothing three months later. And even a perfectly recorded hash proves less than people assume. Know the boundary between what it demonstrates and what it cannot. That is the difference between using hashes as evidence and waving them around.
This article, part of our Chain of Custody series, covers the discipline. It covers the hash-verify-record pattern, the honest limits of a matching value, and choosing an algorithm you can defend. It explains the format of a verification record and manifests for batches. It also covers where hash records should live so they survive scrutiny. The mechanics of hashing itself — what the algorithms do and how to run them — are covered in hashes explained. Here we take the mechanics as given and focus on the evidence.
The Pattern: Hash at Origin, Verify at Every Hop, Record Every Result
The custody use of hashing is a three-beat rhythm, repeated along the file's whole journey.
Hash at origin. Compute the fingerprint the moment the file is complete, before it moves anywhere — ideally as the final step of the process that created it. This value is the anchor: every later comparison refers back to it. An origin hash computed hours after creation is weaker, because the unhashed interval is exactly the kind of gap custody exists to eliminate.
Verify at every hop. Each time the file lands somewhere new — the transfer server, the DMZ relay, the partner's intake — recompute and compare against the origin value. A hop that is never verified is a hop where alteration cannot be ruled out; the whole journey is only as attested as its least-checked segment. This is the same ritual as end-to-end transfer verification, applied at each link instead of only the two ends. For payloads large enough that every re-hash costs real minutes, see verifying large transfers.
Record every result. This is the custody-specific part, and the one most shops skip. A check that leaves no record has no evidentiary existence — you cannot cite a glance. Every verification becomes a log entry carrying six things. These are the timestamp, the file, the algorithm, the computed value, the comparison result, and the actor that ran the check. Six fields, one line, appended to a hash log that never gets edited afterward. Never. Not to fix a typo, not to tidy the columns, not once.
Do that consistently and something quietly powerful happens: the "what changed?" column of your custody record fills itself in with mathematics. Any interval between two recorded MATCH entries is an interval where content alteration is ruled out, no matter who had access during it.
What a Matching Hash Actually Proves
This is exactly where challengers probe, so precision matters. A match between the value recorded at origin and the value computed at a later checkpoint proves one thing: the bytes at the checkpoint are identical to the bytes that existed when the origin hash was computed. Nothing was added, removed, or altered in between — not by transfer corruption, not by a truncated upload, not by a helpful colleague "fixing" a header, not by an attacker. For a properly chosen algorithm, the chance of two different files sharing a fingerprint by accident is so small that no practical process will ever witness it. You, your successor, and the server will all retire first.
That single guarantee, repeated at every hop, supports the claims a custody record actually needs. The file the partner processed is the file we exported. The archived copy is authentic to the original. The content survived the unlogged interval intact. That guarantee converts the weakest kind of dispute — my copy against your copy, memory against memory — into a computation either side can rerun.
What a Matching Hash Cannot Prove
Now the boundary. Five claims regularly get attached to hash evidence that hashes do not support, and a challenger who knows the difference will dismantle a record that leans on them.
- That the content was ever correct. A hash fingerprints bytes; it does not judge them. If the export job produced a file with the wrong rows, the hash faithfully certifies the wrong rows all the way to the archive. Origin correctness is the application's problem, not the fingerprint's.
- Who created or sent the file. Anyone can compute a hash of anything. Binding content to a party's identity takes a digital signature — a different tool for a different question, unpacked in integrity vs authenticity.
- That nobody read or copied it. Hashes speak to alteration only. A file can be exfiltrated wholesale and its hash will match forever. Confidentiality evidence comes from access logs and permissions, not fingerprints.
- When anything happened. A hash value contains no clock. Timing comes from the log entry the hash sits in. That entry is only as credible as the machine's clock and the log's integrity, the territory of timestamps and evidence.
- That the recorded origin value is itself honest. If the person who altered the file could also rewrite the hash log, the match proves nothing. Hash records need their own custody — the reason storage gets its own section below.
The boundary in table form:
| Claim | Does a matching hash support it? | What actually supports it |
|---|---|---|
| Content unchanged between two recorded checks | Yes — this is the guarantee | The two logged verification events |
| Transfer completed without corruption | Yes, for that hop | Hash check on arrival vs origin |
| The exported data was correct | No | Application-level checks at origin |
| A specific party sent it | No | Digital signatures, authenticated sessions |
| Nobody viewed or copied the file | No | Access logs, permission records |
| The event happened at the logged time | No — the log carries the time | Synchronized clocks, protected logs |
| The hash record itself is trustworthy | No — it must earn trust separately | Separate, restricted, append-only storage |
Remember: a hash answers exactly one custody question — "what changed?" — and answers it superbly. The other two questions, who and when, come from authenticated logs and synchronized clocks. Presenting hash evidence as if it answered all three is the fastest way to lose a technically-informed audience.
Choosing an Algorithm You Can Defend
Custody work is adversarial by definition — the record exists for the day someone hostile examines it. That decides the algorithm question. SHA-256 is the working default. No practical method exists for constructing two different files with the same SHA-256 fingerprint, so a match cannot be explained away as manufactured. The algorithm is also universally available, which matters when a partner or examiner needs to recompute your values on their own tooling.
MD5 deserves an honest sentence rather than a reflexive ban. Its weakness is specific: collisions — pairs of different files with identical MD5 fingerprints — can be deliberately constructed. Random corruption will still virtually never produce one, so MD5 remains serviceable for detecting accidental damage in transit. But custody is about ruling out deliberate acts, and an algorithm whose matches can be engineered gives a challenger exactly the opening they want. The practical rule: MD5 for accidental-corruption detection where it is already entrenched, SHA-256 anywhere tampering matters — which is everywhere custody matters. Expect the question, too. When hash evidence surfaces in a dispute, legal teams typically ask whether the algorithm is one whose collisions are known to be constructible. And "we used the industry-standard choice" is the answer you want available.
Two consistency rules finish the job. Record the algorithm name alongside every value — a bare hex string in a log is ambiguous evidence. And use one algorithm across the whole flow, origin to archive, so every checkpoint compares like with like. (Sixty-four hex characters with no label is a very confident-looking mystery.)
Recording a Verification Event Properly
The tools are the ordinary ones. On Windows, PowerShell's Get-FileHash or the built-in certutil; on Linux, sha256sum. What upgrades them from spot-check to evidence is wrapping them in a few script lines. Those lines append a structured entry to a log instead of printing to a screen that forgets. A workable format, one line per event:
timestamp(UTC) | host | actor | file | algorithm | value | result Jul 15 20:41:09 | ledger01 | svc_ledger | statements_jul15.csv | SHA256 | 4E1FA0...B39C | ORIGIN Jul 15 21:03:12 | transfer01 | xferbot | statements_jul15.csv | SHA256 | 4E1FA0...B39C | MATCH Jul 15 21:36:50 | transfer01 | xferbot | statements_jul15.csv | SHA256 | 4E1FA0...B39C | MATCH (pre-send)
For reference, the underlying commands the wrapper script calls — one per platform, all producing the same value for the same bytes:
PowerShell: Get-FileHash -Algorithm SHA256 statements_jul15.csv cmd.exe: certutil -hashfile statements_jul15.csv SHA256 Linux/macOS: sha256sum statements_jul15.csv
The first entry, marked ORIGIN, is the anchor; each later entry states its comparison result explicitly. Log full values rather than the shortened forms shown here for print — truncation is for human eyes, never for records. A mismatch gets logged too, loudly, as MISMATCH with both values. A failed check that vanishes because the script only records successes is a custody record lying by omission.
Acme's verify script logged only successes, by accident of a redirect that sent errors to the console instead of the file. For eleven months the hash log was a wall of MATCH, which everyone found reassuring. Then a partner reported missing rows. The night in question turned out to be the one where an upload had truncated at 61 of 62 kilobytes. The script had printed MISMATCH to a window nobody was watching and written nothing to the log at all. The hash log, produced as evidence that the file arrived intact, showed a clean run with no entry for that date. The fix was one line. The explanation took considerably longer.
Where do these checks run? At origin, as the export's final step. At each hop, as a step in the transfer job itself — which is the honest way to automate custody hashing. Transfer tools do not need built-in hashing when they can run your script at the right moments. In Sysax FTP Automation, for example, a scheduled transfer's pre-processing step runs the verify script before the file leaves. The post-processing step records the result and can fire an email notification on failure. So the rhythm of hash-verify-record happens on every run, unattended, in the same order, without anyone remembering to do it.
Manifest Patterns for Batches
Real flows rarely move one file. For a batch, the tool is a manifest — one file listing each payload file with its hash. The formats and creation commands are covered in checksum files and manifests; custody adds three habits on top of the mechanics.
- Treat the manifest as the batch's custody anchor. One document now attests the exact membership and content of the whole delivery. That also catches the failure a per-file check misses: the file that never arrived at all.
- Hash the manifest itself, and record that value in your hash log. The manifest is the linchpin, so it gets its own fingerprint. One recorded value now anchors the entire batch, and a swapped or edited manifest becomes detectable.
- Let the partner's verification become your receipt. Send the manifest with the batch; the partner verifies every file against it and reports the result. That report — "all files verified against your manifest" — is delivery evidence and content evidence in one artifact, worth retaining verbatim alongside your own logs.
Where Hash Records Live So They Survive Scrutiny
Recall the fifth limitation: a hash record is only as trustworthy as its own custody. The moment you imagine a hostile reviewer, the storage requirements write themselves.
Separate from the files they describe. The core rule. If the same account that can modify a payload file can also rewrite the hash log entry for it, the log rules nothing out. Keep hash logs on a different system — or at minimum under different permissions — than the transfer folders, so falsifying content and record requires two distinct compromises. On a flow through a transfer server, the server's own activity log provides a further independent witness. Consider a session log from Sysax Multi Server showing no uploads to a folder between two verification events. It corroborates the hash log's claim that nothing changed there — two sources, separately maintained, telling one story.
Append-only in practice. Entries get added, never edited. Grant the scripts' service accounts append rights, and keep rewrite rights away from daily-driver accounts. Ship copies onward so no single machine holds the only version. These are the same disciplines as protecting any security log, covered in log tamper resistance and centralizing logs.
Retained as long as questions can arrive. A hash log that rolls over after ninety days is useless in a dispute that surfaces after a year. Set the hash log's retention to match the dispute horizon of the data it covers. For financial and health flows that is typically years. In words: keep the records for at least seven years if the underlying files carry that kind of obligation. Decide deliberately rather than by default. The rollover settings that make this decision for you when nobody else does are covered in log growth and service logs health.
Readable by the people who will ask. Evidence that only one admin can find is evidence that goes on vacation with them. Auditors and investigators should be able to get hash records through a documented path, without anyone hand-editing what they see on the way out. I have been that admin, and the vacation was not improved by the phone calls.
A Hash Trail in Action
A dispute is where the pieces come together. Months after the July statements run, a supplier claims the file Northfield sent contained a duplicated page of line items. The admin pulls two sources. One is the hash log entries shown above. The other is the delivery receipt in which the supplier's own intake reported SHA256 4E1FA0...B39C VERIFIED on arrival. The chain now reads: origin value recorded at export; MATCH on the transfer server; MATCH immediately before sending. Then the same value was verified by the recipient's own system on receipt.
Whatever produced the duplicated page, it did not happen to this file between export and delivery. Every interval is bracketed by recorded, matching fingerprints, and the last bracket is the supplier's own report. The investigation moves, correctly, to the export logic and to the supplier's post-receipt processing. That is hash evidence doing its real job: not winning an argument, but eliminating segments of the journey from suspicion so attention lands where the problem actually is. The full assembly of such a record — logs, receipts, and hashes in one document — is worked through in documenting a file's journey.
The Version to Tell a Colleague
Hashes are the recomputable part of a custody record. Hash at origin, verify at every hop, and log every result with its time, actor, algorithm, and outcome. A matching SHA-256 proves the bytes did not change between two recorded checks. It proves nothing about correctness, identity, confidentiality, or timing, which come from logs, signatures, and clocks. Use SHA-256 wherever tampering matters, and keep MD5 only for accidental-corruption detection. Store the hash log where whoever could alter a file could not also alter its record. Do that routinely and most "did it change?" disputes end in a lookup. The rest end in a lookup and a short conversation about where the script sends its errors.
Next in the series: the concept behind all of this in chain of custody, explained if you skipped ahead. Find the design that makes the whole rhythm automatic in custody by design.
Frequently Asked Questions
Is MD5 ever acceptable in a custody record?
How often should I hash a file that just sits in an archive?
Do I need to hash the encrypted file, the plaintext, or both?
What if two hash tools give me different-looking results for the same file?
Should the hash log include files that failed verification?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
