Verifying Nothing Was Left Behind
The final pass has run, the exit code is in the success family, and the new server holds what looks like everything. Here is the uncomfortable question that decides whether a migration is finished: how do you know? "The copy finished without errors" is a statement about the tool. "Every file made it, byte for byte, permissions intact" is a statement about the data. Only the second one lets you retire a server that held years of the company's work.
Verification is the discipline of replacing hope with proof, and it is cheaper than its reputation. This article builds the method in layers — counts and bytes, difference reports, honest hash sampling, metadata spot checks. Then it shows the two artifacts that make it real. One is a verification report a change board will accept. The other is a worked residue chase that runs a discrepancy down to zero unexplained items.
This is the proof chapter of our Bulk Internal Moves with Robocopy and Rsync series. It assumes the setup from the earlier articles: a seeded, delta-converged copy, and a write freeze in effect while these checks run. Verification of a moving target proves nothing.
Why "It Finished Without Errors" Is Not Proof
Copy tools are honest, but they answer a narrower question than the one you care about. A clean exit means "everything I attempted succeeded." The gaps hide in what was never attempted, or attempted and quietly categorized:
- Excluded on purpose, forgotten in practice. Your
/XDand--excludepatterns skipped directories by design. Was every exclusion recorded? Does anyone remember why? - Failed and logged — in a log nobody read. A locked file that exhausted its retries is a FAILED line in a log from four nights ago. It is not an error in tonight's pass, which simply skipped what was never there to compare.
- Deleted at the destination. A mirroring pass with a wrong path once purged something it should not have, and a later pass did not restore it because the source had moved on.
- Changed after the copy. A file modified between the last delta and the freeze — or a freeze that leaked writes — leaves the target holding yesterday's version with today's confidence.
The principle that catches all four: verify independently of the copy tool. Measure the source with one set of eyes, the target with another, and compare the measurements. The tool that did the copying may participate — its dry-run mode is a superb difference engine. But the census numbers must come from direct measurement of both trees. The layers, cheapest first:
| Layer | Question it answers | Cost on four terabytes |
|---|---|---|
| 1. Counts and bytes | Is anything missing or extra, in bulk? | Minutes to tens of minutes (tree walk) |
| 2. Difference report | Exactly which items differ, and how? | One dual tree walk |
| 3. Content hashes | Are the bytes truly identical? | Full: many hours (reads every byte twice). Sampled: minutes |
| 4. Metadata spot checks | Did permissions, owners, timestamps survive? | Minutes (sampled) |
Layer 1: Counts and Bytes
The census is the backbone: total files, total directories, total bytes, measured the same way on both sides. On Windows, PowerShell gives exact numbers; on Linux, find and du do:
# Windows (run against each side; save the output)
$m = Get-ChildItem \\server\share -Recurse -File -Force |
Measure-Object -Property Length -Sum
"{0} files, {1} bytes" -f $m.Count, $m.Sum
(Get-ChildItem \\server\share -Recurse -Directory -Force).Count
# Linux (each side)
find /srv/data -type f | wc -l # files
find /srv/data -type d | wc -l # directories
du -sb /srv/data # apparent size in bytes
The comparison is meaningful only if you respect four caveats. First, freeze first: counts taken on a live share drift while you take them; the authoritative census happens inside the cutover window, source frozen. Second, subtract the exclusion ledger: if 62 directories were excluded by design, the source census must be adjusted before comparing. That is why every exclusion was written down when it was added. Third, measure the same thing: du without -b reports disk blocks, which differ across filesystems even for identical data. Compare apparent sizes, and include hidden and system files on both sides (-Force above). Fourth, know your special cases: hard links copied without preservation become multiple independent files (counts match, bytes inflate on the target). And sparse files can legitimately occupy different disk space while remaining byte-identical. Both are explained in bulk move pitfalls.
Count directories separately from files — the two numbers fail differently, and the difference diagnoses. A file gap with matching directory counts points at locked or forbidden files. A directory gap points at structural skips. Its classic cause is an engine run with the skip-empty-directories option (/S instead of /E). That silently drops every empty folder in the tree. Empty folders are cheap to lose and maddening to explain three weeks later when an application expects its drop folder to exist.
Matching counts and bytes do not prove the contents are right — that is layer 3's job. But a mismatch here is the cheapest possible tripwire. The sizes of the mismatch (91 files? one file? a terabyte?) immediately shapes the investigation.
Layer 2: The Dry Run as Difference Report
Both migration engines can compare trees without copying — and pointed at a frozen source and a finished target, that comparison is your difference report, item by item.
With robocopy, /L lists what a pass would do without doing it. Run your exact mirror command with /L and a log, and read the classes. New File means present at source, missing at target — a gap. Newer/Older mean the copies differ by timestamp. EXTRA File means the target holds something the source does not. A clean report shows nothing but the summary, with zeros in the Copied and Extras columns:
robocopy \\oldserver\eng \\newserver\eng /MIR /L /NP /NJH /FP /NDL /LOG:C:\miglogs\eng_verify_diff.log # clean output ends like: # Total Copied Skipped Mismatch FAILED Extras # Files : 3214853 0 3214853 0 0 0
With rsync, the equivalent is the dry run with itemized changes: rsync -aH -x -n -i --delete /srv/data/ root@newserver:/srv/data/. Silence is success; any output line is a difference, with the leading code telling you what kind. The >f code means a file that would transfer. A code with t means a timestamp-only difference. The *deleting code means an extra on the target. Decoding itemize output fluently is covered in rsync dry runs and verification.
Run the difference report after the final frozen pass, as the cutover runbook schedules it. At that moment the truthful result is "nothing to do" — and every line that appears anyway goes straight onto the residue ledger for the chase below.
Layer 3: Content Hashes, Honestly
Counts prove presence; hashes prove identity. A hash is a short fingerprint computed from a file's entire contents. Identical fingerprints on both sides mean identical bytes, full stop (the mechanics are in hashing explained). The catch is arithmetic: hashing reads every byte. So fully hashing four terabytes means reading four terabytes on each side — roughly the cost of another seed pass, hours upon hours. Demanding 100 percent hashing of a large migration is usually how verification gets skipped altogether. The honest alternative is sampling plus targeting:
- A random sample, stratified by size. A few hundred files drawn across small, medium, and large buckets. Random matters: it catches systemic problems (truncation, encoding damage) wherever they hide.
- The crown jewels, in full. Finance, contracts, the folders whose loss would be existential — hash every file. Small subtrees, total coverage, disproportionate reassurance.
- Everything that ever failed. Every file that appeared in any FAILED or error line across all pass logs gets hashed on both sides. These files had eventful journeys; they have earned the scrutiny.
- The largest files. The biggest movers are the likeliest to have been interrupted and resumed; hash the top few dozen.
# Windows - hash one file on each side and compare
Get-FileHash -Algorithm SHA256 \\oldserver\eng\finance\ledger.xlsx
Get-FileHash -Algorithm SHA256 \\newserver\eng\finance\ledger.xlsx
# (certutil -hashfile file SHA256 works without PowerShell)
# Linux - hash a whole sampled subtree, then check on the target
cd /srv/data && find finance -type f -exec sha256sum {} + > /tmp/finance.sha
cd /srv/newdata && sha256sum -c /tmp/finance.sha | grep -v ': OK$'
The manifest pattern in that Linux snippet is to generate fingerprints on one side and verify them on the other. It scales from a folder to a whole share and leaves an audit trail. See checksum files and manifests for proper coverage. If your move ran over rsync, note that rsync -c -n (checksum dry run) is a full-content comparison in one command. It is priced accordingly, since it too reads everything on both sides. Whatever you choose, record what was hashed, with which algorithm, and the result: sampled verification is legitimate exactly insofar as it is documented.
Remember: full hashing is a budget decision, not a virtue. Four terabytes hashed both sides can add a day. A documented sample of a few hundred files plus total coverage of the critical folders and every logged failure catches the same classes of damage in minutes.
Layer 4: Metadata Spot Checks
Files can arrive byte-perfect and still be broken for users, because permissions, owners, or timestamps did not survive. Reuse your hash sample set and compare the metadata around it. On Windows, icacls \\server\share\dir /save out.txt /t exports the ACLs of a subtree to a file. Export the same subtree on both sides and compare the files. The dir /q command shows owners at a glance. On Linux, compare ls -l and, where ACLs are in play, getfacl -R output for sampled directories. Timestamps deserve one deliberate glance too. If every folder on the target claims it was modified the night of the migration, directory timestamps were not preserved. That is cosmetic, but the kind of cosmetic that breaks date-based cleanup scripts later.
Interpret findings by shape. One folder with wrong permissions is a local fix. Every file owned by the migration account, or every ACL flattened to the destination default, is a systemic flag error. The copy ran without /COPYALL or without root. The honest cure is rerunning a security-fix pass with the right flags (see the robocopy article on /SECFIX). It is not hand-patching ACLs for a week.
The Verification Report
Verification ends in a document — the artifact that lets a change board approve retiring the source, and the answer to any later "did everything really move?" A working skeleton:
MIGRATION VERIFICATION REPORT - engineering share Scope : \\oldserver\eng -> \\newserver\eng State : source frozen (read-only) throughout all checks Checked : Sat 00:10 - Sat 01:05 Verifier: SAM Approver: RB 1. CENSUS (independent measurement) SOURCE TARGET Files _______ _______ Directories _______ _______ Bytes (apparent) _______ _______ Adjustments: exclusion ledger v3 (62 dirs, documented) 2. DIFFERENCE REPORT Tool/command: robocopy /MIR /L (log: eng_verify_diff.log) Result: 0 to copy, 0 mismatched, 0 extras [ ] clean 3. RESIDUE LEDGER (every discrepancy, dispositioned) path | first seen | reason | disposition (fixed/excluded/accepted) 4. CONTENT SAMPLE 240 random files (3 size buckets) + finance subtree (100%) + 25 files from failure logs + 30 largest files Algorithm: SHA-256 Mismatches: 0 5. METADATA SAMPLE ACL export compared on 12 subtrees; owners on 40 files Directory timestamps spot-checked Findings: none 6. SIGN-OFF Verification passed / passed-with-notes; logs archived at: ______
The report is short on purpose. Its power is in what it forces: independent numbers written down, every difference dispositioned, and a named human signing that the residue ledger reached zero unexplained lines.
Archive the evidence with the report: every pass log, the difference-report output, the hash manifests, the exclusion ledger. Compressed, the whole bundle is a few megabytes — nothing against the comfort it buys when, months later, someone asks where a file went. The honest answer to "did it move?" is then a lookup, not an argument. Keep the bundle at least until the source hardware is disposed of, and longer if the data falls under a retention policy.
Chasing the Residue: A Worked Example
Here is how the chase actually goes. Final frozen pass complete; census says the source holds 3,214,882 files and the target 3,214,853. Gap: 29 files. Nobody panics — the ledger opens.
Step one: the difference report names them. The robocopy /L run lists 25 would-be copies and 0 extras. Cross-checking the pass logs explains the other four immediately. They are junction points excluded by /XJ, present in the exclusion ledger. 29 becomes 25 unexplained.
Step two: group by error, not by file. Of the 25, the final pass log shows 17 Access is denied failures, all under one former executive's home folder. Cause: an explicit deny entry on the folder, blocking even the migration account. Fix: rerun that subtree with backup mode (/B, which reads past ACLs by privilege), then re-verify. 25 becomes 8.
Step three: the stubborn stragglers. The remaining 8 are mailbox archive files, each failing with a sharing violation — something still holds them open during the freeze. The open-files view on the old server points at an archiving service nobody froze. Stop the service, rerun the subtree, watch the difference report drop to zero. 8 becomes 0.
Step four: close the loop. Census rerun: counts and bytes now match exactly after ledger adjustments. The 29 lines in the residue ledger each carry a disposition — 4 excluded-by-design, 25 fixed-and-verified. Zero unexplained. That is what "nothing left behind" means: not a feeling, a ledger that adds up.
Every chase visits the same suspects, which is why the pitfalls article reads like a field guide to residue. Those suspects are locked files, permission walls, path-length oddities, hard links, and the writer process nobody froze. Suppose a migration leg crossed between sites through a transfer server — say an SFTP drop via Sysax Multi Server. In that case, its activity log is one more reconciliation source, a per-file record of what actually crossed. That is precisely the kind of evidence trail described in verifying transfers end to end.
Proof Is What Lets You Let Go
The migration is not done when the copy finishes; it is done when the report is signed. Census both sides independently, and drive the difference report to empty. Hash a documented sample plus the files with eventful histories. Spot-check the metadata, and chase the residue ledger to zero unexplained. Then — and only then — the old server can be retired without a knot in anyone's stomach.
Verification also outlives the migration. If what remains is a standing nightly mirror between sites, give that job its own ongoing checks. That means schedules, retries, and failure notifications rather than one-time reports. A scheduler like Sysax FTP Automation emails on failed runs, which is the everyday cousin of tonight's sign-off. For the night this all happens in sequence, keep the cutover runbook beside this article. For the engines' own verification switches, see the rsync deep-dive.
Frequently Asked Questions
Isn't a clean robocopy or rsync exit code enough proof?
Do I have to hash every file to be sure?
Why do my file counts differ even though the copy looks complete?
When exactly should verification run?
What if a discrepancy cannot be fixed tonight?
From the Sysax team: we build secure file transfer software for Windows. Sysax Multi Server is an FTP, FTPS, SFTP, and HTTPS server. Sysax FTP Automation handles scheduled, scripted transfers. Free trials are on the download page.
