Your data inventory is a list of originals. Attackers don't care which copy they get.
Every sensitive dataset you own has one version that gets all of the attention. It lives in a production database, or in whatever SaaS platform a business unit bought back in 2021, and it's the version that made it onto the architecture diagram and into your records of processing. Encryption wraps that copy. Access reviews cover it, the retention clock runs against it, and your logging can tell you who opened it last Tuesday. Every control you built is aimed at a single address.
And then people do their jobs. An analyst pulls an extract to answer a board question and saves it locally. A vendor asks for a representative sample so they can configure their platform, so someone drops 40,000 rows into a shared folder that was only ever meant to last a week. A BI tool quietly caches results into a store of its own. The finance manager runs the quarterly report on a Friday afternoon and emails it to two colleagues, who each save their own version.
Nobody in that sequence did anything wrong. Every one of them made a copy your controls never saw.
The hiding places repeat, environment after environment. Staging tables in the warehouse that nobody ever bothered to drop. Lower environments seeded with production data for a migration test two years ago, still sitting there because deleting them was never anyone's job. The one that catches people off guard is the ticketing system, where a support engineer pasted a live customer record into a thread to reproduce a bug, resolved the ticket, and never thought about it again. That record is still searchable by everyone who can read the queue.
The copies degrade in predictable ways. No access review touches a spreadsheet on a laptop. Retention schedules don't reach into a shared drive, and nothing logs who opened the vendor's sandbox last March. Same rows, a fraction of the protection, and none of it shows up in the risk score you reported last quarter.
Most teams file this under hygiene. Something to clean up when the audit calendar loosens. What's actually sitting in front of you is a measurement problem. You scored one dataset. Your real exposure is that score multiplied by however many copies exist, and you don't know the multiplier.
The count itself is the interesting part. Copies form where the sanctioned path runs too slow. Every extract on a desktop is quiet evidence that someone needed data faster than your process would deliver it. Find seven copies and you've found a control gap. Find seventy and you've found a broken workflow that your team has been routing around for years.
So, here's the MondayMove
Pick your most sensitive dataset and find every copy of it living outside the system of record.
Start with the one you'd least want to explain in a breach notification, and count.
Discussion