WEEK 32 · DATA

Find every copy of your most sensitive dataset.

Your controls wrap the original. The exposure lives in the copies nobody logged.

Your data inventory is a list of originals. Attackers don't care which copy they get.

Every sensitive dataset you own has one version that gets all of the attention. It lives in a production database, or in whatever SaaS platform a business unit bought back in 2021, and it's the version that made it onto the architecture diagram and into your records of processing. Encryption wraps that copy. Access reviews cover it, the retention clock runs against it, and your logging can tell you who opened it last Tuesday. Every control you built is aimed at a single address.

And then people do their jobs. An analyst pulls an extract to answer a board question and saves it locally. A vendor asks for a representative sample so they can configure their platform, so someone drops 40,000 rows into a shared folder that was only ever meant to last a week. A BI tool quietly caches results into a store of its own. The finance manager runs the quarterly report on a Friday afternoon and emails it to two colleagues, who each save their own version.

Nobody in that sequence did anything wrong. Every one of them made a copy your controls never saw.

The hiding places repeat, environment after environment. Staging tables in the warehouse that nobody ever bothered to drop. Lower environments seeded with production data for a migration test two years ago, still sitting there because deleting them was never anyone's job. The one that catches people off guard is the ticketing system, where a support engineer pasted a live customer record into a thread to reproduce a bug, resolved the ticket, and never thought about it again. That record is still searchable by everyone who can read the queue.

The copies degrade in predictable ways. No access review touches a spreadsheet on a laptop. Retention schedules don't reach into a shared drive, and nothing logs who opened the vendor's sandbox last March. Same rows, a fraction of the protection, and none of it shows up in the risk score you reported last quarter.

Most teams file this under hygiene. Something to clean up when the audit calendar loosens. What's actually sitting in front of you is a measurement problem. You scored one dataset. Your real exposure is that score multiplied by however many copies exist, and you don't know the multiplier.

The count itself is the interesting part. Copies form where the sanctioned path runs too slow. Every extract on a desktop is quiet evidence that someone needed data faster than your process would deliver it. Find seven copies and you've found a control gap. Find seventy and you've found a broken workflow that your team has been routing around for years.

So, here's the MondayMove

Pick your most sensitive dataset and find every copy of it living outside the system of record.

Start with the one you'd least want to explain in a breach notification, and count.

Friday Follow-Up

Where the count breaks.

The place your count stops is a better finding than the number you reach.

MondayMove gives you one concrete action every Monday. FridayFollowUp closes the loop.

Each Friday, a short dispatch on what practitioners actually found when they ran the week's move: where they got stuck, what surprised them, and what to do next. Not sanitized case studies. Field notes. Practitioner to practitioner.

Monday's move was to pick your most sensitive dataset and find every copy of it living outside the system of record.

My guess is the first wall has nothing to do with counting. It's agreeing on which system is the system of record. Security names one platform, the data team names another, and both sides have a case. Have the argument. Just know it can eat a day before anybody counts anything.

The second thing I expect is that you stop at what you can query. Warehouse tables, storage buckets, file shares, whatever your scanning already reaches. Email attachments, ticket threads, and laptops stay invisible, and you finish holding a number that describes your tooling more accurately than it describes your data.

I also don't think you'll finish. I didn't. We could name seven or eight copies of a given dataset and explain exactly why each one existed, and past that the sheer volume beat us. More were out there. We had no realistic way to identify them. If that's where you land, you landed where I did, and it's the normal result rather than a failure.

Here's the failure mode I'd watch hardest. You find that every copy you identified is authorized, documented, part of a real pipeline, and you decide there's nothing to fix. That's the trap. Approved and countable are separate properties, and only one of them helps you when legal asks at two in the morning whether a specific client's data was in scope.

So write down two things this week and put a date on them. The number you reached, and the exact point where the count stopped being possible. The number goes stale inside a month. The reason you couldn't go further will still be true next year, and it's pointing straight at the thing worth fixing.

Then send me the second one. That's the part I actually want to read.

Keep going. See what a week can do.

No correct answers here. This is practitioner-to-practitioner. The more honest the responses, the more useful this gets for everyone reading on Monday morning.

See you then.

Discussion