It’s AI Time. Do You Know Where Your Data Is?
Your files may be right where you left them. The information inside them could be moving through prompts, summaries, databases, and agents. Can you trace the path?

Do you know where your data is?
It’s an easy question to answer with confidence. Until someone asks you to show the path.
Picture this. Someone in finance pastes a client spreadsheet into an unapproved public chatbot to “clean up the formatting.” An ops lead spins up an agent over the weekend that nobody in IT knows exists. That agent calls a tool, the tool writes to a database, and another agent reads from it.
Where did the sensitive data go?
If nobody can trace those handoffs, “we know where our data lives” starts sounding a lot like a guess.
I watched a breakdown this week on AI data exposure, and it got me thinking about how much harder that question becomes when the data changes shape.
A customer record gets converted into an embedding, a numerical representation used for retrieval. A contract becomes a summary. A spreadsheet becomes context for an agent that passes information to three more agents.
The original file might still be sitting exactly where you left it. But information from it can now exist in other forms, in other places.
That’s the part we need to pay attention to.
Data loss prevention tools still matter. But recognizing sensitive content and tracing its entire journey are different jobs. A control that catches a spreadsheet upload doesn’t automatically tell you what happened to the summary, the stored context, or the next tool call.
So the question isn’t only, “Which AI tools are we using?”
It’s, “Can we follow sensitive information from its source, through each transformation and handoff, to where it gets stored or shared?”
Seeing the laptop is one piece. Seeing the cloud storage is another. Seeing an agent’s prompts is another. Somebody has to connect those pieces.
This is why I keep pushing the boring stuff.
Know where your AI runs. Know what it’s allowed to touch. Know which tools and outside services it can call. Keep a record of what it does. Put a human at the decision points where sensitive information could be shared or consequential changes made.
And yes, I’m a big believer in keeping control of your data. But running AI inside your own walls doesn’t answer every question either. You still need to understand permissions, logs, and what can leave through a connected tool.
If you can’t draw the path your data takes through your AI, there’s a gap in your AI strategy.
Hope is not an audit trail.
If someone asked you tomorrow where your most sensitive data has been used in AI, could you show them?
Or would you be guessing too?
