A cybersecurity research firm, Harmonic, found that roughly 8.5% of employee prompts submitted to tools like ChatGPT and Copilot contained sensitive data — customer information making up the largest share, followed by employee personal details and financial or legal information. More than half of those exposures happened on free-tier platforms, which routinely use submitted queries to train their models further. Multiply that pattern across a workforce of any real size, and it becomes clear that generative AI isn't just a new productivity tool sitting inside a business — it's a new, largely unmanaged channel through which personal data is leaving the organization every day.

 

This is a distinct problem from the broader question of how DPDP applies to AI systems generally. Generative AI creates specific risk patterns — prompt-level data leakage, hallucinated personal information, and customer-facing chatbots quietly becoming data controllers in their own right — that deserve their own answer.

 

The Prompt Box Is a Data Export You Didn't Approve

Every time an employee pastes a customer's details into a generative AI tool to draft a response, summarize a complaint, or "clean up" a dataset, that's personal data leaving the organization's own systems and entering a third party's infrastructure — often one hosted outside India, governed by terms most employees have never read. Under the DPDP Act, that's processing, and if the underlying data belongs to a customer or employee, the business collecting it remains the accountable Data Fiduciary regardless of which tool an individual employee happened to reach for.

 

The well-publicized incident involving Samsung engineers pasting proprietary source code into ChatGPT — which was then potentially incorporated into the platform's training data — is a useful illustration of exactly this pattern, even though that case involved trade secrets rather than personal data specifically. The mechanism is identical either way: once information crosses into a third-party AI tool's systems, the originating organization has effectively lost control over what happens to it next.

 

Where the Risk Actually Concentrates

Risk PatternWhat It Looks LikeWhy It Matters Under DPDP
Employee prompt leakageStaff pasting customer records, contracts, or HR data into a public AI tool to get help drafting somethingThe business remains the Data Fiduciary for that data, but has no way to enforce security safeguards on a third-party platform
Free-tier training exposureData submitted to a free consumer AI tool gets used to further train the underlying modelThe data's use has extended well beyond the purpose it was originally collected for, with no fresh consent obtained
Customer-facing AI chatbotsA business deploys an LLM-based chatbot on its website or app, and it logs full conversation transcriptsThe chatbot vendor may be processing personal data as an independent Fiduciary if it reuses those logs for its own purposes
Hallucinated personal informationA generative AI tool produces confident but factually incorrect information about a real, identifiable individualRaises the Act's accuracy expectations for any business that relies on or republishes that output

 

When a Customer-Facing Chatbot Becomes Its Own Data Problem

Deploying an AI-powered chatbot for customer support is now a common, sensible business decision — but it introduces a data relationship that's easy to overlook. Every conversation a customer has with that chatbot is a fresh instance of personal data collection, often including account details, complaints, or transaction history volunteered mid-conversation. That data needs its own notice under Section 5, describing what's collected and why, and its retention needs the same purpose-based limits as any other customer data.

 

The added wrinkle is the chatbot vendor's own role. If the underlying AI platform retains those conversation logs to improve its general-purpose model — rather than solely to serve that one business's chatbot — the vendor has stepped outside the boundaries of a straightforward Data Processor relationship. That's the same accountability question that runs through most AI vendor relationships: check what the contract actually permits, not what the marketing page implies.

 

Hallucination Is an Accuracy Problem, Not Just a Reliability One

Generative AI tools are well known for producing plausible-sounding but incorrect information — and when that information relates to an identifiable, real person, it stops being a mere reliability quirk and becomes a data accuracy issue. A business that uses a generative AI tool to draft a customer communication, generate a report referencing an individual, or summarize records, and republishes a fabricated detail about that person, has put inaccurate personal data into circulation under its own name. The DPDP Act's broader expectation that businesses maintain accurate records doesn't pause simply because an AI system, rather than a human, introduced the error — any output referencing a real individual still needs to be checked before it's relied upon or shared further.

 

Building a Practical Generative AI Policy

Given these specific risk patterns, a generative AI usage policy needs to go further than a generic "use AI responsibly" memo. Worth including:

  • A clear list of what can and can't be pasted into public AI tools — customer records, employee data, and anything identifying a real person should be treated as off-limits for consumer-tier tools by default
  • Preference for enterprise-tier or contracted AI tools where the vendor agreement explicitly excludes using submitted data for model training, over free consumer versions with vague or permissive terms
  • A verification step before publishing AI-generated content that references any real individual, treating hallucinated details as a compliance risk, not just an embarrassment
  • Explicit notice and scoped retention for any customer-facing AI chatbot, matching the same standard applied to any other data collection point
  • Basic monitoring or data-loss-prevention controls, where feasible, to catch obvious instances of sensitive data being pasted into unauthorized tools
     

The Broader Point: Convenience Doesn't Pause Accountability

None of this is an argument against using generative AI — the productivity gains are real, and businesses adopting these tools thoughtfully aren't doing anything wrong. The issue is that the DPDP Act's accountability structure doesn't have an exception for "the employee didn't realize the tool worked that way." Whoever collected the personal data in the first place remains responsible for what happens to it, even three steps removed, inside a chat window an IT team may not even know is being used.

 

The Blind Spot Most Compliance Programs Haven't Reached Yet

Ask most compliance teams to list every system that touches customer data, and they'll produce a reasonably complete answer — CRM, HRIS, payment gateway, the usual suspects. Ask them which employees used ChatGPT last week, on what data, and under whose account, and the answer is usually silence. That's not a failure of effort. It's that generative AI adoption happened at the individual level, tool by tool, browser tab by browser tab, faster than most governance processes were built to track.

 

Closing that gap doesn't start with a policy document — it starts with finding out what's actually happening. A short, honest survey of which AI tools different teams are already using, followed by a look at whether those tools are free-tier or contracted, is often enough to reveal where the real exposure sits. From there, the fix is usually less about restricting AI use and more about redirecting it — toward tools with contracts that actually say something, and away from the free version someone bookmarked eighteen months ago.

If your organization hasn't done that basic inventory yet, it's worth doing before writing the policy, not after.