DPDP Compliance for AI Products: What Counts as "Personal Data" When Your Pipeline Runs on Prompts and RAG Context
Most DPDP explainers treat personal data like a database row. If your product is an AI feature - a chatbot, a RAG search tool, an agent reading tickets - that mental model breaks. Here's where personal data actually lives in an AI pipeline.

Most DPDP explainers treat personal data like a database row - a name in a users table, an email in a customers table, something you can point to, find, and delete. If your product is an AI feature - a support chatbot, a RAG-based search tool, an agent that reads customer tickets - that mental model breaks almost immediately.
The Act doesn’t care that the data moved from a column to a prompt. Section 2(t) defines personal data as “any data about an individual who is identifiable by or in relation to such data” - full stop, no carve-out for unstructured text. Section 2(x) defines “processing” to include collection, storage, use, and retrieval by any “automated” means, and Section 2(b) defines automated as any digital process operating on its own. An LLM inference call is squarely inside that definition. So is embedding a document into a vector store. So is caching a model’s response.
Where personal data actually lives in an AI pipeline
Walk a typical RAG-based support tool through the Act’s definitions and you find personal data in more places than the schema suggests:
- Input prompts. A user types “my order #4521, I’m Priya Sharma, it hasn’t arrived” into a chat box. That string is now personal data the moment it’s logged, regardless of whether it ever touches a structured table.
- Retrieved context. If your RAG system embeds customer records, support tickets, or contracts into a vector store, those embeddings are a new representation of the same personal data - and under the Act’s broad definition of “data,” a representation counts.
- Model outputs. A generated response that repeats a customer’s details back, or a summary that includes a name and complaint, is personal data in a new artifact - often one that lives in logs your data-mapping exercise never looked at.
- Fine-tuning and eval sets. If production conversations get pulled into a training or evaluation dataset, personal data has moved into your model weights and your test fixtures. Nobody’s DPDP checklist usually covers either.
Why the standard advice doesn’t transfer cleanly
The usual compliance playbook - map your data, get consent, honor erasure requests within your retention window - assumes personal data sits in locations you can enumerate and query. Prompts, RAG chunks, cached completions, and fine-tuning corpora don’t sit still the way a customers table does. A few places this bites:
- Erasure gets harder, not impossible. Section 12 gives Data Principals a right to erasure. Deleting a row is easy. Finding and removing every embedding derived from that row, every cached response that quoted it, and every fine-tuning example that included it is an engineering problem most teams haven’t scoped - and if the data made it into model weights through fine-tuning, “erasure” in the strict sense may not be achievable at all without retraining.
- Purpose limitation applies to training, not just serving. Consent captured for “help me with my order” doesn’t automatically cover “use this conversation to fine-tune our next model.” Section 6 requires consent limited to the specified purpose - repurposing support logs as training data is a new purpose that needs its own basis.
- Access logging extends to your AI stack. Rule 6 requires monitoring and access logs for personal data processing, retained for at least a year. If your vector DB, prompt cache, and inference logs aren’t in that logging scope, your security-safeguard obligations have a gap most teams don’t know exists until an incident forces the question.
What to actually build
- Treat prompt/completion logs as personal data by default - same retention and access-control rules as your primary database, not a separate “just logs” exception.
- Map every place a user’s data can land in your AI pipeline: input, retrieved context, cached output, eval set, fine-tuning corpus - not just the request/response pair.
- Design erasure to reach the vector store, not just the row it was derived from. If you can’t currently delete a specific customer’s embeddings without a full re-index, that’s worth knowing before someone files a request.
- Get purpose-specific consent before repurposing conversational data for training - and keep that consent separately trackable from your general product consent.
- Extend your access-logging and breach-detection tooling to cover the AI stack specifically, not just the application database.
None of this is exotic. It’s the same discipline DPDP already expects of any data pipeline - it just hasn’t caught up to where AI teams’ data actually flows yet.
VANGUARD inspects every prompt, retrieved chunk, and response in real time, giving you the access logs and audit trail DPDP’s security-safeguard obligations require - across your AI stack, not just your application database. Run a free runtime risk assessment to see where your pipeline currently stands.



