The Practice

An organisation's data science or people analytics team builds an internal AI model to predict employee attrition risk. To train the model, they use three years of historical HR data — including performance ratings, leave records, salary progression, manager feedback notes, and disciplinary records — of current and former employees. No employee was informed of this use, and the data was originally collected for performance management and payroll purposes.

 

Questions Raised for Compliance Review

  1. Is the use of historical employee personal data for AI model training a separate processing activity requiring a distinct lawful basis under DPDP?
  2. What risks arise from the use of sensitive employment data in a predictive AI system, particularly one that generates individual risk scores?
  3. What is the compliant approach to building people analytics AI within the boundaries of DPDP?
     

Is This Permitted Under DPDP?

Yes — this is a separate and distinct processing activity requiring its own lawful basis. The data was collected for performance management, payroll, and operational HR purposes. Using it to train a predictive AI model is a materially different purpose — one the employees had no knowledge of, did not consent to, and which generates inferences about them that they have no opportunity to review or contest. This is a clear purpose limitation violation under DPDP Section 5.

 

Where the Breach Risks Sit

  • Purpose limitation breach — Performance ratings were collected to support appraisal decisions. Leave records were collected for payroll and absence management. Salary progression data was collected for compensation administration. None of these purposes extends, by implication or inference, to the training of a machine learning model that scores each employee's likelihood of leaving the organisation.
  • Sensitive data in model training — Disciplinary records, medical leave patterns, and manager feedback notes are among the most sensitive categories of employment data. Their inclusion in a training dataset — even for an internal model — creates a record of sensitive processing that has no declared basis and cannot be audited against a consent framework.
  • Automated profiling without disclosure — The model generates an attrition risk score for each employee. This is automated profiling — a processing activity DPDP requires be disclosed to the data principal. Employees who are quietly scored as high attrition risks, and who subsequently receive different treatment as a result, have no knowledge that this is occurring and no mechanism to challenge it.
  • Inclusion of former employee data — Former employees whose data is included in the training set have no active employment relationship with the organisation. Their data is being used for a purpose that arose after their departure, without any mechanism to notify them or obtain consent.

 

The Ideal Compliant Approach

  1. Conduct a Data Protection Impact Assessment before building the model. Any AI system that profiles employees and generates individual risk scores is a high-risk processing activity under DPDP. A DPIA must be completed — assessing the purpose, the data used, the risks to individuals, and the mitigations in place — before the model is built or deployed.
  2. Establish a lawful basis before using existing data. The organisation must either obtain informed consent from employees for this specific use — which requires transparently explaining what the model does and how scores are used — or identify another lawful basis under DPDP. Retrospective reclassification of existing data is not a recognised approach.
  3. Disclose the existence of the model to employees. If the model is deployed, employees must be informed through the organisation's privacy notice or a specific HR communication that a people analytics model is in use, what data it draws on, and how outputs influence HR decisions.
  4. Build in human review for all model-influenced decisions. Attrition risk scores must not be the sole basis for any HR decision — whether a promotion, a development opportunity, or a retention intervention. A human reviewer must assess the score in context, and employees must have a pathway to understand and contest any decision that affects them.

 

DPDP Risk Summary

ElementStatusRecommended Action
Historical HR data used for AI trainingPurpose limitation violationEstablish new lawful basis; obtain employee consent for this use
Employees not informed of model existenceTransparency obligation not metUpdate employee privacy notice to disclose people analytics use
Sensitive data (disciplinary, medical leave) in modelHigh-risk processing, no declared basisRemove sensitive categories or obtain explicit consent for inclusion
Former employee data includedNo ongoing lawful basis post-employmentExclude former employees or obtain separate consent
No DPIA conductedRequired for high-risk AI processingComplete DPIA before model training or deployment
Automated profiling with no human reviewNon-compliant automated decision-makingMandate human review for all model-influenced HR decisions