India's DPDP Rules are in force, the Data Protection Board is operational, and a January 2026 consultation proposed halving the compliance runway. For AI teams, the hard question is no longer whether you collected consent — it is whether you can prove that consent covered model training.

The runway just got shorter. Probably.

The Digital Personal Data Protection Rules, 2025 were notified in November 2025 under Gazette G.S.R. 846(E), with commencement dates for the DPDP Act set out in G.S.R. 843(E). Enforcement is phased: provisions establishing the Data Protection Board took effect immediately, the consent manager framework lands twelve months in — November 2026 — and the full set of substantive obligations becomes enforceable in mid-May 2027.

That was the plan. In January 2026, a MeitY stakeholder consultation proposed compressing the eighteen-month window to twelve. The proposal has not been gazetted, and it may never be. But compliance advisors are already telling enterprises to treat November 2026 as the prudent planning baseline, because a typical enterprise programme needs nine to twelve months to go from gap assessment to audit readiness. If your roadmap is pencilled against May 2027 and the acceleration lands, you will be discovering the gap at exactly the moment the Board starts asking for evidence.

₹650 crore, not ₹250 crore

Most coverage quotes the headline figure: up to ₹250 crore for failing to implement reasonable security safeguards. That undersells the exposure, because the Data Protection Board can impose penalties per violation, and a single incident routinely produces several.

One breach of personal data used in an AI training pipeline can stack three failures: inadequate security safeguards (up to ₹250 crore), failure to notify affected data principals (up to ₹200 crore), and failure to notify the Board (up to ₹200 crore). Cumulative ceiling: ₹650 crore from one bad week. The Board determines actual amounts after investigation, but the arithmetic is the point. AI compliance in India is no longer a policy exercise; it is a balance-sheet risk.

Why training data is the exposed flank

Operational systems like your CRM and support desk will be cleaned up first, because that is where consent tooling naturally plugs in. Training datasets are different. They are copied, merged, relabelled, versioned and shipped to vendors, and the consent that authorised the original collection rarely follows the record through any of it.

The DPDP framework is strict about purpose limitation: personal data may be processed only for the specified purpose to which the data principal consented and retained only while that purpose is being served. Consent collected to deliver a service does not silently extend to training a model. If a radiology platform collected scans to provide diagnostics, using those scans to train the next model version is a separate purpose, one that needs its own notice, its own consent record, and its own retention logic.

Withdrawal makes it harder. When a data principal withdraws consent, the obligation does not stop at the production database. Teams need to know which dataset versions, which annotation batches, and which model checkpoints that record touched. Few can answer that today.

There is a second exposure aimed squarely at AI builders. Organisations designated as Significant Data Fiduciaries must undertake due diligence to verify that their technical measures, explicitly including algorithmic software, do not pose a risk to data principals' rights, and face restrictions on transferring certain data and traffic data outside India. If your platform processes health, financial or other sensitive personal data at scale, assume designation is plausible and that your models themselves, not just your databases, will be inside the audit perimeter.

What "proving consent" looks like in practice

When the Board investigates, a privacy policy is not evidence. Evidence is record-level and time-stamped. For training data, that means:

  • Consent lineage per record. Every datum in a training set maps back to a consent artefact naming model training as a purpose, not a generic "service improvement" clause.
  • Notice records. Proof of what the data principal was actually shown, in which language, at the point of collection.
  • Activity logs retained for one year. Rule 6 mandates log retention; for AI teams that should extend to dataset access, export and labelling events.
  • Rights-response machinery. Rule 14(3) gives you ninety days to respond to a data principal's request — which presumes you can locate their data across dataset versions at all.
  • Withdrawal propagation. A documented, repeatable process for tracing a withdrawn record through training corpora and deciding what happens to models already trained on it.

The vendor handoff is where evidence dies

Most Indian AI teams do not label data alone. The moment a dataset leaves your environment for a data annotation company, you are managing a Data Processor relationship — and under DPDP, the fiduciary's liability does not transfer with the files. Your consent evidence has to survive the handoff intact.

That changes what vendor diligence means. A data processing agreement is the floor, not the standard. The questions that matter: Does consent and purpose metadata travel with each record into the labelling environment? Is every annotation event logged against an identifiable record? Can the vendor demonstrate SOC 2 Type II controls and ISO 27001:2022 certification rather than merely claim them? Can they hand you an audit trail the Board could read, not just a delivery folder of labelled files?

If your annotation partner cannot answer those questions, your compliance programme has a hole exactly where your most sensitive processing happens.

If you've done GDPR or the EU AI Act, you're closer than you think

Teams that built lawful-basis documentation for GDPR AI training data will recognise the shape of this: purpose specification, records of processing, demonstrable accountability. DPDP's consent-centric design is stricter in places — legitimate interest is not the escape hatch it is in Europe — but the evidence discipline transfers.

The same is true of the EU AI Act. High-risk systems face Annex IV technical documentation obligations, now enforceable from 2 December 2027 following the Digital Omnibus. Teams selling into both markets should build one evidence layer, not two: dataset provenance, consent lineage and annotation audit trails serve Brussels and New Delhi alike. As we have said before about staggered regulatory timelines — you can defer an obligation; you cannot backdate evidence.

What to do before November 2026

  • Inventory training datasets containing personal data of Indian data principals, including derived and vendor-held copies.
  • Map each dataset to its consent basis. Flag every set where the stated purpose does not explicitly cover model training.
  • Stand up record-level logging for dataset access, export and annotation, with one-year retention.
  • Re-paper vendor relationships so consent metadata and audit obligations are contractual, testable and inspected.
  • Rehearse a withdrawal. Pick one record, trace it through your pipeline, and time how long the answer takes. That number is your real readiness metric.
Penalty stack diagram: one breach in an AI training pipeline can combine inadequate security safeguards (up to 250 crore rupees), failure to notify data principals (200 crore) and failure to notify the Board (200 crore), for a cumulative ceiling of 650 crore rupees.
Figure 2. Penalties are imposed per violation, so one incident stacks. A single training-pipeline breach can combine three failures — security (₹250 cr), principal notification (₹200 cr), Board notification (₹200 cr) — for a ₹650 crore ceiling. The Board sets actual amounts.

Compliance posture

Five frameworks. One annotation backbone that passes Legal, Security, and Procurement.

ISO 27001:2022
CERTIFIED
SOC 2
CERTIFIED
HIPAA
COMPLIANT
GDPR
COMPLIANT
DPDP
READY
The five record-level evidence requirements for proving training-data consent: consent lineage per record, notice records, activity logs retained one year, rights-response machinery within ninety days, and withdrawal propagation through training corpora.
Figure 3. A privacy policy is not evidence. When the Board investigates, evidence is record-level and time-stamped: consent lineage, notice records, one-year activity logs (Rule 6), rights-response within ninety days (Rule 14(3)), and a repeatable withdrawal-propagation process.

Next step

See consent-aware annotation in practice.

LabelFort attaches consent and purpose metadata to every record, logs every annotation event, and produces audit trails built for regulators — DPDP, GDPR and EU AI Act alike. Bring one of your datasets and we will show you what evidence-grade labelling looks like on your own data.