Skip to content
DPDP ActLocalisationConsent managementPrivacy engineeringNotices

Consent in 22 Languages Is a Versioning Problem

The DPDP Act's language duty sits in Sections 5(3) and 6(3), not the 2025 Rules — and 22 scheduled languages means 11 scripts and a notice-version ledger.

Mukul Verma Founder, ConsentEra 9 min read

We built multilingual consent notices twice. The first build was a translation pipeline, and it was wrong — because Section 6(10) of the DPDP Act asks a question no translation pipeline can answer: which language version was this person shown, on which date, and what did it say, word for word. Rebuilding it as a versioned ledger cost a schema change, a migration, and the discovery that the worst failures here do not error at all.

The obligation is in the Act. It is not in the Rules.

Much of the compliance material circulating in India has this backwards. The Digital Personal Data Protection Rules, 2025 do not contain the language obligation at all. I ran a full-text search over the official 41-page MeitY gazette, G.S.R. 846(E): zero occurrences of “Eighth Schedule”, zero of “English”, zero of “translation”. The only language words in Rule 3 are “clear and plain language”.

The duty sits in the parent statute, twice. Section 5(3) says the Data Fiduciary “shall give the Data Principal the option to access the contents of the notice … in English or any language specified in the Eighth Schedule to the Constitution.” Section 6(3) does it again for the consent request itself, which must be presented “in a clear and plain language, giving her the option to access such request” in English or a scheduled language.

Two artefacts, separately regulated. A localisation ticket that covers the notice body but not the consent screen’s own purpose labels and toggle text has covered half the obligation.

The second correction matters more for scoping. The words are “the option to access”, not “shall publish in all” — an availability obligation, choice sitting on the principal’s side. Overstating that is the commonest error in the market; understating it is a mistake too, for reasons in the next section.

Third: these are the 22 scheduled languages of the Eighth Schedule, not “official languages” and not “national languages”. India has no national language. With English, the maximum fan-out is 23 legally distinct artefacts.

The clock: the Rules were notified on 13 November 2025 and commence in tranches. Consent Manager registration under Rule 4 lands on 13 November 2026, about three and a half months away; Rule 3 and the rest of the substantive regime on 13 May 2027. Nothing in the notice regime binds anyone today, which is exactly why the ledger is worth building now: it is the one component you cannot retrofit, because it must already have been running when the consent was collected.

870 million people are going to exercise the option

IAMAI and Kantar’s ICUBE 2024 counts 886 million active internet users in India, of whom 870 million — 98 per cent — used it in Indic languages during 2024. Rural is the larger half, 488 million against 397 million urban, and 57 per cent of urban users say they prefer Indic-language content.

The tail matters more. The same study splits urban preference as English 43 per cent, Hindi 24, Tamil 6, Telugu 4, Marathi 3, Gujarati 3, Malayalam 3, Bengali 2, Kannada 2 — and “Others” at 10 per cent. That bucket beats every named language except English and Hindi, and it is where the low-resource scheduled languages sit.

A correction on the figures everyone still quotes: the 234 million Indian-language users against 175 million English, and the 536 million projection, are from the KPMG and Google study of April 2017, and it covered eight languages. A trend line, not a current measurement — and no demand research exists for the other fourteen. Sanskrit is on the list, and an RTI reply put 24,821 people declaring it as mother tongue in the 2011 Census. The obligation is not proportional to demand.

22 languages is 11 scripts, and three of them are right-to-left

Twenty-two string files is the wrong mental model. Per the Unicode CLDR coverage data, Devanagari is the default script for eight scheduled languages (Hindi, Marathi, Nepali, Konkani, Sanskrit, Bodo, Dogri, Maithili), the Bengali script for three (Bengali, Assamese, Manipuri), and Perso-Arabic for three — Urdu, Sindhi and Kashmiri. Eight more scripts carry one language each: eleven writing systems in total. Right-to-left is therefore not the single Urdu edge case people budget for — it is three of the twenty-two, and it moves bidirectional handling, field alignment and the accept control.

Then the trap that cost us a migration: one language does not map to one locale. CLDR carries ks against ks_Deva, sd against sd_Deva, mni in the Bengali script against mni_Mtei in Meetei Mayek, sat in Ol Chiki against sat_Deva, and kok against kok_Latn. A Kashmiri notice in Perso-Arabic and one in Devanagari are not the same document and cannot share a primary key. Our first cut keyed on language alone.

Where the i18n stack quietly runs out of data

CLDR sets a “modern” target coverage level for only 15 of the 22 scheduled languages. The other seven sit at “basic”: Bodo at 46.5 per cent modern-level completeness, Maithili 37.5, Kashmiri 36.6, Dogri 16.0, Sanskrit 10.7, Manipuri 9.4, Santali 9.2.

Two locales are worse than thin. Manipuri in Meetei Mayek sits at 0.5 per cent and Santali in Devanagari at 0.4, and neither ships in ICU at all. Dates, numbers and plural rules fall back silently. That is the dangerous word: the glyphs render, the page looks finished, and a retention date inside a Meetei Mayek notice comes out in an English pattern with nothing logged anywhere.

The mainstream Indic scripts were encoded in Unicode 1.1 in 1993; Ol Chiki only in 5.1 in 2008 and Meetei Mayek in 5.2 in 2009. Noto covers both and Android has shipped Ol Chiki since 5.0 in 2014, on about 95 per cent of Indian handsets. Rendering is largely solved. The locale data is not, and no translation vendor can sell you that gap closed.

Nobody has said which version wins

The Act designates no authoritative language version. Indian law knows how to do the opposite: Article 348(1)(b) makes the authoritative texts of all Acts of Parliament and State Legislatures English. The DPDP Act does not replicate that for notices — a contrast to draw, not a rule to import.

So each variant binds whoever saw it. A mistranslated purpose is not a copy bug — consent attaches to the purpose as described in the notice served, so a purpose that reads wider in Bengali than in English is a consent defect in Bengali.

No Indian authority on the equivalence of translated consent exists yet, and cannot until the Board’s jurisdiction begins on 13 May 2027. The two closest published standards are analogies. The Article 29 Working Party’s transparency guidelines require that the “phraseology and syntax makes sense in the second language(s) so that the translated text does not have to be deciphered or re-interpreted”. The RBI’s Key Facts Statement circular of 15 April 2024 is closer to Indian practice: the KFS “shall be written in a language understood by such borrowers”, explained, with an acknowledgement of understanding obtained. Neither binds a Data Fiduciary here. Both set comprehension, not string parity, as the bar.

The question an inspection actually asks

Section 6(10) is the sentence that reorganised our architecture: where a question arises in a proceeding, the fiduciary “shall be obliged to prove that a notice was given … and consent was given … in accordance with the provisions of this Act”.

In practice: proving that on 3 March 2027 this Data Principal was served this Marathi text, equivalent to the English core in force that day. Not that a Marathi notice existed. That one did.

The retention arithmetic makes it harder. A registered Consent Manager must keep notices as records for at least seven years under the First Schedule, Part B — a duty that binds Consent Managers, not every Data Fiduciary. A Data Fiduciary’s own floor under Rule 6(1)(e) is one year for logs and personal data. So the log recording which version was rendered can expire six years before the record it explains. We set notice-version retention against the longest consent lifetime plus a limitation buffer, never against the log floor — and the Third Schedule’s three-year erasure clocks for e-commerce and social media platforms above two crore registered users run from the principal’s last approach, so a superseded version stays load-bearing long after the copy team has forgotten it exists.

And the exposure is not ₹250 crore. That entry attaches only to breach of the Section 8(5) security-safeguards obligation. A Section 5 or 6 language failure falls under entry 7, “breach of any other provision”, which “may extend to fifty crore rupees”.

The schema we ended up with

The core of it is unglamorous.

notice_version
  notice_id     uuid
  version       int
  language      text        -- BCP-47 subtag: hi, ur, sat
  script        text        -- ISO 15924: Deva, Arab, Olck, Mtei
  core_hash     bytea       -- SHA-256 over the canonicalised operative core
  render_hash   bytea       -- SHA-256 over the canonicalised rendered body
  status        text        -- draft|in_review|approved|published|superseded
  published_at  timestamptz
  supersedes    uuid
  PRIMARY KEY (notice_id, version, language, script)

Four rules hold it together. Versions are immutable and content-addressed, so a correction is a new row and never an UPDATE. The consent event stores render_hash and the rendered bytes, plus the language and script actually served — read from the render, not from the locale the browser asked for. A change to the legally-operative core (purposes, data categories, retention, rights, recipients) opens a new version in every language; a change to explanatory chrome does not. And publication is gated: a purpose cannot go live in one language ahead of the others without recording a fallback event naming what the principal got instead.

India has a template for this shape already — the DEPA consent artefact mandates a unique id, schema-name, schema-version, content-type, date-time and expiry-time in its headers, signed with JWS.

Three design failures taught us those rules, and all three were ours. The first was storing a pointer to a CMS page rather than a snapshot of the rendered text: edit the Marathi copy and every consent record pointing at it silently changes what it claims to have displayed, with nothing left to reconstruct the original. It does not error, and no alert fires — the only way we found it was going looking. The second was the missing script subtag, above.

The third was translation lag. Publish an English purpose the moment it clears review and the Tamil version follows whenever its own review finishes — and in the gap, one language is serving old text against a new purpose id. Every component behaves correctly on its own; the system is still wrong. That gap is why publication is now gated on the whole language set, not on the first version ready.

What we have not solved

Equivalence review at the tail. There is no reliable supply of reviewers fluent in Santali or Dogri and competent in Indian privacy law, at any price we have found. We serve those languages on request with a recorded review turnaround and a pending-review state on the version. It is honest and it is not good enough.

Treating English as source of truth is our policy, not the law’s. The statute is silent, so if a Bengali version and the English diverge and a principal relied on the Bengali, we expect the Bengali to govern for that principal. We document that position rather than pretend the Act settled it.

We can prove what we rendered. We cannot prove it was understood, and RBI’s acknowledgement model — the only published operational answer — does not survive a consent screen that must resolve in a few hundred milliseconds.

The review fan-out is arithmetic: languages times purposes times every copy change. Splitting the operative core from the chrome cut re-review volume by roughly 70 per cent on our own notices, but where that line falls is a judgement call, and drawing it wrong means under-reviewing something that mattered.

Rule 3 commences on 13 May 2027, a little over nine months out. The translation is the cheap part. The ledger that can answer, years later, which of 23 versions a person actually saw is the part worth starting now, and it is where most of our engineering time goes.

Common questions

Do we have to translate every notice into all 22 scheduled languages?

The Act gives the Data Principal the option to access the notice in English or any language in the Eighth Schedule. How you meet that option is an operational decision, and most organisations scope it to the languages their users actually use plus a route to produce others on request. What you cannot do is treat the option as satisfied by an English-only notice.

Which language version governs if the translations diverge?

Nothing in the Act or the Rules designates an authoritative version, which is precisely the problem. Until that is settled, the defensible position is to treat each rendered language version as its own immutable artefact, hash it, and record which one the person was actually shown — so the question of divergence is answered with evidence rather than argument.

Is machine translation enough for a consent notice?

It is a starting point, not an answer. A notice is a legal artefact whose purpose descriptions have to mean the same thing in every language, and equivalence needs review by someone who reads the language. The engineering consequence is worse than the cost: every purpose change fans out across twenty-three versions, each of which needs its own review before it can be published.

What do we need to store to prove what someone saw?

The version, not the string. Keep notice versions immutable, hash the canonicalised rendered artefact rather than the template, and have the consent record carry both the notice version id and its own copy of that hash. Then drift between what was shown and what is stored is detectable instead of silent.

Sources

  1. Digital Personal Data Protection Act, 2023 (No. 22 of 2023) — India Code, Government of India
  2. Digital Personal Data Protection Rules, 2025 — G.S.R. 846(E), MeitY
  3. Languages included in the Eighth Schedule — Department of Official Language, Ministry of Home Affairs
  4. Unicode CLDR — Locale Coverage chart
  5. Unicode Consortium — Supported Scripts
  6. IAMAI & Kantar — Internet in India 2024 (ICUBE 2024)
  7. KPMG in India & Google — Indian Languages: Defining India's Internet (April 2017)
  8. Constitution of India — Article 348
  9. Article 29 Working Party / EDPB — Guidelines on Transparency (WP260 rev.01)
  10. Reserve Bank of India — Key Facts Statement for Loans & Advances, 15 April 2024
  11. DEPA — Consent Artefact specification
  12. ThePrint — Only 24,821 Indians identified as Sanskrit speakers in 2011 Census (RTI)
  13. Typotheque — Tracing Meitei Mayek & Ol Chiki Letters

Mukul Verma

Founder, ConsentEra

Mukul builds ConsentEra's consent and data-protection platform. He writes about what the DPDP framework actually demands of an engineering team — the parts that turn out to be hard once you try to ship them.

See your consent flow,
sealed and verifiable.

A 30-minute walkthrough mapped to your data, your notices, and your DPDP readiness deadline.