Mohammed Mudassir Uddin

BlogAugust 2026dataprivacyIndian languages

A data node where the text never leaves the device

Collecting Telugu, Tamil, Hindi and Bengali text for training without shipping anyone's health records to a server. The rule is one line; enforcing it took four separate guards.

Models deployed for Indian healthcare and agriculture were trained on text in which English is about 46% and Telugu is a rounding error. The usual fix is to collect more data and fine-tune, which means somebody ships raw health records from a village in Telangana to a server. Nobody in that village got to weigh that trade-off.

So the collection node I published on 30 August has one rule: raw text stays on the device it was typed into. What moves upstream is scores, counts and, later, gradient updates. The node runs at a partner site, an ASHA sub-centre, an agricultural extension office or a legal aid clinic, in four languages (Telugu, Tamil, Hindi, Bengali) and four domains (health, agriculture, legal, education).

ON THE DEVICE UPSTREAM Raw text its own table, never sent Five local checks format, language, fit, PII, score Egress check search the outgoing bytes for the raw text, twice 2G Scores and counts later: gradient updates ON THE DEVICE Raw text its own table, never sent Five local checks format, language, fit, PII, score Egress check search the outgoing bytes for the raw text, twice 2G UPSTREAM Scores and counts later: gradient updates
What crosses the boundary. The check runs before a record is queued and again before the socket.

Four guards, because one is easy to delete

  1. The record type that goes on the wire has a raw_text property that returns nothing and raises if anything tries to set it.
  2. Raw text lives in a separate SQLite table, not a column, so reading it is always a deliberate join you can spot in a diff.
  3. The Avro schema types raw_text as null, so both encoders throw if handed a value. That turned out to be the strongest guard, and it was free.
  4. Before queueing, and again before the socket, the node searches the finished bytes for the raw capture.

Built for a Pi on 2G

The target is a Raspberry Pi 4 or a mid-range Android box with 4 GB, on a 2G link. Nothing heavy loads at import. The default embedder is a character n-gram hasher with no weights, so a node can boot and verify before anyone has copied 500 MB onto its SD card over GPRS. Sync batches are sized by compressed bytes rather than record count, because a one-line field note and an hour of classroom transcript differ in size by a factor of 100. There are 188 tests.

What I only noticed late

The privacy term in the quality score works as a gate. A record still waiting for human review can never reach the corpus, so the corpus grows at the speed of reviewer hours, not collection. That is the right behaviour, and it is also the real bottleneck.