Models deployed for Indian healthcare and agriculture were trained on text in which English is about 46% and Telugu is a rounding error. The usual fix is to collect more data and fine-tune, which means somebody ships raw health records from a village in Telangana to a server. Nobody in that village got to weigh that trade-off.
So the collection node I published on 30 August has one rule: raw text stays on the device it was typed into. What moves upstream is scores, counts and, later, gradient updates. The node runs at a partner site, an ASHA sub-centre, an agricultural extension office or a legal aid clinic, in four languages (Telugu, Tamil, Hindi, Bengali) and four domains (health, agriculture, legal, education).
Four guards, because one is easy to delete
- The record type that goes on the wire has a
raw_textproperty that returns nothing and raises if anything tries to set it. - Raw text lives in a separate SQLite table, not a column, so reading it is always a deliberate join you can spot in a diff.
- The Avro schema types
raw_textas null, so both encoders throw if handed a value. That turned out to be the strongest guard, and it was free. - Before queueing, and again before the socket, the node searches the finished bytes for the raw capture.
Built for a Pi on 2G
The target is a Raspberry Pi 4 or a mid-range Android box with 4 GB, on a 2G link. Nothing heavy loads at import. The default embedder is a character n-gram hasher with no weights, so a node can boot and verify before anyone has copied 500 MB onto its SD card over GPRS. Sync batches are sized by compressed bytes rather than record count, because a one-line field note and an hour of classroom transcript differ in size by a factor of 100. There are 188 tests.
What I only noticed late
The privacy term in the quality score works as a gate. A record still waiting for human review can never reach the corpus, so the corpus grows at the speed of reviewer hours, not collection. That is the right behaviour, and it is also the real bottleneck.