Indic Voice Collection Where Each Language Gets Its Own Design
TLDR If you ship voice for India, a language label is not coverage. You need speakers who match your users, sessions that match your product, and transcripts you can train on without cleanup.
At posterior.xyz we collect speech in English, Hindi, Kannada, Telugu, Tamil, Malayalam, Marathi, and Sanskrit, plus more on request. Each language gets its own plan. That is the whole point.
Why one plan never fits all Indian languages
Hindi code-switching in Delhi sounds nothing like Tamil spoken in Madurai. Reading fluency in Mumbai looks nothing like spontaneous speech from a first-time smartphone user in Kochi.
So we never copy a speaker grid from one language to another. You tell us who uses your product. We build a separate collection design for each language around that reality.
Speaker matrices you approve before recording
Before we record anything, you review a speaker matrix per language.
Typical dimensions include
- Region and mother tongue influence within the language
- Age bands and gender balance tied to your user base
- Device familiarity
- Reading fluency versus spontaneous speaking style
- Recording setting, with studio control or field realism or a mix of both
You sign off on the matrix. Then recruiting starts. What ships later traces back to that sheet, so your eval reflects deployment instead of convenience sampling.
This part is kinda unglamorous and it decides everything. Get the matrix right and the model has a chance. Get it wrong and no amount of tuning saves you.
Task designs that match your deployment
A voice assistant breaks in different places than an IVR flow. We shape prompts and sessions around your failure points.
| Task type | What we record | Good for |
|---|---|---|
| Scripted prompts | Phonetically balanced sentences and controlled vocabulary | ASR coverage and pronunciation |
| Conversational tasks | Free talk with turn taking, pauses, repairs, and interruptions | Assistants and agents that hold a conversation |
| Channel separated recordings | Multi speaker sessions with clean per speaker tracks | Diarization and meeting transcription |
| Wake word and commands | Varied phrasing at different distances and noise levels | Detection with fewer false accepts |
| IVR style flows | Menu paths with confirmations, corrections, and retries | Call automation that survives real callers |
Each session ships with structure notes on mic setup, session length, and prompt order, so you can split train and test the way you want.
Translation that keeps tone intact
Parallel prompts across languages need more than literal word swaps. A polite request in Marathi can sound flat or abrupt when moved word for word into Hindi, and your model picks up the wrong habit.
You send source prompts with desired formality. Native reviewers adapt phrasing so politeness, urgency, and natural phrasing carry over. Where reviewers disagree, we flag the choice for you instead of smoothing it over. You receive source text and adapted text side by side, so you can audit every call.
Transcription rules you set and version
Transcript noise looks like model error. We fix that by agreeing on rules before annotation starts.
You set conventions for
- Code switching and borrowed words heard in daily speech
- Fillers, repetitions, and partial words
- Numerals, punctuation, and product terms
- Addresses and named entities
- Unintelligible spans and overlapping speech
We train annotators on your guide, check adherence in sampling rounds, and version the guide. When you change a rule, we reprocess affected files to the new version and note what changed.
Common questions
Which languages do you cover now?
English, Hindi, Kannada, Telugu, Tamil, Malayalam, Marathi, and Sanskrit. We add more on request when you bring the speaker definition.
Can you mix studio and field audio?
Yes. Many teams use studio takes for baseline coverage and field takes for noise tolerance. You set the ratio per language.
Do you support audio plus video?
Yes. The default pipeline is audio only. When your product needs face or gesture context, we add synced video with the same speaker matrix.
How does licensing work?
You get rights scoped to model training and evaluation. We confirm consent, payment, and usage terms per speaker before delivery.
Start with a paid pilot
Send us your speaker matrix and three sample tasks. We record a small set, transcribe to your conventions, and deliver files you can train on. Both sides check fit before scaling.
