SilencioNetwork/indic-languages-speech
Indic Spontaneous Speech — Silencio Spontaneous speech in 7 Indic languages from 161 speakers, one clip each. Hindi, Urdu, Bengali, Marathi, Nepali, Sindhi and Gujarati, recorded by speakers born in India, Pakistan, Bangladesh, Nepal and the diaspora, and labelled with 34 self-reported regional varieties. Audio and speaker metadata only. Human-validated transcription is available on request. Hours 2.00 Clips 161 Speakers 161 (one clip each) Languages 7… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/indic-languages-speech.
Indic Spontaneous Speech — Silencio
Spontaneous speech in 7 Indic languages from 161 speakers, one clip each. Hindi, Urdu, Bengali, Marathi, Nepali, Sindhi and Gujarati, recorded by speakers born in India, Pakistan, Bangladesh, Nepal and the diaspora, and labelled with 34 self-reported regional varieties. Audio and speaker metadata only. Human-validated transcription is available on request.
Hindi, Urdu and Bengali together have well over a billion speakers, and public speech data for them is dominated by read sentences from a small number of speakers. This release is the opposite shape: one clip per speaker across 161 speakers, all of it unscripted, with regional variety, declared language profile, age band, gender and recording conditions attached to every clip. Hindi and Urdu appear side by side, which is useful because they are close in speech and usually separated in datasets.
Contributors answer an open prompt in their own words, on their own devices, wherever they are. This is spontaneous speech, not read sentences.
This is a sample, not the catalogue. It is drawn to show format, audio quality and metadata, and its speaker mix does not represent what Silencio holds. The off-the-shelf catalogue behind it is far more diverse across regions, age groups, accents and recording conditions. Men are 77% of the speakers in this sample. Across the catalogue the gender split is roughly 40% female to 60% male, with men in the majority in most languages. Licensed subsets can be drawn to a specified gender, age or regional balance.
Load it
from datasets import load_dataset
# everything (default)
ds = load_dataset("SilencioNetwork/indic-languages-speech", split="test")
print(ds[0]["language"], ds[0]["dialect"], ds[0]["country"], ds[0]["proficiency"])
# or one language at a time: "hindi", "urdu", "bengali", "other"
hindi = load_dataset("SilencioNetwork/indic-languages-speech", "hindi", split="test")
# datasets v4 returns a torchcodec AudioDecoder:
sample = ds[0]["audio"].get_all_samples()
audio, sr = sample.data, sample.sample_rateRequires pip install "datasets>=4.0" and FFmpeg ≥ 4.
What is in it
By language
By country of birth
Regional variety — self-reported speaker background, not a classification of the recorded speech
Demographics
Recording conditions
Language subsets
The release ships as named subsets, so you can pull one language without downloading the rest. The preview at the top of this page has a dropdown for them.
Splits
Single split, test, in every subset; 161 rows in all. No train/dev/test partition is provided: at this scale a partition would leave each part too small to be meaningful. Every clip is from a different speaker, so any split you construct is speaker-disjoint by construction, with no speaker leakage to control for.
Fields
How this sample was drawn
The export behind this release held 5.1 hours. This listing is a curated 2-hour subset, chosen to show the widest spread of languages, regional varieties and recording conditions rather than the largest number of hours:
- Clips between 20 and 150 seconds, so no single speaker dominates the release.
- Selection favoured unseen languages and unseen regional varieties first, then audio quality.
- Recordings that failed the audio screen, stopped at the five-minute recording limit, ran under 15 seconds, or whose speaker declared no proficiency in the recorded language were left out and not substituted.
- Every clip was checked for level, clipping, silence and signal-to-noise. The lowest SNR estimate in the release is 20 dB. Audio is shipped at source quality, not resampled or denoised.
What this is for
- Spontaneous-speech evaluation across South Asia. Real speech with hesitations, restarts and English code-mixing, in languages whose public data is mostly read.
- Hindi–Urdu discrimination, with both in one release under the same protocol.
- Regional variety and accent classification. 34 self-reported varieties across Delhi, Lucknow, Jaipur, Bihar, Kolkata, Dhaka, Rajshahi, Chittagong, Lahore, Karachi, Islamabad and more.
- Speaker identification and verification. One clip per speaker, so splits are speaker-disjoint by construction.
- Robustness and fairness auditing by language, region, age, gender and capture condition.
- ASR evaluation once reference text is added. Human-validated transcription is available on request over this sample or any larger subset.
Limitations
- Small, and a sample. 2 hours is for evaluation and for checking the format, not for training.
- Gender skew. 124 of 161 speakers are men. Balanced subsets can be drawn to order.
- Uneven language counts. Hindi, Urdu and Bengali carry almost all of the release; Marathi, Nepali, Sindhi and Gujarati appear with one or two speakers each and are indicative only.
- No transcripts in this release. See
transcriptabove. - Self-reported labels. Language, regional variety, first language and proficiency are declared by contributors, not assessed.
- Mixed channel layout. Audio is shipped as recorded; 2 clips are stereo with one silent channel. Downmix to mono before analysis.
Provenance and consent
Every recording is contributed by an opted-in participant through the Silencio platform, under a consent record covering AI/ML training use. Contributors can request deletion, and deletion propagates to downstream releases. Full provenance documentation is available to licensees.
Speaker and recording identifiers in this release are pseudonymised afresh, so they do not match Silencio's internal IDs or those in any other Silencio release.
License
cc-by-nc-4.0: free for research and non-commercial use with attribution.
Attribution string: Silencio Network, Indic Spontaneous Speech, 2026. CC BY-NC 4.0.
Non-commercial covers research, evaluation and publication. Benchmarking a commercial product model against this data is a commercial use and needs a licence. Ask; it is usually granted for evaluation. Model weights trained on this sample inherit the non-commercial restriction.
Commercial licensing, including terms for models trained on this data: [silencio.network/contact](https://www.silencio.network/contact)
Citation
@misc{silencio_indic_2026,
title = {Indic Spontaneous Speech — Silencio},
author = {Silencio Network},
year = {2026},
url = {https://huggingface.co/datasets/SilencioNetwork/indic-languages-speech}
}Indic languages off the shelf
This release is a 2-hour sample. Catalogue depth for the languages in it, September 2026, published in bands rather than exact hour counts:
South Asian English is also a large holding: India is one of the biggest single origins in the English catalogue, alongside Pakistan, Bangladesh, Nepal and Sri Lanka. See Accents of English.
Available by language, regional variety, demographic profile and recording condition, as spontaneous or read speech. Human-validated transcription with word-level alignment is available over any subset, scoped to your needs.
Silencio: the catalogue behind this sample
This listing is a small fraction of what Silencio provides. Silencio specialises in underrepresented languages, accents and niche domains, delivered at scale. It draws on the largest global community of speech contributors and holds the largest off-the-shelf catalogue of niche-language speech available.
Recorded and available off the shelf. Audio already collected, with metadata, licensable today. Catalogue position as of September 2026:
Contributor community available on demand. Registered, consented contributors who can be activated for a specific brief:
Anything not off the shelf can be sourced through the community. A language, an accent or dialect region, a demographic, a recording condition, a speech style or a specialist domain can be collected to a client's brief. Deep coverage across Africa, South-East Asia, South Asia and the Middle East.
Human-validated transcription is available for every language, scoped to each client's needs: script and orthography conventions, normalisation rules, word- or segment-level alignment, speaker labelling and turnaround. Every delivered clip is checked by a native-speaker reviewer. Published samples with human-validated transcripts: Kenyan Swahili, Cebuano, Tagalog / Filipino, Yoruba, Hausa and Amharic.
Proprietary and first-party. Every recording is collected directly by Silencio from consenting contributors. Nothing is scraped, and these recordings are not available in any other dataset on the internet.
Ethical sourcing, with provenance records. Every recording carries a consent record covering AI/ML training use, and contributor-level provenance documentation is available to licensees, including for EU AI Act training-data summaries. Contributors can withdraw consent, and withdrawal propagates to subsequent releases.
More at silencio.network.
For volume licensing, bespoke transcription or commissioned collection: hello@silencio.network
