CoolFace
Datasetpublic

SilencioNetwork/indic-languages-speech

Indic Spontaneous Speech — Silencio Spontaneous speech in 7 Indic languages from 161 speakers, one clip each. Hindi, Urdu, Bengali, Marathi, Nepali, Sindhi and Gujarati, recorded by speakers born in India, Pakistan, Bangladesh, Nepal and the diaspora, and labelled with 34 self-reported regional varieties. Audio and speaker metadata only. Human-validated transcription is available on request. Hours 2.00 Clips 161 Speakers 161 (one clip each) Languages 7… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/indic-languages-speech.

sourceHugging Facecc-by-nc-4.0updated 1d agoView on Hugging Face
0likes38downloads
Dataset Card

Indic Spontaneous Speech — Silencio

Spontaneous speech in 7 Indic languages from 161 speakers, one clip each. Hindi, Urdu, Bengali, Marathi, Nepali, Sindhi and Gujarati, recorded by speakers born in India, Pakistan, Bangladesh, Nepal and the diaspora, and labelled with 34 self-reported regional varieties. Audio and speaker metadata only. Human-validated transcription is available on request.

Hours2.00
Clips161
Speakers161 (one clip each)
Languages7
Countries of birth8
Regional varieties34
First languages12
Native speakers of the recorded language (declared)149 of 161
Audio48 kHz mono and stereo WAV
Mean clip length44.8 s
TranscriptsHuman-validated transcription available on request
Licencecc-by-nc-4.0

Hindi, Urdu and Bengali together have well over a billion speakers, and public speech data for them is dominated by read sentences from a small number of speakers. This release is the opposite shape: one clip per speaker across 161 speakers, all of it unscripted, with regional variety, declared language profile, age band, gender and recording conditions attached to every clip. Hindi and Urdu appear side by side, which is useful because they are close in speech and usually separated in datasets.

Contributors answer an open prompt in their own words, on their own devices, wherever they are. This is spontaneous speech, not read sentences.

This is a sample, not the catalogue. It is drawn to show format, audio quality and metadata, and its speaker mix does not represent what Silencio holds. The off-the-shelf catalogue behind it is far more diverse across regions, age groups, accents and recording conditions. Men are 77% of the speakers in this sample. Across the catalogue the gender split is roughly 40% female to 60% male, with men in the majority in most languages. Licensed subsets can be drawn to a specified gender, age or regional balance.

Load it

python
from datasets import load_dataset

# everything (default)
ds = load_dataset("SilencioNetwork/indic-languages-speech", split="test")
print(ds[0]["language"], ds[0]["dialect"], ds[0]["country"], ds[0]["proficiency"])

# or one language at a time: "hindi", "urdu", "bengali", "other"
hindi = load_dataset("SilencioNetwork/indic-languages-speech", "hindi", split="test")

# datasets v4 returns a torchcodec AudioDecoder:
sample = ds[0]["audio"].get_all_samples()
audio, sr = sample.data, sample.sample_rate

Requires pip install "datasets>=4.0" and FFmpeg ≥ 4.

What is in it

By language

LanguageClips%
Hindi6137.9%
Urdu5534.2%
Bengali3924.2%
Marathi21.2%
Nepali21.2%
Sindhi10.6%
Gujarati10.6%

By country of birth

Country of birthSpeakers%
India7345.3%
Pakistan5131.7%
Bangladesh3018.6%
Malaysia21.2%
Nepal21.2%
Algeria10.6%
United States10.6%
Afghanistan10.6%

Regional variety — self-reported speaker background, not a classification of the recorded speech

Regional varietySpeakers%
India - Delhi2716.8%
Bangladesh - Dhaka2113.0%
Pakistan - Lahore1710.6%
Pakistan - Islamabad/Rawalpindi169.9%
Pakistan - Karachi127.5%
Bangladesh - Rajshahi95.6%
India - Jaipur Hindi85.0%
India - Bihari-influenced Hindi63.7%
India (West Bengal) - Kolkata63.7%
India - Lucknow (Khariboli)53.1%
Andhra Pradesh - Rayalaseema21.2%
India - South Indian English21.2%
22 further3018.6%

Demographics

GenderSpeakers%
male12477.0%
female3723.0%
Age bandSpeakers%
25-347848.4%
18-245433.5%
35-442113.0%
45-5985.0%

Recording conditions

DeviceClips%
Mobile13684.5%
Desktop2515.5%

Language subsets

The release ships as named subsets, so you can pull one language without downloading the rest. The preview at the top of this page has a dropdown for them.

SubsetContentsClips
all (default)every clip161
hindiHindi61
urduUrdu55
bengaliBengali39
otherMarathi, Nepali, Sindhi, Gujarati6

Splits

Single split, test, in every subset; 161 rows in all. No train/dev/test partition is provided: at this scale a partition would leave each part too small to be meaningful. Every clip is from a different speaker, so any split you construct is speaker-disjoint by construction, with no speaker leakage to control for.

Fields

ColumnDescriptionValues in this release
audioAudio payload, stored at source rate and channel count48 kHz mono and stereo WAV
speaker_idPseudonymous speaker identifier. Coherent within this dataset; deliberately not linkable to other Silencio releases161 distinct
languageLanguage of the recording7 distinct
transcriptEmpty in this releaseHuman-validated transcription available on request
transcript_typeProvenance of the transcriptconstant: none
genderSelf-reportedfemale, male
countrySpeaker's country of birth8 distinct
mother_tongueSpeaker's self-reported first language12 distinct
dialectSelf-reported regional variety of the speaker's own background, NOT a classification of the recorded speech34 distinct
osOperating system of the recording deviceAndroid, Linux, Windows, macOS
deviceRecording device classDesktop, Mobile
durationSeconds156 distinct
script_typeElicitation styleconstant: free_speech
age_bandSelf-reported age, banded18-24, 25-34, 35-44, 45-59
native_speakerTrue where the speaker declares the recorded language as first language or at native level2 distinct
proficiencySpeaker's self-declared proficiency in the recorded languagebasic, conversational, fluent, native

How this sample was drawn

The export behind this release held 5.1 hours. This listing is a curated 2-hour subset, chosen to show the widest spread of languages, regional varieties and recording conditions rather than the largest number of hours:

  • Clips between 20 and 150 seconds, so no single speaker dominates the release.
  • Selection favoured unseen languages and unseen regional varieties first, then audio quality.
  • Recordings that failed the audio screen, stopped at the five-minute recording limit, ran under 15 seconds, or whose speaker declared no proficiency in the recorded language were left out and not substituted.
  • Every clip was checked for level, clipping, silence and signal-to-noise. The lowest SNR estimate in the release is 20 dB. Audio is shipped at source quality, not resampled or denoised.

What this is for

  • Spontaneous-speech evaluation across South Asia. Real speech with hesitations, restarts and English code-mixing, in languages whose public data is mostly read.
  • Hindi–Urdu discrimination, with both in one release under the same protocol.
  • Regional variety and accent classification. 34 self-reported varieties across Delhi, Lucknow, Jaipur, Bihar, Kolkata, Dhaka, Rajshahi, Chittagong, Lahore, Karachi, Islamabad and more.
  • Speaker identification and verification. One clip per speaker, so splits are speaker-disjoint by construction.
  • Robustness and fairness auditing by language, region, age, gender and capture condition.
  • ASR evaluation once reference text is added. Human-validated transcription is available on request over this sample or any larger subset.

Limitations

  • Small, and a sample. 2 hours is for evaluation and for checking the format, not for training.
  • Gender skew. 124 of 161 speakers are men. Balanced subsets can be drawn to order.
  • Uneven language counts. Hindi, Urdu and Bengali carry almost all of the release; Marathi, Nepali, Sindhi and Gujarati appear with one or two speakers each and are indicative only.
  • No transcripts in this release. See transcript above.
  • Self-reported labels. Language, regional variety, first language and proficiency are declared by contributors, not assessed.
  • Mixed channel layout. Audio is shipped as recorded; 2 clips are stereo with one silent channel. Downmix to mono before analysis.

Provenance and consent

Every recording is contributed by an opted-in participant through the Silencio platform, under a consent record covering AI/ML training use. Contributors can request deletion, and deletion propagates to downstream releases. Full provenance documentation is available to licensees.

Speaker and recording identifiers in this release are pseudonymised afresh, so they do not match Silencio's internal IDs or those in any other Silencio release.

License

cc-by-nc-4.0: free for research and non-commercial use with attribution.

Attribution string: Silencio Network, Indic Spontaneous Speech, 2026. CC BY-NC 4.0.

Non-commercial covers research, evaluation and publication. Benchmarking a commercial product model against this data is a commercial use and needs a licence. Ask; it is usually granted for evaluation. Model weights trained on this sample inherit the non-commercial restriction.

Commercial licensing, including terms for models trained on this data: [silencio.network/contact](https://www.silencio.network/contact)

Citation

bibtex
@misc{silencio_indic_2026,
  title  = {Indic Spontaneous Speech — Silencio},
  author = {Silencio Network},
  year   = {2026},
  url    = {https://huggingface.co/datasets/SilencioNetwork/indic-languages-speech}
}

Indic languages off the shelf

This release is a 2-hour sample. Catalogue depth for the languages in it, September 2026, published in bands rather than exact hour counts:

LanguageCatalogue band
UrduTier 1, 10,000 h and above
HindiTier 2, 5,000 to 10,000 h
BengaliTier 3, 1,500 to 5,000 h
Telugu, Marathi, Sindhi, Kannada, OdiaTier 4, 500 to 1,500 h
Further South Asian languagesLess than 500 h, and collected to order

South Asian English is also a large holding: India is one of the biggest single origins in the English catalogue, alongside Pakistan, Bangladesh, Nepal and Sri Lanka. See Accents of English.

Available by language, regional variety, demographic profile and recording condition, as spontaneous or read speech. Human-validated transcription with word-level alignment is available over any subset, scoped to your needs.

Silencio: the catalogue behind this sample

This listing is a small fraction of what Silencio provides. Silencio specialises in underrepresented languages, accents and niche domains, delivered at scale. It draws on the largest global community of speech contributors and holds the largest off-the-shelf catalogue of niche-language speech available.

Recorded and available off the shelf. Audio already collected, with metadata, licensable today. Catalogue position as of September 2026:

Hours, off the shelfaround 500,000
Countries of contributor originaround 180
Catalogue refreshWeekly at record level, quarterly at catalogue level

Contributor community available on demand. Registered, consented contributors who can be activated for a specific brief:

Contributors available on demandaround 2,500,000
Countries180+
Languages that can be collectedaround 350

Anything not off the shelf can be sourced through the community. A language, an accent or dialect region, a demographic, a recording condition, a speech style or a specialist domain can be collected to a client's brief. Deep coverage across Africa, South-East Asia, South Asia and the Middle East.

Human-validated transcription is available for every language, scoped to each client's needs: script and orthography conventions, normalisation rules, word- or segment-level alignment, speaker labelling and turnaround. Every delivered clip is checked by a native-speaker reviewer. Published samples with human-validated transcripts: Kenyan Swahili, Cebuano, Tagalog / Filipino, Yoruba, Hausa and Amharic.

Proprietary and first-party. Every recording is collected directly by Silencio from consenting contributors. Nothing is scraped, and these recordings are not available in any other dataset on the internet.

Ethical sourcing, with provenance records. Every recording carries a consent record covering AI/ML training use, and contributor-level provenance documentation is available to licensees, including for EU AI Act training-data summaries. Contributors can withdraw consent, and withdrawal propagates to subsequent releases.

More at silencio.network.

For volume licensing, bespoke transcription or commissioned collection: hello@silencio.network