CoolFace
Datasetpublic

SilencioNetwork/complete-voiceai-speech-dataset

Silencio Voice AI Sample Dataset Speaker-attributed spontaneous speech. 363 labelled contributors across 111 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment. Hours 16.05 Clips 1,305 Speakers 363 Origin varieties 111 Languages 21 Configs 44 Speaker metadata origin region / variety, mother tongue, gender, device, OS… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset.

sourceHugging Facecc-by-nc-4.0updated 21h agoView on Hugging Face
1likes249downloads
Dataset Card

Silencio Voice AI Sample Dataset

Speaker-attributed spontaneous speech. 363 labelled contributors across 111 self-reported origin varieties, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment.

Hours16.05
Clips1,305
Speakers363
Origin varieties111
Languages21
Configs44
Speaker metadataorigin region / variety, mother tongue, gender, device, OS, recording environment
Audio48 kHz stereo WAV, as captured
TranscriptsASR-generated, 64% coverage, not human-reviewed — human-validated transcription available on demand
Licencecc-by-nc-4.0

Transcripts are ASR-generated and cover 64% of clips. They have not been human-reviewed.

Human-validated transcription is available on demand for this data. Word-level forced alignment included, as shipped in cebuano-speech and tagalog-filipino-speech. Available for this sample, for a larger subset of the same language, or for commissioned collection — [silencio.network/contact](https://www.silencio.network/contact)
This is a sample, not the catalogue. It is drawn to show format, audio quality and metadata, and its speaker mix does not represent what Silencio holds. The off-the-shelf catalogue behind it is far more diverse across regions, age groups, accents and recording conditions. Men are 86% of the speakers in this sample. Across the catalogue the gender split is roughly 40% female to 60% male, with men in the majority in most languages. Licensed subsets can be drawn to a specified gender, age or regional balance.

What this is for

Transcripts here are unreviewed ASR output — usable as pseudo-labels for semi-supervised work, but not a human-validated ASR training set. What it does carry is speaker-level labels on real-world audio, which supports a range of work needing no transcripts at all:

  • Accent and dialect classification. 111 self-reported origin varieties across 363 speakers, labelled per clip. No transcripts needed.
  • Spoken language identification. 21 languages in one release, recorded through the same app under the same conditions, so language is not confounded with recording setup.
  • Speaker identification and verification. 363 labelled speakers. Speaker-disjoint splits can be constructed directly from speaker_id.
  • Age and gender estimation. Self-reported birth year and gender on every speaker.
  • Robustness and fairness auditing. Device, operating system and recording environment are recorded per clip, so error rates can be broken down by capture condition as well as by speaker attribute.
  • Self-supervised pretraining. Unlabelled real-world speech is the input these methods want; the absence of transcripts is not a limitation here.

If you need transcripts, they are available on demand over this data — see below.

Load it

python
from datasets import load_dataset

ds = load_dataset("SilencioNetwork/complete-voiceai-speech-dataset", "amharic_ethiopia", split="train")
print(ds[0])

Requires pip install "datasets>=4.0" and FFmpeg ≥ 4.

Configs: amharic_ethiopia, english_algeria, english_australia, english_belarus, english_china, english_egypt, english_french_speaking, english_haiti, english_ireland, english_jamaica, english_kenya, english_mandarin, english_medical, english_nigeria, english_pakistan, english_poland, english_russia, english_south_africa, english_uganda, english_ukraine, english_united_kingdom, english_united_states, french_canada, french_global, german_germany, global_medical, gujarati_india, hausa_nigeria, igbo_nigeria, italian_italy, japanese_japan, kannada_india, malayalam_india, mandarin_chinese_china, marathi_india, odia_india, portuguese_brazil, russian_russia, spanish_mexico, spanish_mx_mexico, telugu_india, turkish_turkey, vietnamese_vietnam, yoruba_nigeria

Speaker and recording metadata

Speaker origin / varietySpeakers%
Nigeria - Nigerian Standard179.3%
United States - Californian168.8%
Canada (Québec) - Montréalais147.7%
Germany - Hochdeutsch (Standard)137.1%
Mainland China - Beijing126.6%
unknown116.0%
France - Lyonnais105.5%
Nigeria - Ibadan94.9%
Nigeria - Lagos Yoruba94.9%
Nigeria - Pidgin-influenced84.4%
Nigeria - Kano84.4%
United States - General American84.4%
Ethiopia - Addis Ababa73.8%
France - Parisian63.3%
France - Normand63.3%
United States - Midwestern63.3%
Mexico - Chilango (CDMX)63.3%
United States - New York City63.3%
Nigeria - Katsina52.7%
Mainland China - Shandong52.7%
GenderSpeakers%
male31386.2%
female4712.9%
prefernotto_say20.6%
non_binary10.3%
Ethnicity (self-reported)Speakers%
White13436.9%
Black or African American10228.1%
Asian7019.3%
Other185.0%
Hispanic or Latino185.0%
Middle Eastern102.8%
Mixed Race71.9%
Native American20.6%
Prefer not to say20.6%
DeviceClips%
Mobile94672.5%
Desktop35927.5%
EnvironmentClips%
home1,03179.0%
office1209.2%
outdoor1158.8%
public_transport120.9%
cafe90.7%
other80.6%
library50.4%
restaurant30.2%
gym10.1%
shopping_mall10.1%

Ethnicity is self-reported by the contributor at enrolment, using a fixed category list; it is not inferred from the audio and it is not a label of the speech. It is included because accent and speaker-attribute fairness work needs it, and it is collected under the same consent as the rest of the metadata.

Human-validated transcription — available on demand

This release is not human-transcribed. Silencio provides human transcription with word-level forced alignment on demand, over this sample or over a larger subset of the same language, and as part of commissioned collection.

These releases show exactly what that deliverable looks like — human transcript text, machine forced alignment, per-token start and end times:

To request transcription over this data or any other Silencio subset: [silencio.network/contact](https://www.silencio.network/contact)

Limitations

  • Sample scale. 1,305 clips, 363 speakers, 16.05 hours. This is a demonstration sample, not a training corpus.
  • Partial, unreviewed transcripts. 64% coverage, ASR-generated.
  • Gender imbalance. 313 of 363 speakers are male. Do not use this release for gender-comparative work. Balanced cohorts are available through the collection programme.
  • No acoustic annotation. Recording environment, background-noise class and SNR are not annotated in this release. Available for commissioned collection.
  • Audio is 48 kHz stereo WAV as captured. Resample and downmix before batching.
  • No baseline. No reference WER is published with this release.

Provenance and consent

Every recording is contributed by an opted-in participant through the Silencio platform, under a consent record covering AI/ML training use. Contributors can request deletion, and deletion propagates to downstream releases. Full provenance documentation is available to licensees.

License

cc-by-nc-4.0 — free for research and non-commercial use with attribution.

Attribution string: Silencio Network, Silencio Voice AI Sample Dataset, 2026. CC BY-NC 4.0.

Non-commercial covers research, evaluation and publication. Benchmarking a commercial product model against this data is a commercial use and needs a licence — ask, it is usually granted for evaluation. Model weights trained on this sample inherit the non-commercial restriction.

Commercial licensing: [silencio.network/contact](https://www.silencio.network/contact)

Citation

bibtex
@misc{silencio_complete_voiceai_speech_dataset_2026,
  title  = {Silencio Voice AI Sample Dataset},
  author = {Silencio Network},
  year   = {2026},
  url    = {https://huggingface.co/datasets/SilencioNetwork/complete-voiceai-speech-dataset}
}

Silencio: the catalogue behind this sample

This listing is a small fraction of what Silencio provides. Silencio specialises in underrepresented languages, accents and niche domains, delivered at scale. It draws on the largest global community of speech contributors and holds the largest off-the-shelf catalogue of niche-language speech available.

Recorded and available off the shelf. Audio already collected, with metadata, licensable today. Catalogue position as of September 2026:

Hours, off the shelfaround 500,000
Countries of contributor originaround 180
Catalogue refreshWeekly at record level, quarterly at catalogue level

Contributor community available on demand. Registered, consented contributors who can be activated for a specific brief:

Contributors available on demandaround 2,500,000
Countries180+
Languages that can be collectedaround 350

Anything not off the shelf can be sourced through the community. A language, an accent or dialect region, a demographic, a recording condition, a speech style or a specialist domain can be collected to a client's brief. Deep coverage across Africa, South-East Asia, South Asia and the Middle East.

Human-validated transcription is available for every language, scoped to each client's needs: script and orthography conventions, normalisation rules, word- or segment-level alignment, speaker labelling and turnaround. Every delivered clip is checked by a native-speaker reviewer. Published samples with human-validated transcripts: Kenyan Swahili, Cebuano, Tagalog / Filipino, Yoruba, Hausa and Amharic.

Proprietary and first-party. Every recording is collected directly by Silencio from consenting contributors. Nothing is scraped, and these recordings are not available in any other dataset on the internet.

Ethical sourcing, with provenance records. Every recording carries a consent record covering AI/ML training use, and contributor-level provenance documentation is available to licensees, including for EU AI Act training-data summaries. Contributors can withdraw consent, and withdrawal propagates to subsequent releases.

More at silencio.network.

For volume licensing, bespoke transcription or commissioned collection: [silencio.network/contact](https://www.silencio.network/contact)