hamayawayuri/common-voice-26-en-audio
English — English (en) This datasheet is for cv-corpus-26.0-2026-06-12 of the Mozilla Common Voice Scripted Speech dataset for English [English - en]. The dataset contains 2583051 clips representing 3780.95 hours of recorded speech (2784.88 hours validated) from 100172 speakers, recorded from a text corpus of 1,721,897 sentences. Language English is a West Germanic language with origins in England. There are an estimated 1.5 billion English speakers, making it the… See the full description on the dataset page: https://huggingface.co/datasets/hamayawayuri/common-voice-26-en-audio.
English — English (en)
This datasheet is for cv-corpus-26.0-2026-06-12 of the Mozilla Common Voice Scripted Speech dataset for English [English - en]. The dataset contains 2583051 clips representing 3780.95 hours of recorded speech (2784.88 hours validated) from 100172 speakers, recorded from a text corpus of 1,721,897 sentences.
Language
English is a West Germanic language with origins in England. There are an estimated 1.5 billion English speakers, making it the most widely spoken language in the world. English is commonly learned as a second language in many countries.
Accents
Demographic information
The dataset includes the following self-declared age and gender distributions. A coverage summary is shown below each table.
Gender
Self-declared gender information. The table shows clip and speaker counts with percentages. Speakers who did not declare a gender are listed as Unspecified. A dash (-) indicates zero.
Gender declared: 1,578,062 of 2,583,051 clips (61.1%), 21,865 of 100,172 speakers (21.8%)
Age
Self-declared age information. The table shows clip and speaker counts with percentages. Speakers who did not declare an age are listed as Unspecified. A dash (-) indicates zero.
Age declared: 1,663,110 of 2,583,051 clips (64.4%), 23,208 of 100,172 speakers (23.2%)
Data splits for modelling
Clip buckets
Training splits
Training split coverage: 1,180,618 of 1,902,564 validated clips (62.1%)
The dataset contains 1902564 validated, 313382 invalidated, and 367105 unresolved clips. The average clip duration is 5.27 seconds.
Text corpus
Validated sentences: 1,681,666
The corpus contains 1,721,897 sentences: 1,681,666 validated and 40,231 unvalidated (36,278 pending review, 3,953 rejected), with 9,773 reported for review.
Writing system
The English writing system is based off of the latin alphabet.
Symbol table
a b c d e f g h i j k l m n o p q r s t u v w x y z
Sample
There follows a randomly selected sample of five sentences from the corpus.
- For example, a heavily rhythmic speech filled with mnemonic devices enhances memory and recall.
- Opposition to construction came from various citizens' groups and different levels of local government.
- """Perfect Balance"" was produced and engineered by Lionel Hicks."
- It provides people with structure and purpose and a sense of identity.
- I learned this from the late Mr. W. Simpson.
Sources
Text domains
Fields
Clips
Each row of a tsv file represents a single audio clip, and contains the following information:
client_id- hashed UUID of a given userpath- relative path of the audio filesentence- the sentence to be read aloudsentence_id- unique identifier for the sentencesentence_domain- domain classification(s) of the sentenceup_votes- number of people who said audio matches the textdown_votes- number of people who said audio does not match textage- age of the speaker[^1]gender- gender of the speaker[^1]accents- accents of the speaker[^1]variant- variant of the language[^1]locale- locale code of the languagesegment- if sentence belongs to a custom dataset segment, it will be listed here
[^1]: For a full list of age, gender, and accent options, see the demographics spec. These will only be reported if the speaker opted in to provide that information.
validated_sentences.tsv
The validated_sentences.tsv file contains one row per validated sentence in the text corpus:
sentence_id- unique identifier for the sentencesentence- the sentence textvariant- the variant of the languagesentence_domain- the domain(s) the sentence belongs tosource- the source the sentence was collected fromis_used- whether the sentence is still in circulation for recordingclips_count- number of clips recorded for this sentence
unvalidated_sentences.tsv
The unvalidated_sentences.tsv file contains one row per unvalidated sentence in the text corpus:
sentence_id- unique identifier for the sentencesentence- the sentence textvariant- the variant of the languagesentence_domain- the domain(s) the sentence belongs tosource- the source the sentence was collected fromup_votes- number of upvotes the sentence receiveddown_votes- number of downvotes the sentence receivedstatus- current status of the sentence (pendingorrejected)
Get involved
Community links
Discussions
Contribute
Licence
This dataset is released under the Creative Commons Zero (CC-0) licence. By downloading this data you agree to not determine the identity of speakers in the dataset.
