CoolFace
Datasetpublic

hamayawayuri/common-voice-26-en-audio

English — English (en) This datasheet is for cv-corpus-26.0-2026-06-12 of the Mozilla Common Voice Scripted Speech dataset for English [English - en]. The dataset contains 2583051 clips representing 3780.95 hours of recorded speech (2784.88 hours validated) from 100172 speakers, recorded from a text corpus of 1,721,897 sentences. Language English is a West Germanic language with origins in England. There are an estimated 1.5 billion English speakers, making it the… See the full description on the dataset page: https://huggingface.co/datasets/hamayawayuri/common-voice-26-en-audio.

sourceHugging Faceupdated 23d agoView on Hugging Face
0likes122downloads
Dataset Card

English — English (en)

This datasheet is for cv-corpus-26.0-2026-06-12 of the Mozilla Common Voice Scripted Speech dataset for English [English - en]. The dataset contains 2583051 clips representing 3780.95 hours of recorded speech (2784.88 hours validated) from 100172 speakers, recorded from a text corpus of 1,721,897 sentences.

Language

English is a West Germanic language with origins in England. There are an estimated 1.5 billion English speakers, making it the most widely spoken language in the world. English is commonly learned as a second language in many countries.

Accents

CodeAccentClipsSpeakers
usUnited States English573,393 (22.2%)10,970 (11.0%)
englandEngland English204,106 (7.9%)3,445 (3.4%)
indianIndia and South Asia (India, Pakistan, Sri Lanka)152,723 (5.9%)3,029 (3.0%)
canadaCanadian English101,061 (3.9%)1,220 (1.2%)
australiaAustralian English69,363 (2.7%)948 (0.9%)
scotlandScottish English68,131 (2.6%)270 (0.3%)
africanSouthern African (South Africa, Zimbabwe, Namibia)60,292 (2.3%)441 (0.4%)
newzealandNew Zealand English20,768 (0.8%)230 (0.2%)
irelandIrish English11,204 (0.4%)266 (0.3%)
philippinesFilipino7,492 (0.3%)207 (0.2%)
hongkongHong Kong English7,004 (0.3%)190 (0.2%)
singaporeSingaporean English4,722 (0.2%)112 (0.1%)
malaysiaMalaysian English4,235 (0.2%)154 (0.2%)
walesWelsh English3,032 (0.1%)119 (0.1%)
bermudaWest Indies and Bermuda (Bahamas, Bermuda, Jamaica, Trinidad)1,231 (0.0%)74 (0.1%)
southatlandticSouth Atlantic (Falkland Islands, Saint Helena)332 (0.0%)9 (0.0%)
otherOther200,693 (7.8%)1,681 (1.7%)

Demographic information

The dataset includes the following self-declared age and gender distributions. A coverage summary is shown below each table.

Gender

Self-declared gender information. The table shows clip and speaker counts with percentages. Speakers who did not declare a gender are listed as Unspecified. A dash (-) indicates zero.

CodeGenderClipsSpeakers
male_masculineMale, masculine1,117,432 (43.3%)18,517 (18.5%)
female_feminineFemale, feminine457,904 (17.7%)5,563 (5.6%)
transgenderTransgender156 (0.0%)11 (0.0%)
non-binaryNon-binary411 (0.0%)17 (0.0%)
donotwishtosayPrefer not to say2,009 (0.1%)44 (0.0%)
-Unspecified1,004,989 (38.9%)78,307 (78.2%)

Gender declared: 1,578,062 of 2,583,051 clips (61.1%), 21,865 of 100,172 speakers (21.8%)

Age

Self-declared age information. The table shows clip and speaker counts with percentages. Speakers who did not declare an age are listed as Unspecified. A dash (-) indicates zero.

CodeAgeClipsSpeakers
teensTeens151,296 (5.9%)3,255 (3.2%)
twentiesTwenties641,400 (24.8%)11,552 (11.5%)
thirtiesThirties357,504 (13.8%)5,369 (5.4%)
fourtiesFourties240,835 (9.3%)2,651 (2.6%)
fiftiesFifties136,070 (5.3%)1,590 (1.6%)
sixtiesSixties115,750 (4.5%)915 (0.9%)
seventiesSeventies17,522 (0.7%)365 (0.4%)
eightiesEighties2,427 (0.1%)55 (0.1%)
ninetiesNineties306 (0.0%)12 (0.0%)
-Unspecified919,941 (35.6%)76,964 (76.8%)

Age declared: 1,663,110 of 2,583,051 clips (64.4%), 23,208 of 100,172 speakers (23.2%)

Data splits for modelling

Clip buckets

BucketClips
Validated1,902,564 (73.7%)
Invalidated313,382 (12.1%)
Other367,105 (14.2%)

Training splits

SplitClips
Train1,147,812 (60.3%)
Dev16,403 (0.9%)
Test16,403 (0.9%)

Training split coverage: 1,180,618 of 1,902,564 validated clips (62.1%)

The dataset contains 1902564 validated, 313382 invalidated, and 367105 unresolved clips. The average clip duration is 5.27 seconds.

Text corpus

Validated sentences: 1,681,666

CategoryCount
Unvalidated sentences40,231
Pending sentences36,278
Rejected sentences3,953
Reported sentences9,773

The corpus contains 1,721,897 sentences: 1,681,666 validated and 40,231 unvalidated (36,278 pending review, 3,953 rejected), with 9,773 reported for review.

Writing system

The English writing system is based off of the latin alphabet.

Symbol table

a b c d e f g h i j k l m n o p q r s t u v w x y z

Sample

There follows a randomly selected sample of five sentences from the corpus.

  1. 1.For example, a heavily rhythmic speech filled with mnemonic devices enhances memory and recall.
  2. 2.Opposition to construction came from various citizens' groups and different levels of local government.
  3. 3."""Perfect Balance"" was produced and engineered by Lionel Hicks."
  4. 4.It provides people with structure and purpose and a sense of identity.
  5. 5.I learned this from the late Mr. W. Simpson.

Sources

SourceSentences
wiki1,537,302 (92.9%)
sentence-collector61,569 (3.7%)
covost2-xx_en28,881 (1.7%)
Other26,425 (1.6%)

Text domains

CodeDomainClipsSpeakers
generalGeneral677 (0.0%)334 (0.3%)
agriculture_foodAgriculture and Food169 (0.0%)115 (0.1%)
automotive_transportAutomotive and Transport8 (0.0%)7 (0.0%)
financeFinance44 (0.0%)33 (0.0%)
service_retailService and Retail31 (0.0%)24 (0.0%)
healthcareHealthcare26 (0.0%)22 (0.0%)
historylawgovernmentHistory, Law and Government126 (0.0%)94 (0.1%)
media_entertainmentMedia and Entertainment118 (0.0%)92 (0.1%)
nature_environmentNature and Environment64 (0.0%)46 (0.0%)
newscurrentaffairsNews and Current Affairs13 (0.0%)13 (0.0%)
technology_roboticsTechnology and Robotics104 (0.0%)75 (0.1%)
language_fundamentalsLanguage Fundamentals11 (0.0%)10 (0.0%)

Fields

Clips

Each row of a tsv file represents a single audio clip, and contains the following information:

  • —client_id - hashed UUID of a given user
  • —path - relative path of the audio file
  • —sentence - the sentence to be read aloud
  • —sentence_id - unique identifier for the sentence
  • —sentence_domain - domain classification(s) of the sentence
  • —up_votes - number of people who said audio matches the text
  • —down_votes - number of people who said audio does not match text
  • —age - age of the speaker[^1]
  • —gender - gender of the speaker[^1]
  • —accents - accents of the speaker[^1]
  • —variant - variant of the language[^1]
  • —locale - locale code of the language
  • —segment - if sentence belongs to a custom dataset segment, it will be listed here

[^1]: For a full list of age, gender, and accent options, see the demographics spec. These will only be reported if the speaker opted in to provide that information.

validated_sentences.tsv

The validated_sentences.tsv file contains one row per validated sentence in the text corpus:

  • —sentence_id - unique identifier for the sentence
  • —sentence - the sentence text
  • —variant - the variant of the language
  • —sentence_domain - the domain(s) the sentence belongs to
  • —source - the source the sentence was collected from
  • —is_used - whether the sentence is still in circulation for recording
  • —clips_count - number of clips recorded for this sentence
unvalidated_sentences.tsv

The unvalidated_sentences.tsv file contains one row per unvalidated sentence in the text corpus:

  • —sentence_id - unique identifier for the sentence
  • —sentence - the sentence text
  • —variant - the variant of the language
  • —sentence_domain - the domain(s) the sentence belongs to
  • —source - the source the sentence was collected from
  • —up_votes - number of upvotes the sentence received
  • —down_votes - number of downvotes the sentence received
  • —status - current status of the sentence (pending or rejected)

Get involved

Community links

Discussions

Contribute

Licence

This dataset is released under the Creative Commons Zero (CC-0) licence. By downloading this data you agree to not determine the identity of speakers in the dataset.