CoolFace
Datasetpublic

Scicom-intl/clean-speech-raw-sources

Clean Speech — Raw Source Archives (mirror) Durable public mirrors of speech corpora whose original home is not Hugging Face (external academic hosts disappear, move, or go offline). HF is used only as a faster/durable mirror — the original source links are below and remain the canonical home. Every archive is byte-for-byte unmodified from its origin and retains its original licence and attribution. 42 corpora · 297 GB · 171+ files. Scope note. This mirror began as… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/clean-speech-raw-sources.

sourceHugging Facecc-by-nc-sa-4.0updated 1mo agoView on Hugging Face
0likes205downloads
Dataset Card

Clean Speech — Raw Source Archives (mirror)

Durable public mirrors of speech corpora whose original home is not Hugging Face (external academic hosts disappear, move, or go offline). HF is used only as a faster/durable mirror — the original source links are below and remain the canonical home. Every archive is byte-for-byte unmodified from its origin and retains its original licence and attribution.

42 corpora · 297 GB · 171+ files.

Scope note. This mirror began as high-sample-rate (>=44.1 kHz) material for Scicom-intl/TTS-Clean44k. It has since broadened: it now also holds 16 kHz ASR corpora (Samromur, the Google/OpenSLR ASR sets, and others). Check the per-corpus notes before assuming sample rate.

Why mirror at all

Four of the corpora tracked alongside this repo already have dead or moved origin links, and several hosts here are slow enough that a large download is a multi-hour gamble (repository.clarin.is serves no HTTP range requests at all, so a dropped transfer restarts from zero). Mirroring makes the archives fast, resumable, and durable.

How to use

Tooling lives in malaysia-ai/dataset under multilingual-tts/prepare_raw_sources.py:

bash
python3 prepare_raw_sources.py sources                     # registry + layouts
python3 prepare_raw_sources.py probe    --source <name>    # size/layout, no bulk download
python3 prepare_raw_sources.py mirror   --source <name>    # origin -> this repo
python3 prepare_raw_sources.py external --source <name> --from-mirror --run-prepare

--from-mirror pulls from here instead of the origin. Because HF serves ranged reads, you can also inspect an archive's layout remotely — read the zip central directory, or inflate a single metadata.tsv — without downloading gigabytes.

Ingestion status

  • —verified — layout confirmed by reading the real archive; safe to ingest.
  • —unverified — plausible layout, not yet exercised end-to-end.
  • —mirror-only — archive preserved, internal layout not confirmed. Do not assume an index->audio mapping; confirm it first.

Contents

OpenSLR (Google crowdsourced + Thorsten) — verified

Utterance ids are <langgender>_<speaker>_<uttnum>, so these carry real speaker labels and need no speaker clustering.

folderlanghoursGBlicenceorigin
sundanese-asr_openslr36su~33323.00CC BY-SA 4.0src
javanese-asr_openslr35jv~29618.90CC BY-SA 4.0src
sinhala-asr_openslr52si~22414.68CC BY-SA 4.0src
thorsten-german_openslr95de~243.00CC0src
peruvian-spanish_openslr73es~91.94CC BY-SA 4.0src
javanese-tts_openslr41jv~91.89CC BY-SA 4.0src
sundanese-tts_openslr44su~91.47CC BY-SA 4.0src
venezuelan-spanish_openslr75es~111.04CC BY-SA 4.0src
sinhala-tts_openslr30si~150.70CC BY-SA 4.0src
puertorico-spanish_openslr74es~10.21CC BY-SA 4.0src

CLARIN-IS (Icelandic)

The Samromur releases share one layout: audio at audio/<speaker_id>/<speaker_id>-<utt_id>.flac, 16 kHz mono, with a UTF-8 metadata.tsv (id, speaker_id, filename, sentence, sentence_norm, gender, age, ...). For TTS prefer `sentence` over sentence_norm: the latter is lowercased with punctuation stripped, discarding prosodic signal. malromur and icelandic-psc are separate corpora with their own layouts — see the caveats below.

folderlanghoursGBlicenceorigin
malromur_clarin202is~13610.22CC BY 4.0src
samromur-l2_clarin263is~1507.94CC BY 4.0src
samromur-21.05_clarin189is~1007.05CC BY 4.0src
samromur-mimic_clarin264is~673.60CC BY 4.0src
samromur-queries_clarin180is~211.03CC BY 4.0src
icelandic-psc_clarin197is~210.89CC BY 3.0src

Zenodo — layout verified

folderlanghoursGBlicenceorigin
som-tts-thai_zenodo21530909th~303.63CC BY 4.0src
thorsten-de-2022.10_zenodo7265581de~341.39CC BY 4.0src
uspdatro-romanian_zenodo7898233ro~40.42CC BY 4.0src
thorsten-hessisch_zenodo10511260de~30.27CC BY 4.0src
dendi-parakou_zenodo6591186ddn~10.14CC BY 4.0src

Other origins

folderlanghoursGBlicenceorigin
thai-elderly-speech_githubth~175.07CC BY-SA 4.0src
siwis-french_edinburghfr~102.87CC BY 4.0src

Zenodo — mirror-only (layout NOT confirmed)

Preserved as-is; confirm the layout before ingesting.

folderlanghoursGBlicenceorigin
tundra-multilingual-tts_zenodo12543428multi?16.00CC BY 4.0src
earnings25-finance-en_zenodo18762168en~50012.04CC BY 4.0src
uvigo-gl-voices_zenodo8027725gl?9.63CC BY 4.0src
pakistan-multilingual_zenodo19323537ur?8.35CC BY 4.0src
lada-ukrainian-tts_zenodo7396774uk~106.78Apache 2.0src
chulalongkorn-spoken-thai_zenodo17366698th?6.23CC BY-NC-SA 4.0src
cv298-japanese-tts_zenodo21119791ja?5.20CC0src
runyankore-rukiga_zenodo20478456nyn?3.94CC BY 4.0src
coala-dutch_zenodo8413584nl?2.47CC0src
coala-english_zenodo8268928en?1.78CC0src
coala-italian_zenodo8413135it?1.76CC0src
french-tts-blizzard_zenodo13918615fr?1.68CC BY 4.0src
gronings-nasal-besemah_zenodo7946870gos?1.02CC BY 4.0src
dioula-words_zenodo13958016dyu?0.11CC BY 4.0src
baule_zenodo6705861bci?0.05CC BY 4.0src
alsatian-character-speech_zenodo10613199gsw?0.01CC BY 4.0src

Earlier additions

  • —BibleTTS (OpenSLR 129) — bibletts_openslr129/*.tgz (92.8 GB). 48 kHz mono studio, up to 80 h single-speaker per language: Akuapem Twi, Asante Twi, Ewe, Hausa, Lingala, Yoruba. CC BY-SA 4.0. Source: https://www.openslr.org/129/ · Meyer et al., Interspeech 2022.
  • —RAVDESS — RAVDESS_Audio_Speech_Actors_01-24.zip (0.21 GB). 48 kHz, 1440 files, 24 actors (emotional). CC BY-NC-SA 4.0. Source: https://zenodo.org/records/1188976 · Livingstone & Russo, PLoS ONE 13(5).
  • —DAPS — DAPS_daps.tar.gz (16.1 GB). 44.1 kHz, 20 speakers. CC BY-NC-SA 4.0. Source: https://zenodo.org/records/4660670 · Mysore, IEEE SPL 22(8). Not ingestible as {audio, text} pairs — long unaligned reads with no per-utterance transcripts; needs segmentation + forced alignment first.

Per-corpus caveats

  • —`samromur-mimic_clarin264` — ships 32,892 TTS-generated mp3 in repeat_audio/ beside 91,408 human flac in audio/, and metadata.tsv indexes both. Exclude repeat_audio/ or you will train on synthetic speech.
  • —`icelandic-psc_clarin197` — 12 long parliamentary recordings with Transcriber .trs files, not per-utterance segments. Needs segmentation first.
  • —`runyankore-rukiga_zenodo20478456` — 265 wav with no index file; the transcripts appear to live in the bundled .docx (mirrored for that reason).
  • —`pakistan-multilingual_zenodo19323537` — the part inspected had no index file; transcripts may be in another part.
  • —`french-tts-blizzard_zenodo13918615` — 39 long wav + 3 csv, not per-utterance clips.
  • —*`coala-** — distributed as .rar; needs unrar/7z`.
  • —Split archives — samromur-21.05 and thai-elderly-speech are multi-volume. Info-ZIP volumes (.zip + .z01..) and byte-splits (.001/.002..) both need 7z; Python's zipfile reads neither.

Licensing

Licences are those of the original authors and vary per corpus (CC0, CC BY 4.0, CC BY-SA 4.0, CC BY-NC-SA 4.0, Apache 2.0) — see each table above and follow the origin link for the authoritative terms and citation. Several corpora are non-commercial. The repo-level tag is the most restrictive of the set and is not a substitute for checking the individual corpus.

Mirrored for research reproducibility. All attribution belongs to the original authors.