dic
Datasets
All datasets matching “dic”phonemizer-dicts
Phonemizer Dicts
Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup.
Files
en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA
Provenance
Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here.
Regeneration
See scripts/generate-espeak-dict.py in the tts-rd-team repo.
Crenis-DICA-DataTesting new way to compress imagesAll file contains 20,000 WebP images in 1536 resolutions instead
wavepulse-radio-raw-transcripts
WavePulse Radio Raw Transcripts
Dataset Summary
WavePulse Radio Raw Transcripts is a large-scale dataset containing segment-level transcripts from 396 radio stations across the United States, collected between June 26, 2024, and Dec 29th, 2024. The dataset comprises >250 million text segments derived from 750,000+ hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The summarized version of these transcripts is available here. For… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-raw-transcripts.wavepulse-radio-summarized-transcripts
WavePulse Radio Summarized Transcripts
Dataset Summary
WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.FR_Drugs_tiret_dictation_augmented
FR_Drugs_tiret_dictation_augmented
Dictee de listes de medicaments au format "tiret" (patterns du dataset
FR_Drugs_dictation_with_dashes_pattern, molecules completees depuis les sources
autocorrect/medical), synthetisee en TTS Coqui XTTS v2 (600 voix clonees) puis
augmentee acoustiquement.
element
valeur
extraits
141960
moteur
Coqui XTTS v2, 600 voix clonees (16 kHz mono)
base
audio/coqui/<shard>/
parasite seul (65%)
audio_babble_only/
echo seul (25%)… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/FR_Drugs_tiret_dictation_augmented.ILRDF_Dict_Rukai
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ILRDF_Dict_Rukai
This is a noncanonical compatibility mirror. Use FormosanBank/ILRDF_Dicts for the complete canonical dataset and stable download contract.
This mirror… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ILRDF_Dict_Rukai.
