CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01egcortes /asr-jargon-specialized-vocabulary A Dataset for Evaluating ASR on Specialized Vocabulary Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026). Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code Configs Config Language Description synthetic_terms_en English Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms synthetic_terms_pt Portuguese Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes635 downloads3mo agoHugging Face02TCabbage /gsat-vocab-sentences-tts GSAT Vocabulary TTS Audio Text-to-speech audio files for GSAT (General Scholastic Ability Test) English vocabulary. Structure audio/ - MP3 audio files organized by hash prefix (e.g., audio/ab/abcd1234....mp3) index.jsonl - Index file mapping hashes to text and TTS engine used Engines Kokoro (af_heart voice) - Used for lemmas (single words/phrases) Supertonic (M1 voice) - Used for example sentences Audio Format Format: MP3 Sample rate: 24kHz… See the full description on the dataset page: https://huggingface.co/datasets/TCabbage/gsat-vocab-sentences-tts.audiotext-to-speech10K<n<100K0 likes316 downloads9mo agoHugging Face03SwiftieJerry /english-vocabulary-materialsgated English Vocabulary Teaching Materials(雅思与初高中词汇教学资料) 中学段的英语词汇教学资料:雅思分级词汇(预备班 / 一阶 / 二阶 / 三阶)的词汇本、词测本、配套听力录音与听说读讲义,外加初高中词表。 原始材料是 PDF、MP3 和 Excel —— 词表分散在 Excel 的多张工作表里,音频按中文文件名散落各目录,PDF 里的词测没法检索。这份仓库做了两件事:60 个原始文件原样归档不做改动,另外从 Excel 抽出 8257 条结构化词条存成 CSV/JSONL,可以直接读进来做背诵、默写、出题或全文检索。 ⚠️ 版权提醒 这批材料整理自绿新的教研资料,不是原创数据集,著作权归原权利人所有。仓库采用 CC BY-NC-ND 4.0 并开启 gated access:禁止商业使用、再分发、公开镜像与演绎;研究用途允许,但须按第 7 节格式署名。完整条款见第 6 节。 1. 数据总览 指标 数值 清单内文件 85(另有… See the full description on the dataset page: https://huggingface.co/datasets/SwiftieJerry/english-vocabulary-materials.audiofill-mask10K<n<100K1 likes42 downloads29d agoHugging Face04BrunoHays /en_librispeech_fleurs_vocab_hints en_librispeech_fleurs_vocab_hints Made from concatenating the test splits of the english config datasets of fleurs and librispeech. For each sample, we used zipf_frequency from wordfreq to extract complex words that were added as a tag like <vocab: splendours> and in the vocab_hints column. Moreover, we randomly sampled 10 vocab_hints from other samples and added them under vocab_distractors column. We filtered all samples were no complex word could be found. All audio samples were… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/en_librispeech_fleurs_vocab_hints.audion<1K0 likes32 downloads5mo agoHugging Face05FormosanBank /ePark_xue_xi_ci_biao_learning_vocabulary FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum. This is… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_xue_xi_ci_biao_learning_vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes29 downloads2mo agoHugging Face06whucedar /amoros_prof_vocab_03audio1K<n<10K0 likes20 downloads1y agoHugging Face07hamaada /arabic-somali-vocabbaudion<1K0 likes11 downloads1y agoHugging Face08whucedar /amoros_prof_vocab_01audion<1K0 likes9 downloads1y agoHugging Face09whucedar /amoros_prof_vocab_02audion<1K0 likes7 downloads1y agoHugging Face10vebaev /nhk_vocabaudion<1K0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.