CoolFace
17 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes888 downloads4mo agoHugging Face02ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes697 downloads3mo agoHugging Face03deepghs /arknights_voices_zh ZH Voice-Text Dataset for Arknights Waifus This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models. Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset. 12431 records, 25.9 hours in total. Average duration is approximately 7.49s. id char_id voice_actor_name voice_title voice_text time sample_rate file_size filename mimetype file_url char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.tabularautomatic-speech-recognition10K<n<100K6 likes432 downloads2y agoHugging Face04HeshamHaroon /arabic-msa-25k-saudi-male-tashkeel Arabic MSA 25K — Saudi Male (Tashkeel) 25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories. Dataset Summary arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA) speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.tabulartext-to-speech10K<n<100K10 likes194 downloads5mo agoHugging Face05baryonlabs /open-ko-s2s-eval-artifacts Open Ko-S2S 평가 산출물 (감사용) ⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다 KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의 ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과 채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다. Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라 ref 원문을 그대로 담고 있습니다. 라이선스 보유자의 검증 절차 AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다. 리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다. 자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.tabularautomatic-speech-recognitionn<1K0 likes192 downloads2mo agoHugging Face06Edge0 /ark-asr-3b-open-asr-leaderboard-results ARK-ASR-3B Open ASR Leaderboard Results Raw JSONL manifests for AutoArk-AI/ARK-ASR-3B on the public English short-form hf-audio/open-asr-leaderboard splits. These manifests were generated on a local 8x RTX 4090 machine and scored with the shared Open ASR Leaderboard scorer: PYTHONPATH=. python - <<'PY' from normalizer.eval_utils import score_results score_results( 'ark_asr/results.AutoArk-AI-ARK-ASR-3B_20260622_official', 'AutoArk-AI/ARK-ASR-3B', ) PY Important:… See the full description on the dataset page: https://huggingface.co/datasets/Edge0/ark-asr-3b-open-asr-leaderboard-results.tabularautomatic-speech-recognition10K<n<100K12 likes185 downloads3mo agoHugging Face07deepghs /arknights_voices_jp JP Voice-Text Dataset for Arknights Waifus This is the JP voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models. Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset. 10905 records, 26.3 hours in total. Average duration is approximately 8.7s. id char_id voice_actor_name voice_title voice_text time sample_rate file_size filename mimetype file_url char_427_vigil_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_jp.tabularautomatic-speech-recognition10K<n<100K4 likes145 downloads2y agoHugging Face08Rabe3 /egyptian-arabic-tts-diacritized Egyptian Arabic TTS Corpus (Diacritized) 97,163 utterances / ~334 hours of Egyptian Arabic speech at 24 kHz, with diacritized transcripts — the short vowels that Arabic script does not write. Why diacritics Arabic is an abjad: short vowels are unwritten, so كتب may be kataba, kutiba, or kutub. A TTS model with no Arabic pretraining cannot infer which, and guesses — which native listeners hear as a foreign accent with constant mispronunciation. This is invisible to… See the full description on the dataset page: https://huggingface.co/datasets/Rabe3/egyptian-arabic-tts-diacritized.tabulartext-to-speech10K<n<100K0 likes142 downloads1mo agoHugging Face09FatimahEmadEldin /Moroccan-Arabic-Multimodal-Emotion-Recognition MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging) A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits. Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.audiotext-to-speech1K<n<10K1 likes84 downloads5mo agoHugging Face10TNSA /Aren ARen — Arabic/English ASR Robustness Set Curated and published by TNSA AI. A small, deliberately hard evaluation set for Arabic and English speech recognition. Every clip exists in three acoustic conditions so you can measure not just how a model scores, but how fast it falls apart as the channel degrades. Built because clean read-speech benchmarks stop discriminating between modern ASR systems long before real deployments stop breaking. Why it exists On clean… See the full description on the dataset page: https://huggingface.co/datasets/TNSA/Aren.audioautomatic-speech-recognitionn<1K0 likes81 downloads1mo agoHugging Face11deepghs /arknights_voices_en EN Voice-Text Dataset for Arknights Waifus This is the EN voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models. Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset. 8216 records, 16.1 hours in total. Average duration is approximately 7.05s. id char_id voice_actor_name voice_title voice_text time sample_rate file_size filename mimetype file_url char_214_kafka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_en.tabularautomatic-speech-recognition1K<n<10K4 likes66 downloads2y agoHugging Face12FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes51 downloads5mo agoHugging Face13deepghs /arknights_voices_kr KR Voice-Text Dataset for Arknights Waifus This is the KR voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models. Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset. 9996 records, 23.1 hours in total. Average duration is approximately 8.34s. id char_id voice_actor_name voice_title voice_text time sample_rate file_size filename mimetype file_url char_4046_ebnhlz_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_kr.tabularautomatic-speech-recognition1K<n<10K2 likes43 downloads2y agoHugging Face14abdo1819 /arabic-english-code-switching-review-annotations Review Annotations for Arabic-English Code-Switching Speech This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts. The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index. Coverage and outcomes The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.tabularautomatic-speech-recognition10K<n<100K0 likes33 downloads2mo agoHugging Face15arvinsingh /welsh-speech-dataset Welsh Speech Dataset A multimodal dataset of 33 speakers producing 10 Welsh phrases, captured using 3DMD technology with audio and dense facial landmark annotations. Dataset Overview Speakers: 33 participants Phrases: 10 Welsh phrases per speaker Sequences: ~330 (33 speakers x 10 phrases) Modalities: Audio recordings (.wav) 3D facial reconstructions (.obj meshes + texture maps) 68-point facial landmarks (ibug68 template) Fluency Scores: Each phrase rated 0-5 (5 =… See the full description on the dataset page: https://huggingface.co/datasets/arvinsingh/welsh-speech-dataset.tabularautomatic-speech-recognitionn<1K0 likes16 downloads8mo agoHugging Face16Khaledtelbahnasy /egyptian-arabic-stt-data Egyptian Arabic STT Dataset Synthetic Egyptian Arabic speech dataset generated by the Synthetic Egyptian Speech Data Pipeline. Samples are human-reviewed and quality-validated using Whisper ASR (WER/CER). Dataset Statistics Metric Value Total samples 50 Total duration 85.2s (0.02h) Dialect validated 50 / 50 Average WER 0.4136 Average CER 0.1642 Topics food_ordering Fields Field Type Description id string Deterministic… See the full description on the dataset page: https://huggingface.co/datasets/Khaledtelbahnasy/egyptian-arabic-stt-data.tabularautomatic-speech-recognitionn<1K1 likes16 downloads5mo agoHugging Face17driodnexus /droidnexus-arabic-editorial-speech-scorecard-mini DroidNexus Arabic Editorial Speech Scorecard Mini A public DroidNexus Labs scorecard dataset for Arabic speech workflows: representative editorial scenarios, latency targets, overlap pressure, and the metric stack that decides whether a transcript is usable. Why this exists This dataset is the first public speech artifact layer for DroidNexus Labs. It publishes representative editorial workloads and evaluation pressure before claiming a full source-audio benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-editorial-speech-scorecard-mini.tabularautomatic-speech-recognitionn<1K0 likes13 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.