CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Parakeet-Inc /J-HARD-TTS-Eval J-HARD-TTS-Eval [!NOTE] For full documentation, detailed benchmark results, and methodology, please refer to the GitHub Repository. Overview J-HARD-TTS-Eval is a benchmark designed to evaluate the robustness of autoregressive Japanese Text-To-Speech (TTS) models. It focuses on specific failure modes such as stability in short sequences, repetition handling, and context completion. Usage You can easily load the dataset using the Hugging Face datasets… See the full description on the dataset page: https://huggingface.co/datasets/Parakeet-Inc/J-HARD-TTS-Eval.audiotext-to-speechn<1K6 likes45 downloads8mo agoHugging Face02NeurologyAI /neuro-parakeet-foodgated neuro-whisper-v1 Dataset Description This is a synthetic dataset for German medical speech recognition, specifically designed for fine-tuning ASR models on neuro-oncology and neurology terminology. The dataset provides a comprehensive coverage of German medical terminology in the neurology and neuro-oncology domains. Data Generation Voice Data: Synthetically generated using Resemble AI Chatterbox TTS Text Data: Medical text generated with Qwen/Qwen3-30B-A3B… See the full description on the dataset page: https://huggingface.co/datasets/NeurologyAI/neuro-parakeet-food.audioautomatic-speech-recognition10K<n<100K4 likes37 downloads8mo agoHugging Face03TieIncred /parakeet-tdt-blind-spots Blind Spots of nvidia/parakeet-tdt-0.6b-v2 This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data. Model Under Test Property Value Model nvidia/parakeet-tdt-0.6b-v2 Parameters 600M Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.audioautomatic-speech-recognitionn<1K0 likes25 downloads7mo agoHugging Face04Trelis /eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920 Evaluation Results: parakeet-tdt-0.6b-v3 Evaluation results from Whisper model evaluation. Summary Model WER CER nvidia/parakeet-tdt-0.6b-v3 37.59% 20.64% Source Data Evaluation Dataset: Trelis/eka-hard Model Evaluated: nvidia/parakeet-tdt-0.6b-v3 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error Rate for… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-eka-hard-20260408-1920.audion<1K0 likes12 downloads6mo agoHugging Face05Trelis /eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926 Evaluation Results: parakeet-tdt-0.6b-v3 Evaluation results from Whisper model evaluation. Summary Model WER CER nvidia/parakeet-tdt-0.6b-v3 11.34% 3.63% Source Data Evaluation Dataset: Trelis/medical-terms-2025 Model Evaluated: nvidia/parakeet-tdt-0.6b-v3 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-medical-terms-2025-20260408-1926.audion<1K1 likes12 downloads6mo agoHugging Face06rdsm /parakeet-stt-redone parakeet-stt-redone What this is 108,276 raw→clean transcript pairs sourced from aldigobbler/stt-correction, re-labeled using GLM-5.1-FP8 as the teacher model with our production cleanup prompt. How it differs from the source dataset aldigobbler/stt-correction this dataset Target Verbatim transcript restoration (lowercase, no punctuation, fillers kept/restored) Polished readable text — punctuated, paragraphed, fillers selectively removed… See the full description on the dataset page: https://huggingface.co/datasets/rdsm/parakeet-stt-redone.texttext-generation100K<n<1M0 likes11 downloads3mo agoHugging Face07Trelis /eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930 Evaluation Results: parakeet-tdt-0.6b-v3 Evaluation results from Whisper model evaluation. Summary Model WER CER nvidia/parakeet-tdt-0.6b-v3 15.94% 10.13% Source Data Evaluation Dataset: Trelis/multimed-hard Model Evaluated: nvidia/parakeet-tdt-0.6b-v3 Columns Column Description audio Audio sample (if available from source dataset) reference Ground truth transcription prediction Model prediction wer Word Error Rate… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-parakeet-tdt-0.6b-v3-multimed-hard-20260408-1930.audion<1K0 likes8 downloads6mo agoHugging Face08gab1k /mmm_project_parakeetimagen<1K0 likes4 downloads9mo agoHugging Face09gab1k /mmm_project_parakeet_intermdataimagen<1K0 likes3 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.