datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
neuro-parakeet-food
neuro-whisper-v1
Dataset Description
This is a synthetic dataset for German medical speech recognition, specifically designed for fine-tuning ASR models on neuro-oncology and neurology terminology. The dataset provides a comprehensive coverage of German medical terminology in the neurology and neuro-oncology domains.
Data Generation
Voice Data: Synthetically generated using Resemble AI Chatterbox TTS
Text Data: Medical text generated with Qwen/Qwen3-30B-A3B… See the full description on the dataset page: https://huggingface.co/datasets/NeurologyAI/neuro-parakeet-food.parakeet-tdt-blind-spots
Blind Spots of nvidia/parakeet-tdt-0.6b-v2
This dataset documents 14 systematically identified blind spots in NVIDIA's parakeet-tdt-0.6b-v2 automatic speech recognition model. The errors span 8 distinct categories and reveal a consistent pattern: the model struggles with inputs outside the distribution of its Western English-centric training data.
Model Under Test
Property
Value
Model
nvidia/parakeet-tdt-0.6b-v2
Parameters
600M
Architecture… See the full description on the dataset page: https://huggingface.co/datasets/TieIncred/parakeet-tdt-blind-spots.
