datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-canary-asr-0324
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 16.1, Voxpopuli und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-canary-asr-0324.CanaryAura
Dataset Card for "Canary Aura"
This is a dataset for...
eval-canary-1b-v2-eka-hard-20260408-1921
Evaluation Results: canary-1b-v2
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/canary-1b-v2
39.85%
22.43%
Source Data
Evaluation Dataset: Trelis/eka-hard
Model Evaluated: nvidia/canary-1b-v2
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for this sample
cer… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-canary-1b-v2-eka-hard-20260408-1921.eval-canary-1b-v2-medical-terms-2025-20260408-1926
Evaluation Results: canary-1b-v2
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/canary-1b-v2
10.11%
3.26%
Source Data
Evaluation Dataset: Trelis/medical-terms-2025
Model Evaluated: nvidia/canary-1b-v2
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for this sample… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-canary-1b-v2-medical-terms-2025-20260408-1926.eval-canary-1b-v2-multimed-hard-20260408-1931
Evaluation Results: canary-1b-v2
Evaluation results from Whisper model evaluation.
Summary
Model
WER
CER
nvidia/canary-1b-v2
15.02%
9.27%
Source Data
Evaluation Dataset: Trelis/multimed-hard
Model Evaluated: nvidia/canary-1b-v2
Columns
Column
Description
audio
Audio sample (if available from source dataset)
reference
Ground truth transcription
prediction
Model prediction
wer
Word Error Rate for this sample
cer… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/eval-canary-1b-v2-multimed-hard-20260408-1931.
