datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/fosple/german-asr-mixed-whisper.asr-german-mixed
Dataset Beschreibung
Allgemeine Informationen
Dieser Datensatz ist eine Kombination aus drei verschiedenen Quellen für die deutsche Sprache: Commonvoice 17.0 und Multilingual librispeech. Die Daten wurden gefiltert, normalisiert und grammatikalisch korrigiert.
Die drei Datensätze wurden erneut transkribiert und mit den entsprechenden Audio-Daten abgeglichen, um genaue Transkriptionen zu erhalten. Anschließend wurde ein Abgleich mit den Originaltranskripten durchgeführt… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/asr-german-mixed.german-asr-mixed-whisper
Dataset Card
Dataset Sources and Licensing
This dataset is a mixture of several German and multilingual speech datasets. For each dataset, the license of the original author applies. Please consult the linked sources for detailed licensing information and terms of use.
Dataset Name
Original Source / Author
Link
TUDA-De German Speech Corpus
LT Group at UHH / TU Darmstadt
https://huggingface.co/datasets/uhhlt/Tuda-De
Mozilla Common Voice
Mozilla Foundation… See the full description on the dataset page: https://huggingface.co/datasets/flozi00/german-asr-mixed-whisper.mixed_shona_datasetASR-GERMAN-MIXED-TEST
Dataset Beschreibung
Dieser Datensatz und die Beschreibung wurde von flozi00/asr-german-mixed übernommen und nur der Test-Split hier hochgeladen, da Hugging Face native erst einmal alle Splits herunterlädt. Für eine Evaluation von Speech-to-Text Modellen ist ein Download von 136 GB allerdings etwas zeit- & speicherraubend, weshalb wir hier nur den Test-Split für Evaluationen anbieten möchten. Die Arbeit und die Anerkennung sollten deshalb weiter bei primeline & flozi00 für die… See the full description on the dataset page: https://huggingface.co/datasets/avemio/ASR-GERMAN-MIXED-TEST.uk-en-code-mixed-asr-2h
uk-en-code-mixed-asr-2h
A 2-hour dataset of Ukrainian-English code-mixed speech for automatic speech recognition, recorded by a single male speaker across 16 speaking-style personas covering software engineering domains (.NET, React, DevOps, project management).
Samples: 448
Total duration: 2h 6m 35s
Language: 446 mixed (code-mixed) + 2 uk (monolingual)
Speaker: 1 male voice, 16 personas (varied pacing, formality, topic; 15 map to Valera, 1 to ValeraAlt)
Audio format: OGG Opus… See the full description on the dataset page: https://huggingface.co/datasets/vnikitin/uk-en-code-mixed-asr-2h.
