datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human-robot-conversation-german
Human-Robot Dataset
The dataset comprises 660+ hours of audio recordings across 20,000+ files for human-robot interactions in the German language. It captures authentic dialogues between humans and artificial conversational agents, specifically designed for training language models and advancing speech recognition systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in speech recognition, natural language processing, and… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/human-robot-conversation-german.dwesui-grupa-1-neurologia
NeuroSpeechPL
Publiczny eksport HuggingFace zawiera wyłącznie redystrybuowalne audio source=natural. Wiersze TTS są celowo wyłączone z publicznego zbioru danych, ponieważ ich source_license zabrania redystrybucji audio. Pełna lokalna ewaluacja opisana w raporcie korzystała zarówno z nagrań naturalnych, jak i TTS.
Repozytorium zbioru danych HF: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Repozytorium kodu:… See the full description on the dataset page: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia.2026-dwesui-g01-neurologia
DWESUI 2026 - Grupa 1 - neurologia
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 1 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Domena: neurologia
Licencja zrodla: nagrania YouTube CC-BY + synteza TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował zbiór… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g01-neurologia.dwesui-grupa-2-kulinarna
G2-Polish-Culinary-ASR-Evaluation-Corpus
Korpus do ewaluacji systemow ASR jezyka polskiego (domena kulinarna) stworzony
w ramach warsztatow Ewaluacja Systemow Rozpoznawania Mowy (UAM WMI, edycja 2026,
zespol 2). Publikowany podzbior to mowa naturalna z wideo kulinarnych YouTube
(licencja CC-BY) - sluzy do badania odpornosci ASR na szum kuchenny oraz dopasowania
domenowego do specjalistycznego slownictwa (zapozyczenia, miary, liczby).
Pelny eksperyment ewaluacyjny zespolu… See the full description on the dataset page: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna.German-Speech-Dataset
🎧 German Speech Dataset
The German Speech Dataset is a high-quality speech audio dataset designed to provide structured and scalable audio data for advanced AI and machine learning systems. It includes 142 hours of audio data across 768 files, delivered in MP3 and WAV formats, with a total size of 327 MB. This carefully curated audio dataset ensures diverse and representative voice data, with 53% male and 47% female speakers, and a balanced age distribution ranging from 18 to 50+… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/German-Speech-Dataset.german-speech-recognition-dataset
German Speech Dataset for recognition task
Dataset comprises 431 hours of telephone dialogues in German, collected from 590+ native speakers across various topics and domains, achieving an impressive 95% sentence accuracy rate. It is designed for research in automatic speech recognition (ASR) systems.
By utilizing this dataset, researchers and developers can advance their understanding and capabilities in transcribing audio, and natural language processing (NLP). - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/german-speech-recognition-dataset.Greek-Speech-Dataset
🎧 Greek Speech Dataset
The Greek Speech Dataset is a structured and high-quality speech audio dataset designed to support modern AI systems that rely on diverse audio data and multilingual voice data. It contains 184 hours of recordings distributed across 592 files, stored in MP3 and WAV formats, with a total size of 330 MB. This carefully curated audio dataset delivers balanced representation across speakers, including 49% female and 51% male participants, and a broad age range… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Greek-Speech-Dataset.GSMAEthioTelecomBench
GSMA EthioTelecomBench: Amharic ASR Benchmark for Telecom Domain
Overview
GSMA EthioTelecomBench is a comprehensive benchmark for evaluating Automatic Speech Recognition (ASR) systems on Amharic speech, with a focus on telecom customer service conversations. This dataset contains evaluation results from 12 models across 7 evaluation splits.
Audio Conditions
The benchmark evaluates models across different audio conditions:
Column Name… See the full description on the dataset page: https://huggingface.co/datasets/SaarAI/GSMAEthioTelecomBench.2026-dwesui-g02-kulinarna
DWESUI 2026 - Grupa 2 - kulinarna (PIEROGA)
Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu
Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2026, tryb dzienny.
Zespol (atrybucja): Grupa 2 (DWESUI 2026)
Zrodlo oryginalne: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna
Domena: kulinarna
Licencja zrodla: nagrania YouTube CC-BY/CC-BY-SA + TTS
Status: kopia publiczna w organizacji kursowej (zespół opublikował… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2026-dwesui-g02-kulinarna.Georgian-Speech-Dataset
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Georgian (ka)
🏷️ Tags
Audio, Speech, Speech Recognition, Georgian, ML, Machine, Machine Learning
📦 Size Category
n < 1K
german-speech-recognition-dataset
German Telephone Dialogues Dataset - 431 Hours
Dataset comprises 431 hours of high-quality audio recordings from 590+ native German speakers, featuring telephone dialogues across diverse topics and domains. With a 95% sentence accuracy rate, this essential dataset is ideal for training and evaluating German speech recognition systems. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of telephone dialogues in German for training… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/german-speech-recognition-dataset.human-robot-conversation-german
Human-Robot Conversation Dataset (German) - 660+ Hours
Dataset (German) contains 660+ hours of audio featuring dialogues between AI and a human in German across 20,000 recordings. The dataset supports conversational AI, speech recognition, and human-robot interaction research, with short M4A audio files (up to 2 minutes) and structured metadata for model training. - Get the data
Dataset characteristics:
Characteristic
Data
Description
Audio of dialogues between AI… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/human-robot-conversation-german.Gujarati-Speech-Dataset
🎧 Gujarati Speech Dataset
The Gujarati Speech Dataset is a high-quality multilingual speech audio dataset designed to support advanced AI systems that rely on diverse audio data and reliable voice data. It comprises 122 hours of recordings across 763 files, provided in MP3 and WAV formats, with a total size of 353 MB. This structured audio dataset ensures balanced representation with 55% female and 45% male speakers, covering an age range from 18 to 50+ years. The dataset language… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Gujarati-Speech-Dataset.gemma-4-public-bench-eval
Gemma 4 (e4b & 12b) — Public ASR Benchmark (Decoded Hypotheses + WER)
Decoded transcripts and word-level error metrics from Gemma 4 Unified
(the encoder-free, natively audio-capable models) run as automatic speech
recognition (ASR) systems on three standard English test sets. Two models are
evaluated — gemma4:e4b (8B params) and gemma4:12b. Everything was
produced locally with ollama; the evaluation tool
(eval_asr.py) is included so the numbers are fully reproducible.
Gemma 4… See the full description on the dataset page: https://huggingface.co/datasets/huckiyang/gemma-4-public-bench-eval.zwesui-grupa-5-it-ai
Wykorzystanie ASR do transkrypcji polskich nagrań o tematyce AI
Korpus do ewaluacji systemów ASR języka polskiego stworzony w ramach warsztatów
Ewaluacja Systemów Rozpoznawania Mowy (UAM WMI, edycja 2026, zespół 5).
Zbiór powstał jako część kursu - publikujemy go publicznie, żeby inni badacze
polskiego ASR mogli z niego korzystać i porównywać wyniki na wspólnym benchmarku.
Cel i pytania badawcze
Cel główny:
Porównanie jakości 3 systemów ASR dla spontanicznej… See the full description on the dataset page: https://huggingface.co/datasets/slapekm/zwesui-grupa-5-it-ai.Gleason
TalkBank CHILDES Gleason Raw
This repository mirrors the Gleason corpus from TalkBank CHILDES as raw files
for reproducible local workflows.
Source corpus page: https://talkbank.org/childes/access/Eng-NA/Gleason.html
DOI: doi:10.21415/T5101R
HF repo: MagicLuke/Gleason
Contents
transcripts/Gleason/{Mother,Father,Dinner}/*.cha
media/{Mother,Father,Dinner}/*.mp3
raw/Gleason.zip
metadata.json
metadata_from_cha.json
recordings_from_cha.csv
Data config
The… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/Gleason.Galician-Speech-Dataset
🎧 Galician Speech Dataset
The Galician Speech Dataset is a high-quality speech audio dataset designed to provide structured and diverse audio data for AI-driven voice technologies. It includes 158 hours of audio data across 718 files, delivered in MP3 and WAV formats, with a total size of 276 MB. This well-organized audio dataset ensures balanced and representative voice data, with 48% female and 52% male speakers, and a wide age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Galician-Speech-Dataset.
