datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
BIRDeep_AudioAnnotations
BIRDeep Audio Annotations
The BIRDeep Audio Annotations dataset is a collection of bird vocalizations from Doñana National Park, Spain. It was created as part of the BIRDeep project, which aims to optimize the detection and classification of bird species in audio recordings using deep learning techniques. The dataset is intended for use in training and evaluating models for bird vocalization detection and identification.
The research code and further information is available at… See the full description on the dataset page: https://huggingface.co/datasets/GrunCrow/BIRDeep_AudioAnnotations.indic-multilingual-asr
Indic Multilingual ASR Dataset
A multilingual ASR dataset covering 13 major Indian languages with 1.1M+ samples.
Usage
from datasets import load_dataset
ds = load_dataset("grushaaaaa/indic-multilingual-asr", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
llama-omni-speech-instruct
Llama3.2 Omni Speech Instruct Dataset
This dataset is created for the sole purpose of enhancing the LLM capability to become multi-modals. This dataset has speech instruction
that a model could use to learn and produce the output thus allowing the model to overcome only text input and extends it capabilities
towards processing speech command as well.
Dataset Details
Dataset Description
This dataset can be used to train an LLM model to allow adaptibility in… See the full description on the dataset page: https://huggingface.co/datasets/gruhit-patel/llama-omni-speech-instruct.tts-indian
TTS Indian Languages Dataset
Speech dataset for Text-to-Speech covering 6 Indian languages, collected and processed from YouTube.
Languages & Speakers
Speaker
Language
Gender
monihara_bengali
Bengali
Male
munir_kashmiri
Kashmiri
Male
nandini_gujarati
Gujarati
Female
sansri_kannada
Kannada
Female
tamil_pokkisham
Tamil
Male
teluguM
Telugu
Male
Pipeline
Audio was collected and processed through these stages:
YouTube Download — yt-dlp… See the full description on the dataset page: https://huggingface.co/datasets/grushaaaaa/tts-indian.libritts_r_train33kdwesui-grupa-1-neurologia
NeuroSpeechPL
Publiczny eksport HuggingFace zawiera wyłącznie redystrybuowalne audio source=natural. Wiersze TTS są celowo wyłączone z publicznego zbioru danych, ponieważ ich source_license zabrania redystrybucji audio. Pełna lokalna ewaluacja opisana w raporcie korzystała zarówno z nagrań naturalnych, jak i TTS.
Repozytorium zbioru danych HF: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia
Repozytorium kodu:… See the full description on the dataset page: https://huggingface.co/datasets/JankesTNJ/dwesui-grupa-1-neurologia.dwesui-grupa-2-kulinarna
G2-Polish-Culinary-ASR-Evaluation-Corpus
Korpus do ewaluacji systemow ASR jezyka polskiego (domena kulinarna) stworzony
w ramach warsztatow Ewaluacja Systemow Rozpoznawania Mowy (UAM WMI, edycja 2026,
zespol 2). Publikowany podzbior to mowa naturalna z wideo kulinarnych YouTube
(licencja CC-BY) - sluzy do badania odpornosci ASR na szum kuchenny oraz dopasowania
domenowego do specjalistycznego slownictwa (zapozyczenia, miary, liczby).
Pelny eksperyment ewaluacyjny zespolu… See the full description on the dataset page: https://huggingface.co/datasets/s479246/dwesui-grupa-2-kulinarna.alpaca_speech_instructkannada-emotional-tts
Grusha Kannada Emotional TTS
A single-speaker Kannada (ಕನ್ನಡ) speech dataset for text-to-speech (TTS) and
expressive / emotional speech synthesis, recorded by a single female speaker
(grusha_kannada). Every utterance is labelled with one of four emotions —
neutral, happy, sad, angry — making the corpus suitable for training expressive
and emotion-controllable TTS models, as well as speech-emotion classification.
Dataset at a glance
Language
Kannada… See the full description on the dataset page: https://huggingface.co/datasets/grushaaaaa/kannada-emotional-tts.zwesui-grupa-4-medyczn
Polski korpus ASR — grupa 4 (mowa medyczna, trzustka i obrazowanie)
Korpus referencyjny do ewaluacji systemów ASR w języku polskim, przygotowany w ramach warsztatów
Ewaluacja Systemów Rozpoznawania Mowy (UAM WMI, edycja 2026, grupa 4).
Domena: mowa medyczna — opisy obrazowania diagnostycznego, objawów i badań laboratoryjnych
związanych z chorobami trzustki. Zbiór łączy segmenty z materiału edukacyjnego (YouTube, CC-BY)
oraz uzupełniające nagrania syntetyczne TTS, transkrybowane… See the full description on the dataset page: https://huggingface.co/datasets/TCA/zwesui-grupa-4-medyczn.Gru_150zwesui-grupa-5-it-ai
Wykorzystanie ASR do transkrypcji polskich nagrań o tematyce AI
Korpus do ewaluacji systemów ASR języka polskiego stworzony w ramach warsztatów
Ewaluacja Systemów Rozpoznawania Mowy (UAM WMI, edycja 2026, zespół 5).
Zbiór powstał jako część kursu - publikujemy go publicznie, żeby inni badacze
polskiego ASR mogli z niego korzystać i porównywać wyniki na wspólnym benchmarku.
Cel i pytania badawcze
Cel główny:
Porównanie jakości 3 systemów ASR dla spontanicznej… See the full description on the dataset page: https://huggingface.co/datasets/slapekm/zwesui-grupa-5-it-ai.Frankensteins-Monster-Gruntsjapanese-medical-tts-datasetgrupa-4-ZWESUI0
Dataset Card: Group 4 - ZWESUI0 (Houseplants)
Dataset Summary
A dataset created for the final project of the Speech Recognition Systems Evaluation Workshop (ZWESUI). The corpus focuses on the evaluation of ASR systems in the specific domain of houseplants and botany.
The dataset contains 523 audio segments originating from two main sources with different acoustic and linguistic characteristics:
Spontaneous speech (YouTube): Excerpts from educational and tutorial… See the full description on the dataset page: https://huggingface.co/datasets/mszulcc/grupa-4-ZWESUI0.gruivan-bunin-grugan-uladzimir-ragautsou
Груган
Metadata
Author: Іван Бунін
Title: Груган
Narrator: Уладзімір Рагаўцоў
Source Group: Аўдыёкнігі
Source: БЛР#аўдыякніга
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size:… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/ivan-bunin-grugan-uladzimir-ragautsou.uladzimir-karatkevich-grubae-i-laskavae
Грубае і ласкавае
Metadata
Author: Уладзімір Караткевіч
Title: Грубае і ласкавае
Narrator:
Source Group: Аўдыёкнігі
Source:
Notes
The original audio files are preserved as-is:
no conversion;
no re-encoding;
no filename changes inside each split folder, except removing one common top-level archive folder when present.
To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders.
Target maximum split size:… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/uladzimir-karatkevich-grubae-i-laskavae.
