datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ramanv-tts-all-raw
ramanv-tts-all-raw
Multi-source speech corpus for ASR/STT training. Real human speech across 60+ languages.
linustechtips-transcript-audio
Dataset Card for "linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset contains all the transcripts plus the audio of the different videos of Linus Tech Tips.
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/linustechtips-transcript-audio.SUMM-RENote: if the data viewer is not working, use the "example" subset.
SUMM-RE
The SUMM-RE dataset is a collection of transcripts of French conversations, aligned with the audio signal.
It is a corpus of meeting-style conversations in French created for the purpose of the SUMM-RE project (ANR-20-CE23-0017).
The full dataset is described in Hunter et al. (2024): "SUMM-RE: A corpus of French meeting-style conversations".
Created by: Recording and manual correction of the corpus was… See the full description on the dataset page: https://huggingface.co/datasets/linagora/SUMM-RE.linto-dataset-audio-ar-tn
LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task
This is the first packaged version of the datasets used to train the Linto Tunisian dialect with code-switching STT
(linagora/linto-asr-ar-tn).
Dataset Summary
Dataset composition
Sources
Data Table
Data sources
Content Types
Languages and Dialects
Example use (python)
License
Citations
Dataset Summary
The LinTO DataSet Audio for Arabic Tunisian is a diverse… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn.viet_bud500
Bud500: A Comprehensive Vietnamese ASR Dataset
Introducing Bud500, a diverse Vietnamese speech corpus designed to support ASR research community. With aprroximately 500 hours of audio, it covers a broad spectrum of topics including podcast, travel, book, food, and so on, while spanning accents from Vietnam's North, South, and Central regions. Derived from free public audio resources, this publicly accessible dataset is designed to significantly enhance the work of developers and… See the full description on the dataset page: https://huggingface.co/datasets/linhtran92/viet_bud500.linto-dataset-audio-ar-tn-augmented
LinTO DataSet Audio for Arabic Tunisian Augmented A collection of Tunisian dialect audio and its annotations for STT task
This is the augmented datasets used to train the Linto Tunisian dialect with code-switching STT linagora/linto-asr-ar-tn.
Dataset Summary
Dataset composition
Sources
Content Types
Languages and Dialects
Example use (python)
License
Citations
Dataset Summary
The LinTO DataSet Audio for Arabic Tunisian Augmented is a dataset that builds on LinTO… See the full description on the dataset page: https://huggingface.co/datasets/linagora/linto-dataset-audio-ar-tn-augmented.LinguaLibre-word-retrieval
Lingua Libre spoken-word retrieval (MTEB)
Single words read aloud by volunteers, paired with the written word, across many
languages.
Recordings come from Lingua Libre, a Wikimedia project, hosted on Wikimedia
Commons, which is free by site policy. Published as cc-by-sa-4.0. Audio is 16 kHz
Opus. Bare punctuation and read sentences are excluded, and each word is kept once.
Built by scripts/data/lingua_libre/create_data.py in the MTEB repo.
audios-lingala-annotatees
Annotated Lingala Dataset – Full Version
Description
This dataset gathers annotated Lingala audio data, intended for open-source automatic speech recognition (ASR) research and for fine-tuning Whisper-type models.
It includes:
the original audio files (viewable directly in the Hugging Face viewer)
text transcriptions
Mel spectrograms
tokenized labels
Overall statistics
Metric
Value
Total volume
5 h 0 min 18 s
Number of audio segments… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees.Lingala_100hrs
Lingala 100hrs
110.7 hours (23,539 rows) of Lingala speech with transcriptions, aggregated
from three publicly available CC-BY-4.0 corpora for ASR research.
Composition
Counts from a full-pass audit on 2026-07-09:
Source
Upstream location
Rows
Splits
AfriVoice (Lingala)
https://huggingface.co/datasets/DigitalUmuganda/AfriVoice
17,544
train (16,144), validation (915), test (485)
LRSC (Lingala Read Speech Corpus)… See the full description on the dataset page: https://huggingface.co/datasets/KasuleTrevor/Lingala_100hrs.audios-lingala-annotatees-v2
Annotated Lingala Audio — canonical corpus
Annotated Lingala speech for open automatic speech recognition research and for
fine-tuning speech models.
This release is a full reconstruction of the corpus from its source
recordings and annotations. It supersedes
Congo-digital-service/audios-lingala-annotatees,
which is deprecated — see Relationship to the previous release below.
What this dataset contains
Each row is one annotated speech segment, carrying the audio… See the full description on the dataset page: https://huggingface.co/datasets/Congo-digital-service/audios-lingala-annotatees-v2.jam-alt-lines
Jam-ALT Lines
Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset.
Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit.
[!tip]
See the Jam-ALT project website for details and the JamendoLyrics community for related datasets.
Dataset flavors
Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/cublya/jam-alt-lines.fleurs
FLEURS
Fleurs is the speech version of the FLoRes machine translation benchmark.
We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages.
Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is
used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/linq1005/fleurs.omnilingual-asr-corpus
Meta Omnilingual ASR Corpus
The Omnilingual ASR Corpus is a collection of spontaneous speech recordings and their transcriptions for 348 under-served languages. The corpus was collected as part of Meta FAIR’s Omnilingual ASR project (blog, model, paper) for the purposes of training automatic speech recognition (ASR) and spoken language identification models.
Data schema
{
`language`: "lij_Latn",
`iso_639_3`: "lij",
`iso_15924`: "Latn",
`glottocode`:… See the full description on the dataset page: https://huggingface.co/datasets/lindonghello/omnilingual-asr-corpus.jam-alt-lines
Jam-ALT Lines
Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset.
Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit.
[!tip]
See the Jam-ALT project website for details and the JamendoLyrics community for related datasets.
Dataset flavors
Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/jamendolyrics/jam-alt-lines.Lingua_Libre_br
[!NOTE]
Dataset origin: https://lingualibre.org/LanguagesGallery/
Description
1h 0min 2s d'audio.Les fichiers .ogg ont été convertis en .mp3.L'extraction date de janvier 2026.
astryd_lindgren_braty_lvinae_sertsa_all
Браты Ільвінае сэрца
Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian)
Аўдыё нарэзана з арыгінальнага запісу ў зыходнай частаце дыскрэтызацыі (native), мона, фрагменты да 30 секунд.
Частка калекцыі Belarusian Audiobooks (native).
Радкоў у датасеце
2,116
Працягласць
6 гадз 31 хв
Частата дыскрэтызацыі
44100 Hz
Каналы
мона
Даўжыня фрагмента
да 30 с
Структура
Кожны радок змяшчае:
audio — аўдыёфрагмент (native SR, мона… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_all.whisper-transcripts-linustechtips
Dataset Card for "whisper-transcripts-linustechtips"
Dataset Summary
This dataset is created by applying whisper to the videos of the Youtube channel Linus Tech Tips. The dataset was created a medium size whisper model.
Languages
Language: English
Dataset Structure
The dataset
Data Fields
The dataset is composed by:
id: Id of the youtube video.
channel: Name of the channel.
channel_id: Id of the youtube channel.
title: Title given to… See the full description on the dataset page: https://huggingface.co/datasets/Whispering-GPT/whisper-transcripts-linustechtips.lingala_real_eval_benchmark_croped
Lingala Real Eval Benchmark Cropped
Small cropped real-world Lingala audio benchmark for testing BLI ASR 0.
The dataset contains short audio clips cropped from longer real-world files, covering different domains such as news, catechesis, comedy, cartoon and interview speech.
This dataset is intended for quick qualitative ASR testing and human review. It is not a training dataset.
Lingala-Speech-Dataset
Lingala Dataset Metadata
Field
Value
📜 License
CC BY-NC-ND 4.0
🎯 Task Categories
Automatic Speech Recognition
🌍 Language
Lingala (ln)
🏷️ Tags
Audio, Speech, Speech Recognition, ML, Machine, Machine Learning, Lingala
📦 Size Category
n < 1K
astryd_lindgren_braty_lvinae_sertsa_output_original
Браты Ільвінае сэрца — арыгінальнае аўдыё
Аўтар / Author: Астрыд ЛіндгрэнМова / Language: Беларуская (Belarusian)
Арыгінальнае аўдыё без апрацоўкі, захаванае ў зыходнай якасці.
Частка калекцыі Ministerskija —
корпус беларускіх аўдыёкніг.
Апрацаваная версія (сегменты ~15 с, выраўнаваная транскрыпцыя):
astryd_lindgren_braty_lvinae_sertsa_output
Доўгасць аўдыё
7h06m
Радкоў у датасеце
1,835
Структура
Кожны радок змяшчае:
audio — арыгінальны… See the full description on the dataset page: https://huggingface.co/datasets/fosters/astryd_lindgren_braty_lvinae_sertsa_output_original.jam-alt-lines
Jam-ALT Lines
Jam-ALT Lines is a line-level version of the Jam-ALT lyrics transcription dataset.
Unlike Jam-ALT, this dataset contains one audio segment for each lyrics line, facilitating research that considers each line as a separate unit.
[!tip]
See the Jam-ALT project website for details and the JamendoLyrics community for related datasets.
Dataset flavors
Lyrics lines may overlap in time, which makes it impossible to have a one-to-one correspondence between… See the full description on the dataset page: https://huggingface.co/datasets/slaycreep/jam-alt-lines.
