datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT)
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/takuM23/multilingual_audio_alignments.emova-alignment-7m
EMOVA-Alignment-7M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment.
This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data.
This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.multilingual_audio_alignments
Multilingual MFA-Aligned Speech Dataset
A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA).
Dataset Description
This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and phonemes.… See the full description on the dataset page: https://huggingface.co/datasets/AAdonis/multilingual_audio_alignments.librispeech-alignments
Dataset Card for Librispeech Alignments
Librispeech with alignments generated by the Montreal Forced Aligner. The original alignments in TextGrid format can be found here
Dataset Details
Dataset Description
Librispeech is a corpus of read English speech, designed for training and evaluating automatic speech recognition (ASR) systems. The dataset contains 1000 hours of 16kHz read English speech derived from audiobooks.
The Montreal Forced Aligner (MFA) was used… See the full description on the dataset page: https://huggingface.co/datasets/gilkeyio/librispeech-alignments.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/new-twi-tts-aligned-ipa.thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.new-twi-tts-aligned-ipa
new-twi-tts-aligned + IPA phonemes
ghanaopendata/new-twi-tts-aligned with a machine-generated IPA phoneme
transcription for every clip, produced with
ghananlpcommunity/ghana-speech-phoneme-asr.
Audio included — this is self-contained, no join with the source dataset needed.
Contents
split
clips
hours
phoneme units
mean units/clip
test
16,140
17.24
663,140
41.1
train
145,258
155.21
5,945,389
40.9
Columns
column
type
meaning… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/new-twi-tts-aligned-ipa.librispeech_mfcc_alignedquran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.shona-bible-bdsc-aligned
Shona Bible Speech Alignment Dataset
Lossless, verse-aligned Shona Bible speech dataset derived from the BDSC
source audio made available by Biblica, Inc. through Open.Bible. This release
contains the complete Bible: 66 books, 1,189 chapters, and 31,284 speech
segments covering approximately 75.55 hours.
Dataset summary
Language: Shona (sna)
Speaker: narrator 1
Speaker sex: male
Books: 66
Clips: 31,284
Audio: approximately 75.55 hours
Audio format: mono 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/manassehzw/shona-bible-bdsc-aligned.emilia-yodas-en-aligned
Emilia-YODAS EN Word-Aligned
Word-level forced-alignment timestamps for the English subset of
amphion/Emilia-Dataset
(Emilia-YODAS split), produced with
Qwen/Qwen3-ForcedAligner-0.6B.
No audio is redistributed — this dataset contains only metadata (IDs,
transcripts already present in Emilia-YODAS, and per-word [start, end]
timestamps). To use it, join on id with the original Emilia-YODAS audio.
Stats
Metric
Value
Utterances
4,516,833
Total audio
11,572.7… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-aligned.librispeech-alignments_clean100
librispeech-alignments_clean100
This is a subset of librispeech-alignments (https://huggingface.co/datasets/gilkeyio/librispeech-alignments) which only includes train_clean_100 and test_clean splits for small experiments and tutorials.
Cite:
@inproceedings{panayotov2015librispeech,
title={Librispeech: an ASR corpus based on public domain audio books},
author={Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev},
booktitle={ICASSP},
year={2015}… See the full description on the dataset page: https://huggingface.co/datasets/ErfanAShams/librispeech-alignments_clean100.libritts-alignedDataset used for loading TTS spectrograms and waveform audio with alignments and a number of configurable "measures", which are extracted from the raw audio.libritts-r-alignedDataset used for loading TTS spectrograms and waveform audio with alignments and a number of configurable "measures", which are extracted from the raw audio.neyshekar-v3-asr-aligned
Neyshekar v3 ASR-Aligned
This is a repaired subset of Neyshekar v3 for Persian ASR work. The public v3
archive contains real audio and real transcripts, but the downloaded
dataset.json filename-to-text mapping does not align for the checked samples.
This export keeps only audio clips whose transcript could be recovered by
matching multiple ASR hypotheses back to the original Neyshekar transcript pool.
It is useful as a curated ASR training/evaluation candidate set, with the… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/neyshekar-v3-asr-aligned.cv_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/cv_chunked
Aligned dataset: instinct-org/cv_chunked_nfa_aligned
Rows: 71097 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_nfa_aligned.CHILDES-Aligned
[!IMPORTANT]
How to access this dataset: the official public release is hosted by TalkBank at
https://talkbank.org/childes/access/Derived/CHILDES-Aligned.html (audio archives +
CSV/JSONL metadata, CC BY-NC-SA 4.0). Please obtain the dataset there.
This Hugging Face copy is retained gated, for internal use; access requests are
approved manually and general requests may be declined — use the TalkBank release instead.
CHILDES-Aligned: Curated Child-Speech Dataset (BEACON)
English… See the full description on the dataset page: https://huggingface.co/datasets/MagicLuke/CHILDES-Aligned.audio_youtube_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/audio_youtube_chunked
Aligned dataset: instinct-org/audio_youtube_chunked_nfa_aligned
Rows: 559484 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_nfa_aligned.medical_asr_aligned_27_04
Medical ASR Aligned Dataset
Aligned Kazakh medical speech dataset from the «ТЕЛЕДӘРІГЕР» (TeleDoctor) TV program on Qazaqstan National Channel.
Dataset Description
Audio-transcript aligned segments of Kazakh-language medical TV broadcasts. Each segment contains the original audio chunk, ASR transcription, human reference transcription, and Character Error Rate (CER).
Only segments with CER < 25% are included.
Stats
Split
Segments
Avg CER
Avg Duration… See the full description on the dataset page: https://huggingface.co/datasets/RakhatM/medical_asr_aligned_27_04.default_voices_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/default_voices_chunked
Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned
Rows: 134236 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.miscellaneous_yt_chunked_speech_restorised_nfa_aligned
miscellaneous_yt_chunked_speech_restorised_nfa_aligned
Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised.
Contents
Parquet shards: 130
Rows: 528,187
Approx hours: 863.88
Audio column: audio with embedded FLAC bytes
Transcript column: transcription
Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments
Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.audiobook_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/audiobook_chunked
Aligned dataset: instinct-org/audiobook_chunked_nfa_aligned
Rows: 1291838 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_nfa_aligned.zy_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/zy_chunked
Aligned dataset: instinct-org/zy_chunked_nfa_aligned
Rows: 534816 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_nfa_aligned.tbp_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/tbp_chunked
Aligned dataset: instinct-org/tbp_chunked_nfa_aligned
Rows: 548483 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/tbp_chunked_nfa_aligned.espeech_podcasts_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/espeech_podcasts_chunked_speech_restorised
Aligned dataset: instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned
Rows: 2467471 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned.yt_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt_chunked_speech_restorised
Aligned dataset: instinct-org/yt_chunked_speech_restorised_nfa_aligned
Rows: 416380 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_nfa_aligned.omni_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/omni_chunked
Aligned dataset: instinct-org/omni_chunked_nfa_aligned
Rows: 2699 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/omni_chunked_nfa_aligned.yt1_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt1_chunked_speech_restorised
Aligned dataset: instinct-org/yt1_chunked_speech_restorised_nfa_aligned
Rows: 261565 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_nfa_aligned.yt2_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt2_chunked_speech_restorised
Aligned dataset: instinct-org/yt2_chunked_speech_restorised_nfa_aligned
Rows: 809612 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_nfa_aligned.yt3_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt3_chunked_speech_restorised
Aligned dataset: instinct-org/yt3_chunked_speech_restorised_nfa_aligned
Rows: 506444 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_nfa_aligned.
