datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.audio_youtube_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/audio_youtube_chunked
Aligned dataset: instinct-org/audio_youtube_chunked_nfa_aligned
Rows: 559484 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audio_youtube_chunked_nfa_aligned.default_voices_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/default_voices_chunked
Aligned dataset: instinct-org/default_voices_chunked_nfa_aligned
Rows: 134236 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/default_voices_chunked_nfa_aligned.miscellaneous_yt_chunked_speech_restorised_nfa_aligned
miscellaneous_yt_chunked_speech_restorised_nfa_aligned
Public, manually gated NFA-aligned Uzbek speech dataset derived from instinct-org/miscellaneous_yt_chunked_speech_restorised.
Contents
Parquet shards: 130
Rows: 528,187
Approx hours: 863.88
Audio column: audio with embedded FLAC bytes
Transcript column: transcription
Alignment columns: nfa_token_alignments, nfa_word_alignments, nfa_segment_alignments, nfa_character_alignments
Access And Use… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/miscellaneous_yt_chunked_speech_restorised_nfa_aligned.audiobook_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/audiobook_chunked
Aligned dataset: instinct-org/audiobook_chunked_nfa_aligned
Rows: 1291838 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/audiobook_chunked_nfa_aligned.zy_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/zy_chunked
Aligned dataset: instinct-org/zy_chunked_nfa_aligned
Rows: 534816 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans
nfa_segment_alignments:… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/zy_chunked_nfa_aligned.espeech_podcasts_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/espeech_podcasts_chunked_speech_restorised
Aligned dataset: instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned
Rows: 2467471 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_speech_restorised_nfa_aligned.yt_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt_chunked_speech_restorised
Aligned dataset: instinct-org/yt_chunked_speech_restorised_nfa_aligned
Rows: 416380 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt_chunked_speech_restorised_nfa_aligned.yt1_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt1_chunked_speech_restorised
Aligned dataset: instinct-org/yt1_chunked_speech_restorised_nfa_aligned
Rows: 261565 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt1_chunked_speech_restorised_nfa_aligned.yt2_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt2_chunked_speech_restorised
Aligned dataset: instinct-org/yt2_chunked_speech_restorised_nfa_aligned
Rows: 809612 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt2_chunked_speech_restorised_nfa_aligned.yt3_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt3_chunked_speech_restorised
Aligned dataset: instinct-org/yt3_chunked_speech_restorised_nfa_aligned
Rows: 506444 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt3_chunked_speech_restorised_nfa_aligned.yt4_chunked_speech_restorised_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/yt4_chunked_speech_restorised
Aligned dataset: instinct-org/yt4_chunked_speech_restorised_nfa_aligned
Rows: 965192 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_nfa_aligned.education_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/education_chunked
Aligned dataset: instinct-org/education_chunked_nfa_aligned
Rows: 386961 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/education_chunked_nfa_aligned.espeech_podcasts_chunked_nfa_aligned
Forced-Aligned STT Dataset
Source dataset: instinct-org/espeech_podcasts_chunked
Aligned dataset: instinct-org/espeech_podcasts_chunked_nfa_aligned
Rows: 231121 successfully aligned rows
This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality
control. Source rows that did not produce usable CTM alignments are excluded from the
published data.
Added columns:
nfa_token_alignments: NeMo token/subword CTM spans
nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_nfa_aligned.
