instinct-org/espeech_podcasts_chunked_nfa_aligned
Forced-Aligned STT Dataset Source dataset: instinct-org/espeech_podcasts_chunked Aligned dataset: instinct-org/espeech_podcasts_chunked_nfa_aligned Rows: 231121 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level CTM spans… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/espeech_podcasts_chunked_nfa_aligned.
01
No card is published for this repository, or it could not be fetched from Hugging Face right now.
