CoolFace
Datasetpublicgated

instinct-org/yt4_chunked_speech_restorised_nfa_aligned

Forced-Aligned STT Dataset Source dataset: instinct-org/yt4_chunked_speech_restorised Aligned dataset: instinct-org/yt4_chunked_speech_restorised_nfa_aligned Rows: 965192 successfully aligned rows This dataset adds NeMo Forced Aligner metadata for STT training and timestamp quality control. Source rows that did not produce usable CTM alignments are excluded from the published data. Added columns: nfa_token_alignments: NeMo token/subword CTM spans nfa_word_alignments: word-level… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/yt4_chunked_speech_restorised_nfa_aligned.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes3downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
instinct-org/yt4_chunked_speech_restorised_nfa_aligned · CoolFace