datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
THE-BLUEPRINT-FOR-AI-ALIGNMENTlibrispeech-codec-22khzeval-forced-alignment
Hebrew Forced Alignment Evaluation Dataset
Human-verified, word-level time-aligned Hebrew speech clips.
To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was
built. The system lets labelers fix the transcript and align each spoken word to the audio,
down to 1ms precision (though annotators typically work at ~10ms granularity).
The audio samples were gathered by randomly sampling from several of ivrit-ai's larger,
published open datasets. The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-forced-alignment.dogplskillmeiwantodiepodcast-1-test-preprocessed
