CoolFace
Datasetpublicgated

ivrit-ai/crowd-whatsapp-yi-whisper-training

Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~19h License: other… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.

sourceHugging Faceotherupdated 10mo agoView on Hugging Face
0likes36downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.