CoolFace
Datasetpublicgated

ivrit-ai/crowd-whatsapp-yi-whisper-training

Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~19h License: other… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.

sourceHugging Faceotherupdated 10mo agoView on Hugging Face
0likes36downloads
settings

This repository belongs to ivrit-ai on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namecrowd-whatsapp-yi-whisper-training
visibilitypublic
licenceother
gatedyes
ownerivrit-ai
Account settings
ivrit-ai/crowd-whatsapp-yi-whisper-training · CoolFace