CoolFace
Datasetpublicgated

ivrit-ai/audio-v2-transcripts

Overview This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on May 18th, 2025. You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt. All files were transcribed using the process.py pipeline, performing: Frame-level VAD Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.

sourceHugging Faceotherupdated 10mo agoView on Hugging Face
1likes882downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.

ivrit-ai/audio-v2-transcripts · CoolFace