CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ivrit-ai /knesset-committeesgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) committee sessions as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings of committee sessions, alongside human-generated protocols. We extract the audio stream, abd produce weakly time stamped segmentation of the protocol text (we… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-committees.automatic-speech-recognition3 likes4.5k downloads4mo agoHugging Face02ivrit-ai /audio-v2gatedThis dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It wa released on April 20th, 2025. You can find the full list of sources in this dataset under the dataset's sources.txt. Paper: https://arxiv.org/abs/2307.08720 If you use our datasets, the following quote is preferable: @misc{marmor2023ivritai, title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development}, author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2.audio-classification10K<n<100K3 likes2k downloads8mo agoHugging Face03ivrit-ai /audio-v2-transcriptsgated Overview This dataset provides full, machine-generated transcriptions for the entire audio-v2 dataset, containing >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on May 18th, 2025. You can find the full list of sources in this dataset under the audio-v2 dataset's sources.txt. All files were transcribed using the process.py pipeline, performing: Frame-level VAD Machine transcription using ivrit.ai's whisper-large-v3-turbo engine with the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-transcripts.audio-classification10K<n<100K1 likes882 downloads10mo agoHugging Face04ivrit-ai /knesset-plenums-whisper-traininggated Dataset Card for ivrit.ai - Knesset Plenums Whisper Training This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset. This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less. Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription. The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audiotext-to-speech100K<n<1M3 likes420 downloads10mo agoHugging Face05ivrit-ai /crowd-recitalgated About This dataset was created by crowd-sourced recording sessions in Hebrew as part of the ivrit.ai Crowd Recital project. Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read. Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below). The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital.automatic-speech-recognition1 likes120 downloads10mo agoHugging Face06ivrit-ai /eval-whatsappgated Dataset Card for ivrit.ai Whatsapp Eval Evaluation dataset of Hebrew Whatsapp voice-messages, expert-transcribed. Dataset Details Dataset Description This dataset containeד Whatsapp hebrew voice recordings gathered around April 2025. The recordings are of volunteer native hebrew speakers using consumer devices in natural environments. Each recording is by a single speaker about a random topic of their choice spoken in a non-scripted natural manner. Each… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-whatsapp.audioautomatic-speech-recognitionn<1K0 likes120 downloads10mo agoHugging Face07ivrit-ai /knesset-plenumsgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps. We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts). The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.audioautomatic-speech-recognition1K<n<10K3 likes108 downloads10mo agoHugging Face08ivrit-ai /crowd-recital-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital Dataset Details Dataset Description License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. - Full license: https://www.ivrit.ai/en/the-license/ - FAQs: https://www.ivrit.ai/en/license-faqs/ Dataset Structure Data Fields Each example in the dataset contains: audio: An audio column containing: bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.audiotext-to-speech1K<n<10K3 likes87 downloads10mo agoHugging Face09ivrit-ai /crowd-recital-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~78h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.audiotext-to-speech10K<n<100K0 likes72 downloads10mo agoHugging Face10ivrit-ai /audio-v2-opusgatedThis dataset contains >20k hours of Hebrew audio, all licensed under the ivrit.ai v1 license. It was released on April 20th, 2025. You can find the full list of sources in this dataset under the dataset's sources.txt. Paper: https://arxiv.org/abs/2307.08720 If you use our datasets, the following quote is preferable: @misc{marmor2023ivritai, title={ivrit.ai: A Comprehensive Dataset of Hebrew Speech for AI Research and Development}, author={Yanir Marmor and Kinneret Misgav and Yair… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-v2-opus.audioaudio-classification10K<n<100K0 likes60 downloads10mo agoHugging Face11ivrit-ai /crowd-recital-yigated About This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project. Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read. Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below). The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.audioautomatic-speech-recognition1K<n<10K0 likes46 downloads10mo agoHugging Face12ivrit-ai /crowd-whatsapp-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~19h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.audiotext-to-speech1K<n<10K0 likes31 downloads10mo agoHugging Face13ivrit-ai /crowd-whatsapp-yigated About This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project. Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot. Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below). The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.audioautomatic-speech-recognition1K<n<10K0 likes24 downloads10mo agoHugging Face14ivrit-ai /crowd-transcribe-v4gatedNote: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai. crowd-transcribe-v4 This is ivrit.ai's 4th crowd-sourced transcribed dataset release. It contains over 250 hours of volunteer-transcribed data, randomly selected from our audio-vad dataset of over 10,000 hours. License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. Full license:… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-transcribe-v4.audioautomatic-speech-recognition100K<n<1M2 likes9 downloads10mo agoHugging Face15ivrit-ai /eval-forced-alignmentgated Hebrew Forced Alignment Evaluation Dataset Human-verified, word-level time-aligned Hebrew speech clips. To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was built. The system lets labelers fix the transcript and align each spoken word to the audio, down to 1ms precision (though annotators typically work at ~10ms granularity). The audio samples were gathered by randomly sampling from several of ivrit-ai's larger, published open datasets. The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-forced-alignment.audioautomatic-speech-recognitionn<1K1 likes9 downloads19h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.