CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ivrit-ai /crowd-transcribe-v5gated License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. Full license: https://www.ivrit.ai/en/the-license/ FAQs: https://www.ivrit.ai/en/license-faqs/ audio100K<n<1M13 likes1k downloads10mo agoHugging Face02ivrit-ai /VoxKnessetgated VoxKnesset Voice recordings of Israeli politicians from Knesset proceedings, annotated with speaker age and demographic metadata. Dataset Summary Total hours (longitudinal subset): 2,307 Plenary sessions: ~1,550 Unique speakers: 393 Members of Knesset Language: Hebrew Recording years: 2009–2025 (16 years) Maximum span per speaker: 15 years Median span per speaker: 3.4 years Speakers with >10 years coverage: 47 (12%) Age range: 28–81 years Split Samples… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/VoxKnesset.audio10K<n<100K3 likes899 downloads7mo agoHugging Face03ivrit-ai /jbdtabular10M<n<100M0 likes633 downloads3mo agoHugging Face04ivrit-ai /knesset-plenums-whisper-traininggated Dataset Card for ivrit.ai - Knesset Plenums Whisper Training This is a whisper-formatted version of the ivrit.ai Knesset Plenums dataset. This dataset was created by splitting long audio recordings, along with their respective transcriptions, into audio slices of 30 seconds or less. Each such slice represents one or more consecutive segments, along with timestamp token data and the previous slice's transcription. The code for this dataset preparation process is available on the… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums-whisper-training.audiotext-to-speech100K<n<1M3 likes420 downloads10mo agoHugging Face05ivrit-ai /audio-vadgatedivrit.ai is a database of Hebrew audio and text content. audio-base contains the raw, unprocessed sources. audio-vad contains audio snippets generated by applying Silero VAD (https://github.com/snakers4/silero-vad) to the base dataset. audio-transcripts contains transcriptions for each snippet in the audio-vad dataset. The audio-base dataset contains data from the following sources: Geekonomy (Podcast, https://geekonomy.net) HaCongress (Podcast, https://hacongress.podbean.com/) Idan Eretz's… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-vad.audioaudio-classification1M<n<10M9 likes152 downloads10mo agoHugging Face06ivrit-ai /hebrew-handwriting-ocr-benchmarkgated Hebrew Handwriting OCR Benchmark A small, human-verified benchmark for OCR / handwritten text recognition (HTR) on modern Hebrew handwriting: 225 gold lines across 10 pages, one page per writer, drawn from the transcriptor.ivrit.ai volunteer transcription corpus. This is a test set. There is no train split, by design — it exists to be held out. It is deliberately small and clean rather than large and noisy: every line was transcribed by at least two volunteers independently and… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/hebrew-handwriting-ocr-benchmark.imageimage-to-textn<1K0 likes148 downloads21d agoHugging Face07ivrit-ai /audio-basegatedivrit.ai is a database of Hebrew audio and text content. audio-base contains the raw, unprocessed sources. audio-vad contains audio snippets generated by applying Silero VAD (https://github.com/snakers4/silero-vad) to the base dataset. v1 data is generated using silero-vad's default parameters. v2 data is generated using min_speech_duration_ms=2000 (milliseconds), and max_speech_duration_s=30 (seconds). audio-transcripts contains transcriptions for each snippet in the audio-vad dataset. You… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-base.audioaudio-classificationn<1K6 likes129 downloads10mo agoHugging Face08ivrit-ai /eval-whatsappgated Dataset Card for ivrit.ai Whatsapp Eval Evaluation dataset of Hebrew Whatsapp voice-messages, expert-transcribed. Dataset Details Dataset Description This dataset containeד Whatsapp hebrew voice recordings gathered around April 2025. The recordings are of volunteer native hebrew speakers using consumer devices in natural environments. Each recording is by a single speaker about a random topic of their choice spoken in a non-scripted natural manner. Each… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-whatsapp.audioautomatic-speech-recognitionn<1K0 likes120 downloads10mo agoHugging Face09ivrit-ai /knesset-plenumsgated About This dataset is derived from raw a/v recordings and human-generated protocols of the Knesset (the Israeli house of representatives) plenums as part of the ivrit.ai project. Consider visiting the preview space for this dataset here Method Data dumps from the Knesset contain A/V recordings, alongside proprietary protocols with timestamps. We extract the audio stream, and clean up timestamp mistakes (such as backward jumps, or out-of-order timestamp artifacts). The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/knesset-plenums.audioautomatic-speech-recognition1K<n<10K3 likes108 downloads10mo agoHugging Face10ivrit-ai /whisper-traininggated Dataset Card for "whisper-training" Note: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai. More Information needed License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. Full license: https://www.ivrit.ai/en/the-license/ FAQs: https://www.ivrit.ai/en/license-faqs/ audioaudio-classification10K<n<100K13 likes106 downloads10mo agoHugging Face11ivrit-ai /eval-d1gatedNote: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai. License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. Full license: https://www.ivrit.ai/en/the-license/ FAQs: https://www.ivrit.ai/en/license-faqs/ audion<1K6 likes106 downloads10mo agoHugging Face12ivrejchik /brain-teasers_ENtext1K<n<10K0 likes106 downloads2y agoHugging Face13ivrit-ai /crowd-recital-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital Dataset Details Dataset Description License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. - Full license: https://www.ivrit.ai/en/the-license/ - FAQs: https://www.ivrit.ai/en/license-faqs/ Dataset Structure Data Fields Each example in the dataset contains: audio: An audio column containing: bytes: The audio data… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-whisper-training.audiotext-to-speech1K<n<10K3 likes87 downloads10mo agoHugging Face14Aya168 /tofu-translit-ivrit_lat-lebnani-hepburngated TOFU — Hebrish / Arabizi / Romaji (Latin-script transliterations) Three Latin-script transliteration arms of the TOFU fictitious-author unlearning benchmark, built for "Script, Not Syntax: Transliteration as a Blind Spot in Multilingual Unlearning" (Tsir Cohen, Rubinstein, Spira — Trustworthy Machine Learning, Tel Aviv University, 2026). Why this exists TOFU (Maini et al., 2024) asks factual questions about invented authors, so it's answerable only from what a… See the full description on the dataset page: https://huggingface.co/datasets/Aya168/tofu-translit-ivrit_lat-lebnani-hepburn.text10K<n<100K0 likes73 downloads19d agoHugging Face15ivrit-ai /crowd-recital-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Recital - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~78h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi-whisper-training.audiotext-to-speech10K<n<100K0 likes72 downloads10mo agoHugging Face16ivrejchik /medmcqa-instructiontext100K<n<1M0 likes55 downloads2y agoHugging Face17ivrit-ai /audio-transcriptsgatedivrit.ai is a database of Hebrew audio and text content. audio-base contains the raw, unprocessed sources. audio-vad contains audio snippets generated by applying Silero VAD (https://github.com/snakers4/silero-vad) to the base dataset. audio-transcripts contains transcriptions for each snippet in the audio-vad dataset. The audio-base dataset contains data from the following sources: Geekonomy (Podcast, https://geekonomy.net) HaCongress (Podcast, https://hacongress.podbean.com/) Idan Eretz's… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/audio-transcripts.textaudio-classification1M<n<10M13 likes50 downloads10mo agoHugging Face18ivrit-ai /crowd-recital-yigated About This dataset was created by crowd-sourced recording sessions in Yiddish as part of the ivrit.ai Crowd Recital project. Volunteers read on normal desktop or mobile setting Wikipedia articles while time-stamping every sentence read. Later this data is normalized by aligning the gathered captions with the audio using Stable Whisper (See Below). The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is being generated.… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-recital-yi.audioautomatic-speech-recognition1K<n<10K0 likes46 downloads10mo agoHugging Face19ivrejchik /medmcqa-conversationtext100K<n<1M0 likes39 downloads2y agoHugging Face20ivrit-ai /audio-labeledgated Dataset Card for "audio-labeled" More Information needed audio100K<n<1M2 likes31 downloads2y agoHugging Face21ivrit-ai /crowd-whatsapp-yi-whisper-traininggated Dataset Card for ivrit.ai - Crowd Whatsapp - Yiddish See more details on the source dataset card. Dataset Details Dataset Description This is a derived dataset for structured for whisper training: Excludes low quality segments (judged by probabilities of the text-audio auto alignment process) Encodes timestamps along segments of text + previous text Audio encoded to 16K sample-rate, mono Total audio duration - ~19h License: other Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi-whisper-training.audiotext-to-speech1K<n<10K0 likes31 downloads10mo agoHugging Face22ivrejchik /medmcqa-benchmarktext100K<n<1M0 likes27 downloads2y agoHugging Face23anote-ai /IVR-pilot-benchmark IntentSpec Benchmark — Data Supplement This archive contains the benchmark data used to compute Intent Violation Rate (IVR) in the paper: 49 tasks, each derived from a HumanEval problem and extended with an ambiguous/gold prompt pair and a decomposed set of executable constraints. Files spec_pairs.jsonl The benchmark itself — one JSON object per line, one line per task. This is the file consumed directly by the evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/IVR-pilot-benchmark.texttext-generationn<1K0 likes26 downloads1mo agoHugging Face24ivrit-ai /crowd-whatsapp-yigated About This dataset was created by crowd-sourced Whatsapp voice recordings in Yiddish as part of the ivrit.ai project. Volunteers read a message sent to them from a predefined set of messages, recording themselves using Whasapp voice message sent to the collecting bot. Later this data is normalized by aligning the captions with the audio using Stable Whisper (See Below). The recording project is an ongoing effort and new data will be appended to this dataset periodically as it is… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-whatsapp-yi.audioautomatic-speech-recognition1K<n<10K0 likes24 downloads10mo agoHugging Face25ivrit-ai /crowd-transcribe-v4gatedNote: If you are looking for our latest dataset and model, please refer to the main README here: https://huggingface.co/ivrit-ai. crowd-transcribe-v4 This is ivrit.ai's 4th crowd-sourced transcribed dataset release. It contains over 250 hours of volunteer-transcribed data, randomly selected from our audio-vad dataset of over 10,000 hours. License The dataset is released under the ivrit.ai License, which enables broad research and commercial use. Full license:… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/crowd-transcribe-v4.audioautomatic-speech-recognition100K<n<1M2 likes9 downloads10mo agoHugging Face26ivrit-ai /eval-forced-alignmentgated Hebrew Forced Alignment Evaluation Dataset Human-verified, word-level time-aligned Hebrew speech clips. To create this dataset, a dedicated labeling system (similar to Praat, but web-based) was built. The system lets labelers fix the transcript and align each spoken word to the audio, down to 1ms precision (though annotators typically work at ~10ms granularity). The audio samples were gathered by randomly sampling from several of ivrit-ai's larger, published open datasets. The… See the full description on the dataset page: https://huggingface.co/datasets/ivrit-ai/eval-forced-alignment.audioautomatic-speech-recognitionn<1K1 likes9 downloads1d agoHugging Face27Tyl3rDrden /ivrit-manifeststabular10M<n<100M0 likes3 downloads4mo agoHugging Face28Tyl3rDrden /ivrit-manifests-textstabular10M<n<100M0 likes3 downloads4mo agoHugging Face29Tyl3rDrden /ivrit-recordingstabular10K<n<100K0 likes3 downloads4mo agoHugging Face30ivrit-ai /jpress-demogatedimage1K<n<10K0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.