CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jupyterjazz /XModBench-MTEB XModBench-Lite for MTEB This repository is a deterministic MTEB normalization of the official RyanWW/XModBench XModBench-Lite release at revision a679188cf062b9810d2e09c2edabc0b1aef9f244. The source contains 6,000 four-choice questions balanced across six canonical modality configurations and five capability families. This MTEB adaptation retains 5,981 questions. It excludes 19 questions that reference five unusable MP4 files in the pinned official archive. Four are truncated:… See the full description on the dataset page: https://huggingface.co/datasets/jupyterjazz/XModBench-MTEB.audio10K<n<100K0 likes564 downloads25d agoHugging Face02RyanWW /XModBenchXModBench Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models 🎉 Accepted at ICLR 2026 What is XModBench? XModBench is the first tri-modal (audio / vision / text) multiple-choice QA benchmark explicitly designed to measure cross-modal consistency — does an omni-language model give the same correct answer when the same semantic content is presented in different modalities? Each item is a 4-choice question with a <context>… See the full description on the dataset page: https://huggingface.co/datasets/RyanWW/XModBench.audiomultiple-choice10K<n<100K3 likes506 downloads4mo agoHugging Face03xmodar /commonvoice-12.0-arabic-voice-converted Dataset Card for Voice Converted Arabic Common Voice 12.0 This dataset is derived from the Common Voice Arabic Corpus 12.0 and includes automatically diacritized transcriptions and phoneme representations for the original augmented audio data. The recordings feature Arabic text read aloud by users, where the text was initially undiacritized, allowing for potential reading errors. The diacritization and phonemes were generated automatically, resulting in a dataset that is valuable… See the full description on the dataset page: https://huggingface.co/datasets/xmodar/commonvoice-12.0-arabic-voice-converted.audioautomatic-speech-recognition100K<n<1M8 likes355 downloads2y agoHugging Face04xmodar /commonvoice-12.0-arabic-voice-converted-speakersThese are the speakers' embeddings and information for the Voice Converted Arabic Common Voice 12.0 dataset. textn<1K0 likes15 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.