CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gavinlaw /chinese-lips-speech-slide-probe Chinese-LiPS Speech + Slide Probe A self-contained probe set for testing whether visual slide context helps simultaneous speech translation — with the input as audio, not transcripts. Why audio matters: feeding a transcript to a text LLM deletes the acoustic ambiguity (homophones, polysemy) that slide context is meant to resolve; the transcript already commits to one reading. Any honest test of "does vision help streaming ST" must consume speech. Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.audiotranslationn<1K0 likes249 downloads2mo agoHugging Face02Sellopale /final-certificatesimagetext-classificationn<1K0 likes57 downloads9mo agoHugging Face03collectivat /una-fraza-al-diya Una fraza al diya Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total. Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/ Citation If you use this dataset, please cite: Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.audiotext-generationn<1K1 likes56 downloads11mo agoHugging Face04CortexSwarm /EdgeMMEval EdgeMMEval Minimal multimodal evaluation dataset for on-device inference testing. Covers functional correctness, accuracy, latency stress, and memory pressure across image, audio, text, multi-turn, combination, structured output, and tool-calling cases. Dataset summary The test split is defined in data/test/metadata.jsonl (200 rows). Each row has a test_id (for example IMG-001, STO-020) and a modality. Modality Samples Focus Image 34 VQA, OCR, description… See the full description on the dataset page: https://huggingface.co/datasets/CortexSwarm/EdgeMMEval.audiovisual-question-answeringn<1K0 likes17 downloads5mo agoHugging Face05ARTPARK-IISc /Vaani-Atypical-Speech-CorpusgatedProject Euphonia is a public initiative led by Google that aims to improve Automatic Speech Recognition (ASR) for individuals with atypical speech. To date, most of Project Euphonia’s work has focused on English, resulting in outcomes such as the Android application Project Relate, which generates personalized speech recognition models in English. In recent years, the project has expanded its data collection efforts to additional languages, including French, Spanish, Japanese, and Hindi. The… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Atypical-Speech-Corpus.audioautomatic-speech-recognition1K<n<10K2 likes16 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.