CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ARTPARK-IISc /VaanigatedVAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.audioautomatic-speech-recognition1M<n<10M157 likes19k downloads7d agoHugging Face02ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads1h agoHugging Face03Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face04Emova-ollm /emova-sft-4m EMOVA-SFT-4M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.imageimage-to-text1M<n<10M6 likes3.1k downloads2y agoHugging Face05Ardea /NEXUS-temporal_hierarchical_multi-modal NEXUS: Neural Evolution for eXtensible Universal Semantics Dataset (Temporal Multimodal Slices) This dataset is a multi-modal, hierarchical, temporal representation derived from HuggingFaceFV/finevideo. It is designed for streaming training where the primary unit is a 10 ms "slice" that aggregates upward into moments (100 ms), seconds (1 s), experiences (10 s), and minutes (60 s). It is meant to represent an extensible stream of "experience" as there are… See the full description on the dataset page: https://huggingface.co/datasets/Ardea/NEXUS-temporal_hierarchical_multi-modal.imageautomatic-speech-recognition10M<n<100M5 likes915 downloads4mo agoHugging Face06OmniEvalKit /omnievalkit-dataset OmniEvalKit Evaluation Datasets Evaluation datasets for OmniEvalKit, a comprehensive evaluation framework for omni-modal (audio + video + image + text) models. Overview Total subsets: 65 Total samples: 315,264 Total size: 620.3 GB (Parquet with embedded audio/image/video) Subsets with embedded video: 15 Subsets requiring external video download: 2 Usage from datasets import load_dataset ds = load_dataset("OmniEvalKit/omnievalkit-dataset", "aishell1_test")… See the full description on the dataset page: https://huggingface.co/datasets/OmniEvalKit/omnievalkit-dataset.audioaudio-classification100K<n<1M0 likes437 downloads6mo agoHugging Face07Emova-ollm /emova-sft-speech-231k EMOVA-SFT-Speech-231K 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-Speech-231K is a comprehensive dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-231K is part of EMOVA-Datasets collection and is used in… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-231k.imageaudio-to-audio100K<n<1M3 likes316 downloads2y agoHugging Face08ARTPARK-IISc /Vaani-Benchmark-V1.0gated Vaani-Benchmark-V1.0 A curated ASR evaluation set drawn from the Vaani project. This benchmark contains 5,050 audio segments from 1,103 speakers across 104 Indian districts, each with three independent human transcriptions. Evaluation Toolkit A standalone toolkit implementing this benchmark's scoring methodology, plus Latin-script normalization for code-switched predictions and one-command publishing of results to a model's HF card, is available at… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0.audioautomatic-speech-recognition1K<n<10K7 likes176 downloads1mo agoHugging Face09MingweiFu /ScreenASR-Bench ScreenASR-Bench Data Item Value Split test Cases 2,002 Audio clips 2,002 Keyframes 2,469 Languages Chinese Structure Field Type Description caseid string Unique case identifier ref string Reference transcription target string Target text in TN form level string Difficulty level: L1, L2, or L3 audio audio Audio clip keyframes list[image] Keyframes associated with the case frame_captions list[string]… See the full description on the dataset page: https://huggingface.co/datasets/MingweiFu/ScreenASR-Bench.audioautomatic-speech-recognition1K<n<10K2 likes143 downloads1d agoHugging Face10collectivat /una-fraza-al-diya Una fraza al diya Ladino language learning sentences prepared by Karen Sarhon of Sephardic Center of Istanbul. Each sentence has translations in Turkish, English, Spanish. Includes audio and image. 307 sentences in total. Source: https://sefarad.com.tr/judeo-espanyolladino/frazadeldia/ Citation If you use this dataset, please cite: Preparing an Endangered Language for the Digital Age: The Case of Judeo-Spanish Preparing an endangered language for the digital age: The… See the full description on the dataset page: https://huggingface.co/datasets/collectivat/una-fraza-al-diya.audiotext-generationn<1K1 likes58 downloads11mo agoHugging Face11Emova-ollm /emova-sft-speech-eval EMOVA-SFT-Speech-Eval 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-Speech-Eval is an evaluation dataset curated for omni-modal instruction tuning and emotional spoken dialogue. This dataset is created by converting existing text and visual instruction datasets via Text-to-Speech (TTS) tools. EMOVA-SFT-Speech-Eval is part of EMOVA-Datasets collection, and the training… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-speech-eval.imageaudio-to-audio1K<n<10K1 likes38 downloads2y agoHugging Face12ARTPARK-IISc /Vaani-Atypical-Speech-CorpusgatedProject Euphonia is a public initiative led by Google that aims to improve Automatic Speech Recognition (ASR) for individuals with atypical speech. To date, most of Project Euphonia’s work has focused on English, resulting in outcomes such as the Android application Project Relate, which generates personalized speech recognition models in English. In recent years, the project has expanded its data collection efforts to additional languages, including French, Spanish, Japanese, and Hindi. The… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Atypical-Speech-Corpus.audioautomatic-speech-recognition1K<n<10K2 likes16 downloads6mo agoHugging Face13tonibirat /Sagarmatha-ASR-Nepali-Diamond-V3gated Dataset Card for Sagarmatha ASR Nepali Diamond V3 Dataset Summary Sagarmatha ASR Nepali Diamond V3 is a large-scale, production-grade Automatic Speech Recognition (ASR) dataset designed for the Nepali language. The corpus contains 265.7 hours of verified, 16 kHz audio paired with strictly normalized Devanagari transcriptions. It was compiled and curated primarily for the fine-tuning of state-of-the-art multilingual acoustic models, including OpenAI's Whisper… See the full description on the dataset page: https://huggingface.co/datasets/tonibirat/Sagarmatha-ASR-Nepali-Diamond-V3.audioautomatic-speech-recognition100K<n<1M0 likes6 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.