CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01InternRobotics /OmniWorld[ICLR 2026] OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling         🎉NEWS [2026.3.21] 🔥 OmniWorld-Game with Metric Scale is now released! Check out our latest model Pi3X (an enhanced version of Pi3), which leverages this data to achieve better performance! [2026.1.26] 🎉 OmniWorld was accepted by ICLR 2026! [2026.1.7] Update OmniWorld-Game, release RH20T-Robot, RH20T-Human, Ego-Exo4D, EgoDex, Epic-Kitchens. [2025.11.11] The OmniWorld is… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/OmniWorld.imagetext-to-video1B<n<10B96 likes69k downloads5mo agoHugging Face02TIGER-Lab /OmniEdit-Filtered-1.2M OmniEdit In this paper, we present OMNI-EDIT, which is an omnipotent editor to handle seven different image editing tasks with any aspect ratio seamlessly. Our contribution is in four folds: (1) OMNI-EDIT is trained by utilizing the supervision from seven different specialist models to ensure task coverage. (2) we utilize importance sampling based on the scores provided by large multimodal models (like GPT-4o) instead of CLIP-score to improve the data quality. 📃Paper | 🌐Website |… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/OmniEdit-Filtered-1.2M.image1M<n<10M132 likes62k downloads2y agoHugging Face03OpenMOSS-Team /OmniAction RoboOmni: Proactive Robot Manipulation in Omni-modal Context 📖 arXiv Paper (Accepted to ICLR 2026 🎉) | 🌐 Website | 🤗 Model | 🤗 Dataset | 🛠️ Github | Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/OmniAction.robotics285 likes62k downloads6mo agoHugging Face04tars-robotics /OmniVitac9 likes28k downloads6mo agoHugging Face05opendatalab /OmniDocBench OmniDocBench English | 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics: Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.image1K<n<10K106 likes26k downloads3mo agoHugging Face06Insta360-Research /OmniRooms UniSHARP: Universal Sharp Monocular View Synthesis Meixi Song1 · Dizhe Zhang1,* · Hao Ren1 · Ruiyang Zhang1 · Bo Du2 · Ming-Hsuan Yang3 · Lu Qi1,2,* 1Insta360 Research · 2Wuhan University · 3University of California, Merced UniSHARP extends SHARP-style photorealistic monocular view synthesis to universal camera systems. Given a single image from a perspective, wide-FoV, fisheye, or panoramic camera, UniSHARP predicts a 3D Gaussian representation and… See the full description on the dataset page: https://huggingface.co/datasets/Insta360-Research/OmniRooms.imagedepth-estimation100K<n<1M6 likes19k downloads3mo agoHugging Face07omnibioai /omnibioai-sif-images OmniBioAI SIF Images 🧬 500+ native ARM64 Singularity (SIF) container images for bioinformatics, built on NVIDIA DGX (aarch64). Tool Categories Category Tools Genomics & Alignment BWA, STAR, HISAT2, Minimap2, Bowtie2 Variant Calling GATK, DeepVariant, Clair3, Mutect2 RNA-seq Salmon, Kallisto, DESeq2, edgeR Single Cell Seurat, Scanpy, Cell Ranger, Harmony Epigenomics MACS2, deepTools, Bismark Metagenomics Kraken2, MetaPhlAn, QIIME2 Proteomics… See the full description on the dataset page: https://huggingface.co/datasets/omnibioai/omnibioai-sif-images.0 likes18k downloads2mo agoHugging Face08OpenGVLab /OmniCorpus-CC 🐳 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text ⭐️ NOTE: Several parquet files were marked unsafe (viruses) by official scaning of hf, while they are reported safe by ClamAV and Virustotal. We found many false positive cases of the hf automatic scanning in hf discussions and raise one discussion to ask for a re-scanning. This is the repository of OmniCorpus-CC, which contains 988 million image-text interleaved documents collected from Common… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/OmniCorpus-CC.textimage-to-text100M<n<1B29 likes17k downloads2y agoHugging Face09OpenGVLab /OmniCorpus-CC-210Mgated 🐳 OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text This repository contains 210 million image-text interleaved documents filtered from the OmniCorpus-CC dataset, which was sourced from Common Crawl. Repository: https://github.com/OpenGVLab/OmniCorpus Paper (ICLR 2025 Spotlight): https://arxiv.org/abs/2406.08418 OmniCorpus dataset is a large-scale image-text interleaved dataset, which pushes the boundaries of scale and diversity by encompassing… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/OmniCorpus-CC-210M.textimage-to-text100M<n<1B35 likes15k downloads2y agoHugging Face10lsmpp /omni-refiner-video0 likes13k downloads8mo agoHugging Face11chuonghm /OmniRet-train OmniRet training dataset OmniRet-train is the training-data release for OmniRet, a unified retrieval model for text, image, video, and audio. This card documents the released snapshot for researchers training or analyzing OmniRet. Dataset summary The release contains 6,405,109 query rows and 7,119,841 candidate rows from 30 datasets. It covers 15 retrieval directions across text (T), image (I), video (V), and audio (A). The OmniRet paper reports this corpus as… See the full description on the dataset page: https://huggingface.co/datasets/chuonghm/OmniRet-train.image10M<n<100M0 likes11k downloads1mo agoHugging Face12KbsdJames /Omni-MATH Dataset Card for Omni-MATH Recent advancements in AI, particularly in large language models (LLMs), have led to significant breakthroughs in mathematical reasoning capabilities. However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.g., OpenAI o1 achieves 94.8% on MATH dataset), indicating their inadequacy for truly challenging these models. To mitigate this limitation, we propose a comprehensive and challenging benchmark specifically designed… See the full description on the dataset page: https://huggingface.co/datasets/KbsdJames/Omni-MATH.text1K<n<10K132 likes11k downloads2y agoHugging Face13OmniAICreator /ASMR-Archive-Processed ASMR-Archive-Processed (WIP) Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended. Work in Progress — expect breaking changes while the pipeline and data layout stabilize. This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.imageautomatic-speech-recognition96 likes10k downloads6mo agoHugging Face14OpenMOSS-Team /OmniAction-LIBERO RoboOmni: Proactive Robot Manipulation in Omni-modal Context 📖 arXiv Paper (Accepted to ICLR 2026 🎉) | 🌐 Website | 🤗 Model | 🤗 Dataset | 🛠️ Github | Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/OmniAction-LIBERO.robotics70 likes10k downloads6mo agoHugging Face15alibaba-pai /OmniThoughtV_Raw_1.8M Dataset Introduction OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Raw_1.8M.text1M<n<10M1 likes9.7k downloads8mo agoHugging Face16Rocky131 /OmniReasoner-SFT OmniReasoner-SFT OmniReasoner-SFT is a mixed-source, research-only supervised fine-tuning dataset for audio-visual and long-video reasoning. It contains two-stage cold-start SFT trajectories with interval selection, zoom-in evidence, and final answers. Contents data/train.jsonl: HF-ready training JSONL with repo-relative media paths. media/: raw and derived media referenced by train.jsonl. manifests/media_manifest.jsonl: media inventory with repo paths, source family… See the full description on the dataset page: https://huggingface.co/datasets/Rocky131/OmniReasoner-SFT.audiovisual-question-answering10K<n<100K0 likes8.5k downloads4mo agoHugging Face17PaddlePaddle /Real5-OmniDocBench Real5-OmniDocBench A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild Leaderboard | Overview | Dataset | Evaluation | Submit Results | Citation Real5-OmniDocBench measures the robustness of document parsing systems under five physical acquisition conditions: Scanning, Warping, Screen-Photography, Illumination, and Skew. It reconstructs the same 1,355 pages from OmniDocBench v1.5 in every condition, producing 6,775 images in total. The one-to-one… See the full description on the dataset page: https://huggingface.co/datasets/PaddlePaddle/Real5-OmniDocBench.documentimage-to-text1K<n<10K38 likes8.3k downloads10d agoHugging Face18omnicad-lab-L3d /Omni-CAD-Subset-Completeimage0 likes8k downloads10mo agoHugging Face19ArtificialAnalysis /AA-Omniscience-Public Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.documentquestion-answeringn<1K49 likes7.8k downloads29d agoHugging Face20Omni-iEEG /Omni-iEEGtabular10K<n<100K3 likes6.7k downloads2mo agoHugging Face21paxini /Omnisharing_DB_SampleData Overview The embodied intelligence industry is currently facing significant development challenges. The most critical issue is the lack of high-quality data, particularly omnimodal data that integrates force and tactile sensing. The PaXini introduces the PX OmniSharing Dataset, built on the PaXini Super EID Factory, enabling large-scale, high-fidelity human data collection across diverse tasks and scenarios. The dataset includes multi-dimensional tactile data, multi-view visual… See the full description on the dataset page: https://huggingface.co/datasets/paxini/Omnisharing_DB_SampleData.8 likes6.5k downloads4mo agoHugging Face22worstchan /Belle_1.4M-SLAM-Omni Belle_1.4M This dataset is prepared for the reproduction of SLAM-Omni. This is a multi-round Chinese spoken dialogue training dataset. For code and usage examples, please refer to the related GitHub repository: X-LANCE/SLAM-LLM (examples/s2s) 🔧 Modifications Data Filtering: We removed samples with excessively long data. Speech Response Tokens: We used CosyVoice to synthesize corresponding semantic speech tokens for the speech response. These tokens, represented as… See the full description on the dataset page: https://huggingface.co/datasets/worstchan/Belle_1.4M-SLAM-Omni.tabularquestion-answering1M<n<10M3 likes6.3k downloads1y agoHugging Face23alibaba-pai /OmniThoughtV_Filter_0.5M Dataset Introduction OmniThoughtV is a large-scale multimodal long-chain-of-thought dataset distilled from the FineVision dataset using Alibaba Cloud's AI platform (PAI) distillation toolkit, EasyDistill. This dataset establishes a transparent and reproducible data distillation pipeline, enabling efficient construction of multimodal reasoning chains of thought. Fine-tuning smaller models with this dataset effectively endows them with stronger reasoning capabilities and enhances… See the full description on the dataset page: https://huggingface.co/datasets/alibaba-pai/OmniThoughtV_Filter_0.5M.1 likes5.7k downloads8mo agoHugging Face24wd21 /omnibox-backups10 likes5.5k downloads30m agoHugging Face25RUC-NLPIR /Omnimodal-Agent-SFT-2K OmniGAIA: Omni-Modal General AI Assistant Benchmark 📄 Paper   •   💻 Code & Demo   •   🤗 Dataset & Model   •   📈 Leaderboard This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.audioquestion-answering1K<n<10K9 likes5k downloads7mo agoHugging Face26hosam12kalad /OmniAction RoboOmni: Proactive Robot Manipulation in Omni-modal Context 📖 arXiv Paper (Accepted to ICLR 2026 🎉) | 🌐 Website | 🤗 Model | 🤗 Dataset | 🛠️ Github | Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real-world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/hosam12kalad/OmniAction.robotics0 likes4.7k downloads6mo agoHugging Face27omnibioai /pubmed-faiss-indexes0 likes4.5k downloads2h agoHugging Face28gpt-omni /VoiceAssistant-400Kaudio100K<n<1M100 likes4.2k downloads2y agoHugging Face29behavior-1k /omnigibson-robot-assets omnigibson-robot-assets This repo contains the robot assets and the state mesh assets for OmniGibson. Pushing updates First, make sure you have updated the VERSION file. Every zipped release must have a higher version. This can go ahead of the OmniGibson version. echo "9999.9.9" > VERSION Then commit the change to main before building the archive, since git archive only packages committed files: git status # Verify it contains the stuff you need git add -A && git… See the full description on the dataset page: https://huggingface.co/datasets/behavior-1k/omnigibson-robot-assets.0 likes4.2k downloads5mo agoHugging Face30yaolily /Timechat-OmniCaptioner-42K TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions 🌟 Overview TimeChat-Captioner is a multimodal model designed to generate detailed, time-aware, and structurally coherent captions for multi-scene videos. It effectively coordinates visual and audio information to provide comprehensive video descriptions. 🌐 Project Page: timechat-captioner.github.io 🏠 Model: TimeChat-Captioner (7B) 📚 Train Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/yaolily/Timechat-OmniCaptioner-42K.videovideo-text-to-text10K<n<100K5 likes4.2k downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.