CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01physical-intelligence /liberoThis dataset was created using LeRobot. Dataset Description This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object, Libero-Goal and Libero-10. All datasets were taken from here and converted into LeRobot format. Homepage: https://libero-project.github.io Paper: https://arxiv.org/abs/2306.03310 License: CC-BY 4.0 Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "panda", "total_episodes": 1693… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/libero.imagerobotics100K<n<1M91 likes42k downloads2y agoHugging Face02arcee-ai /distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO. Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0 text10K<n<100K1 likes24k downloads2y agoHugging Face03argilla /distilabel-intel-orca-dpo-pairs distilabel Orca Pairs for DPO The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved. Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.text10K<n<100K182 likes23k downloads1y agoHugging Face04PrimeIntellect /INTELLECT-3-RLtabular10K<n<100K8 likes3.5k downloads4mo agoHugging Face05Intelligent-Internet /II-Medical-Reasoning-SFT II-Medical-Reasoning-SFT II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice. The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.text1M<n<10M57 likes3.1k downloads1y agoHugging Face06vidore /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes1.7k downloads1y agoHugging Face07geometric-intelligence /ogbench OgBench: Benchmarking Graph Neural Networks on Omics Data OgBench is the first benchmark suite for graph-level prediction in the n ≪ p regime characteristic of omics data, where the number of patient samples n is much smaller than the number of nodes (genes or proteins) p per graph. Datasets This repository contains four preprocessed omics graph classification datasets: Dataset Modality n p Task HERITAGE Proteomics 654 4,977 Exercise responder… See the full description on the dataset page: https://huggingface.co/datasets/geometric-intelligence/ogbench.texttabular-classification100K<n<1M0 likes1.6k downloads12d agoHugging Face08PrimeIntellect /INTELLECT-3-SFTtext1M<n<10M8 likes1.5k downloads10mo agoHugging Face09Delta-Intelligence /OmniEgo D1 Headset Egocentric Whole-body Dataset D1 is a headset multi-camera human motion dataset for humanoid intelligence, embodied AI, whole-body motion understanding, and imitation learning. Overview The D1 dataset is exported from the D1 headset multi-camera human motion capture system developed by Delta Intelligence. Each recorded episode contains synchronized multi-view video streams and whole-body skeleton and headset pose data. The dataset supports research… See the full description on the dataset page: https://huggingface.co/datasets/Delta-Intelligence/OmniEgo.imagerobotics100K<n<1M0 likes1.3k downloads9m agoHugging Face10mteb /syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test. imagedocument-question-answering1K<n<10K0 likes992 downloads8mo agoHugging Face11Intel /SK-VQA Dataset Card for SQ-VQA Dataset Summary SK-VQA is a large-scale synthetic multimodal dataset containing over 2 million visual question-answer pairs, each paired with context documents that contain the information needed to answer the questions. The dataset is designed to address the critical need for training and evaluating multimodal LLMs (MLLMs) in context-augmented generation settings, particularly for retrieval-augmented generation (RAG) systems. It enables training… See the full description on the dataset page: https://huggingface.co/datasets/Intel/SK-VQA.image1M<n<10M1 likes948 downloads1y agoHugging Face12Intelligent-Internet /ii-agent_gaia-benchmark_validationtextn<1K8 likes931 downloads1y agoHugging Face13Intelligent-Internet /GAIA-Subset-Benchmark GAIA Benchmark Subset Model Card This dataset is a subset of the GAIA benchmark, containing 44 web-search-based questions from the validation set. It evaluates multiple AI models on their ability to retrieve and process real-time information using web search and browser tools. Performance metrics include success indicators and detailed reports for each model. A comparative chart summarizing the results will be provided separately. Benchmark Results textn<1K3 likes919 downloads1y agoHugging Face14Intel /SocialCounterfactualsimage100K<n<1M11 likes877 downloads2y agoHugging Face15physical-intelligence /aloha_pen_uncap_diverseThis dataset was created using LeRobot. Dataset Description This dataset is a lerobot conversion of the aloha_pen_uncap_diverse subset of BiPlay. BiPlay contains 9.7 hours of bimanual data collected with an aloha robot at the RAIL lab @ UC Berkeley, USA. It contains 7023 clips, 2000 language annotations and 326 unique scenes. Paper: https://huggingface.co/papers/2410.10088 Code: https://github.com/sudeepdasari/dit-policy If you use the dataset please cite:… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/aloha_pen_uncap_diverse.imagerobotics10K<n<100K8 likes806 downloads2y agoHugging Face16AIM-Intelligence /COMPASS-Policy-Alignment-Testbed-Dataset COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings. What is COMPASS? COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.texttext-generation1K<n<10K12 likes738 downloads27d agoHugging Face17AIM-Intelligence /XL-SafetyBench XL-SafetyBench A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity ⚠️ Content Warning: This dataset contains adversarial prompts and culturally sensitive content for safety and cultural-evaluation research. By using this dataset, you agree to use it solely for research purposes and not for malicious applications. Paper: https://arxiv.org/abs/2605.05662 Eval Code: github.com/AIM-Intelligence/XL-SafetyBench Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.texttext-classification1K<n<10K8 likes565 downloads2mo agoHugging Face18sskapci /domain-intelligence-dataset Domain Intelligence Dataset A large-scale, derived snapshot of the public internet's domain graph: who links to whom, where domains resolve, which nameservers host them, how their DNS records change over time, and computed authority/spam signals on top. Built from three public sources: ICANN CZDS zone files — daily TLD zone snapshots (.com, .net, .org, …) giving the authoritative set of registered domains and their nameserver delegations. CommonCrawl WARC archives — parsed… See the full description on the dataset page: https://huggingface.co/datasets/sskapci/domain-intelligence-dataset.tabulargraph-ml1B<n<10B1 likes551 downloads10d agoHugging Face19K-intelligence /KSAFE-MMgated KSAFE-MM 📑 Paper | 🛠️ Technical Blog 📢 News ⚡️ 2026/06/11: Released on Hugging Face 🤗 📑 2026/05/29: arXiv preprint released 📕 2026/05/20: Technical blog article published ⚠️ CONTENT WARNING This dataset contains potentially harmful and sensitive visual and textual content across the following 11 safety risk categories: Risk Domain Categories Content Safety Risks Hate and Unfairness, Violence, Sexual, Self-harm Socio-economic Risks Political and… See the full description on the dataset page: https://huggingface.co/datasets/K-intelligence/KSAFE-MM.image10K<n<100K27 likes537 downloads4mo agoHugging Face20IntelligenceLab /VideoHallu VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos Zongxia Li*, Xiyang Wu*, Guangyao Shi, Yubin Qin, Hongyang Du, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber [📖 Paper] [🤗 Dataset] [🌍Website] 👀 About VideoHallu Synthetic video generation has gained significant attention for its realism and broad applications, but remains prone to violations of common sense and physical laws. This highlights the need for reliable abnormality… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/VideoHallu.textvisual-question-answering1K<n<10K5 likes517 downloads1y agoHugging Face21PrimeIntellect /INTELLECT-2-RL-Dataset INTELLECT-2 INTELLECT-2 is a 32 billion parameter language model trained through a reinforcement learning run leveraging globally distributed, permissionless GPU resources contributed by the community. The model was trained using prime-rl, a framework designed for distributed asynchronous RL, using GRPO over verifiable rewards along with modifications for improved training stability. For detailed information on our infrastructure and training recipe, see our technical report.… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/INTELLECT-2-RL-Dataset.text100K<n<1M68 likes508 downloads1y agoHugging Face22vidore /syntheticDocQA_artificial_intelligence_test Dataset Description This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications. It includes documents about the Artificial Intelligence. Data Collection Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('artificial intelligence'). From these documents, we randomly sampled 1000 pages. We associated these with 100 questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_artificial_intelligence_test.imagedocument-question-answering1K<n<10K2 likes484 downloads1y agoHugging Face23IntelliProcure /SwissSPARK ⚠️ Caution: The dataset is subject to continuous changes. We are currently actively developing it. Dataset Card: Dataset for a Swiss Sustainable Procurement Analysis & Reporting Kit Dataset Description This dataset is designed to train and evaluate models for detecting sustainability criteria in Swiss public procurement documents (Call for Tenders, CFT). The dataset classifies text segments based on whether they contain specific sustainability… See the full description on the dataset page: https://huggingface.co/datasets/IntelliProcure/SwissSPARK.texttext-classification10K<n<100K0 likes442 downloads17h agoHugging Face24sphita /Intel-WebCorpus-forms 💻 Intel WebCorpus Forms (Enterprise Hardware Q&A) This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums. It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.textquestion-answering100K<n<1M3 likes431 downloads20h agoHugging Face25PrimeIntellect /INTELLECT-2-only-mathtext100K<n<1M0 likes423 downloads1y agoHugging Face26intelsense /see-world-1-CGDimage100K<n<1M0 likes381 downloads2y agoHugging Face27TCLResearchEurope /intelligent_wakeup Intelligent Wakeup A synthetic corpus for device-directed speech detection: multi-speaker conversations in which most speech is not addressed to the voice assistant, with the moments that are clearly marked. Conventional assistants detect a wake word but cannot tell whether what follows is meant for them. This corpus is built to train and evaluate the module that makes that decision from the whole session, not from an isolated command. Project page:… See the full description on the dataset page: https://huggingface.co/datasets/TCLResearchEurope/intelligent_wakeup.audioaudio-classification1K<n<10K0 likes351 downloads11d agoHugging Face28Intelligent-Internet /II-Thought-RL-v0 II-Thought RL v0: A Large-Scale Curated Dataset for Reinforcement Learning See our blog here for additional details. We introduce II-Thought RL v0, the first large-scale, multi-task dataset designed for Reinforcement Learning. This dataset consists of high-quality question-answer pairs that have undergone a rigorous multi-step filtering process, leveraging Gemini 2.0 Flash and Qwen 32B as quality evaluators. In this initial release, we have curated and refined publicly available… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Thought-RL-v0.text100K<n<1M54 likes349 downloads1y agoHugging Face29intelsense /ocr-bn-datagen-v1image1M<n<10M1 likes331 downloads2y agoHugging Face30io-intelligence /realman_aidal_desktop_cleanupThe dataset was collected and open-sourced by IO Intelligence, and exported in the LeRobot format provided by the IO Data Platform. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "custom_arm", "total_episodes": 1099, "total_frames": 246816, "total_tasks": 323, "total_videos": 4396, "total_chunks": 2, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1099"}, "data_path":… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/realman_aidal_desktop_cleanup.tabularrobotics100K<n<1M0 likes326 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.