datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liberoThis dataset was created using LeRobot.
Dataset Description
This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object, Libero-Goal and Libero-10.
All datasets were taken from here and converted into LeRobot format.
Homepage: https://libero-project.github.io
Paper: https://arxiv.org/abs/2306.03310
License: CC-BY 4.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1693… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/libero.distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO.
Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0
distilabel-intel-orca-dpo-pairs
distilabel Orca Pairs for DPO
The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved.
Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.INTELLECT-3-RLII-Medical-Reasoning-SFT
II-Medical-Reasoning-SFT
II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice.
The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
ogbench
OgBench: Benchmarking Graph Neural Networks on Omics Data
OgBench is the first benchmark suite for graph-level prediction in the
n ≪ p regime characteristic of omics data, where the number of
patient samples n is much smaller than the number of nodes (genes or
proteins) p per graph.
Datasets
This repository contains four preprocessed omics graph classification
datasets:
Dataset
Modality
n
p
Task
HERITAGE
Proteomics
654
4,977
Exercise responder… See the full description on the dataset page: https://huggingface.co/datasets/geometric-intelligence/ogbench.INTELLECT-3-SFTOmniEgo
D1 Headset Egocentric Whole-body Dataset
D1 is a headset multi-camera human motion dataset for humanoid intelligence, embodied AI, whole-body motion understanding, and imitation learning.
Overview
The D1 dataset is exported from the D1 headset multi-camera human motion capture system developed by Delta Intelligence. Each recorded episode contains synchronized multi-view video streams and whole-body skeleton and headset pose data.
The dataset supports research… See the full description on the dataset page: https://huggingface.co/datasets/Delta-Intelligence/OmniEgo.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
SK-VQA
Dataset Card for SQ-VQA
Dataset Summary
SK-VQA is a large-scale synthetic multimodal dataset containing over 2 million visual question-answer pairs, each paired with context documents that contain the information needed to answer the questions.
The dataset is designed to address the critical need for training and evaluating multimodal LLMs (MLLMs) in context-augmented generation settings, particularly for retrieval-augmented generation (RAG) systems. It enables training… See the full description on the dataset page: https://huggingface.co/datasets/Intel/SK-VQA.ii-agent_gaia-benchmark_validationGAIA-Subset-Benchmark
GAIA Benchmark Subset Model Card
This dataset is a subset of the GAIA benchmark, containing 44 web-search-based questions from the validation set. It evaluates multiple AI models on their ability to retrieve and process real-time information using web search and browser tools. Performance metrics include success indicators and detailed reports for each model. A comparative chart summarizing the results will be provided separately.
Benchmark Results
SocialCounterfactualsaloha_pen_uncap_diverseThis dataset was created using LeRobot.
Dataset Description
This dataset is a lerobot conversion of the aloha_pen_uncap_diverse subset of BiPlay.
BiPlay contains 9.7 hours of bimanual data collected with an aloha robot at the RAIL lab @ UC Berkeley, USA. It contains 7023 clips, 2000 language annotations and 326 unique scenes.
Paper: https://huggingface.co/papers/2410.10088 Code: https://github.com/sudeepdasari/dit-policy If you use the dataset please cite:… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/aloha_pen_uncap_diverse.COMPASS-Policy-Alignment-Testbed-Dataset
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings.
What is COMPASS?
COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.XL-SafetyBench
XL-SafetyBench
A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
⚠️ Content Warning: This dataset contains adversarial prompts and
culturally sensitive content for safety and cultural-evaluation research.
By using this dataset, you agree to use it solely for research purposes
and not for malicious applications.
Paper: https://arxiv.org/abs/2605.05662
Eval Code: github.com/AIM-Intelligence/XL-SafetyBench
Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.domain-intelligence-dataset
Domain Intelligence Dataset
A large-scale, derived snapshot of the public internet's domain graph: who links to whom, where domains resolve, which nameservers host them, how their DNS records change over time, and computed authority/spam signals on top.
Built from three public sources:
ICANN CZDS zone files — daily TLD zone snapshots (.com, .net, .org, …) giving the authoritative set of registered domains and their nameserver delegations.
CommonCrawl WARC archives — parsed… See the full description on the dataset page: https://huggingface.co/datasets/sskapci/domain-intelligence-dataset.KSAFE-MM
KSAFE-MM
📑 Paper |
🛠️ Technical Blog
📢 News
⚡️ 2026/06/11: Released on Hugging Face 🤗
📑 2026/05/29: arXiv preprint released
📕 2026/05/20: Technical blog article published
⚠️ CONTENT WARNING
This dataset contains potentially harmful and sensitive visual and textual content across the following 11 safety risk categories:
Risk Domain
Categories
Content Safety Risks
Hate and Unfairness, Violence, Sexual, Self-harm
Socio-economic Risks
Political and… See the full description on the dataset page: https://huggingface.co/datasets/K-intelligence/KSAFE-MM.VideoHallu
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos
Zongxia Li*, Xiyang Wu*, Guangyao Shi, Yubin Qin, Hongyang Du, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber
[📖 Paper] [🤗 Dataset] [🌍Website]
👀 About VideoHallu
Synthetic video generation has gained significant attention for its realism and broad applications, but remains prone to violations of common sense and physical laws. This highlights the need for reliable abnormality… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/VideoHallu.INTELLECT-2-RL-Dataset
INTELLECT-2
INTELLECT-2 is a 32 billion parameter language model trained through a reinforcement learning run leveraging globally distributed, permissionless GPU resources contributed by the community.
The model was trained using prime-rl, a framework designed for distributed asynchronous RL, using GRPO over verifiable rewards along with modifications for improved training stability. For detailed information on our infrastructure and training recipe, see our technical report.… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/INTELLECT-2-RL-Dataset.syntheticDocQA_artificial_intelligence_test
Dataset Description
This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications.
It includes documents about the Artificial Intelligence.
Data Collection
Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('artificial intelligence'). From these documents, we randomly sampled 1000 pages.
We associated these with 100 questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_artificial_intelligence_test.SwissSPARK
⚠️ Caution: The dataset is subject to continuous changes. We are currently actively developing it.
Dataset Card: Dataset for a Swiss Sustainable Procurement Analysis & Reporting Kit
Dataset Description
This dataset is designed to train and evaluate models for detecting sustainability criteria in Swiss public procurement documents (Call for Tenders, CFT). The dataset classifies text segments based on whether they contain specific sustainability… See the full description on the dataset page: https://huggingface.co/datasets/IntelliProcure/SwissSPARK.Intel-WebCorpus-forms
💻 Intel WebCorpus Forms (Enterprise Hardware Q&A)
This dataset is a massive, high-fidelity archive of 176,472 technical troubleshooting discussions (containing nearly 1 million individual messages) scraped from the official Intel Community Forums.
It has been meticulously engineered for Large Language Model (LLM) training. Instead of a raw, messy dump of isolated posts, the data has been reconstructed into chronological conversation threads, noise-filtered, deduplicated, and… See the full description on the dataset page: https://huggingface.co/datasets/sphita/Intel-WebCorpus-forms.INTELLECT-2-only-mathsee-world-1-CGDintelligent_wakeup
Intelligent Wakeup
A synthetic corpus for device-directed speech detection: multi-speaker conversations in
which most speech is not addressed to the voice assistant, with the moments that are
clearly marked.
Conventional assistants detect a wake word but cannot tell whether what follows is meant for
them. This corpus is built to train and evaluate the module that makes that decision from the
whole session, not from an isolated command.
Project page:… See the full description on the dataset page: https://huggingface.co/datasets/TCLResearchEurope/intelligent_wakeup.II-Thought-RL-v0
II-Thought RL v0: A Large-Scale Curated Dataset for Reinforcement Learning
See our blog here for additional details.
We introduce II-Thought RL v0, the first large-scale, multi-task dataset designed for Reinforcement Learning. This dataset consists of high-quality question-answer pairs that have undergone a rigorous multi-step filtering process, leveraging Gemini 2.0 Flash and Qwen 32B as quality evaluators.
In this initial release, we have curated and refined publicly available… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Thought-RL-v0.ocr-bn-datagen-v1realman_aidal_desktop_cleanupThe dataset was collected and open-sourced by IO Intelligence, and exported in the LeRobot format provided by the IO Data Platform.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_arm",
"total_episodes": 1099,
"total_frames": 246816,
"total_tasks": 323,
"total_videos": 4396,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1099"},
"data_path":… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/realman_aidal_desktop_cleanup.
