datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BlueLens
Dataset Card for BlueLense
Dataset Summary
Dataset Details
Model
Description
Checkpoint Used
Dataset
Split
GDINO
Features from COCO 2017 training set
groundingdino-swint-ogc
COCO
COCO_TRAIN
GDINO
Features from COCO 2017 validation set
groundingdino-swint-ogc
COCO
COCO_VAL
GDINO
~100 random samples from COCO 2017
groundingdino-swint-ogc
COCO
COCO_MINI
DETR
~100 random samples from COCO 2017
detr_r50_8xb2-150e_coco_20221023_153551-436d03e8… See the full description on the dataset page: https://huggingface.co/datasets/IntelLabs/BlueLens.liberoThis dataset was created using LeRobot.
Dataset Description
This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object, Libero-Goal and Libero-10.
All datasets were taken from here and converted into LeRobot format.
Homepage: https://libero-project.github.io
Paper: https://arxiv.org/abs/2306.03310
License: CC-BY 4.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1693… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/libero.assetsdistilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO.
Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0
distilabel-intel-orca-dpo-pairs
distilabel Orca Pairs for DPO
The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved.
Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.BEDLAM-depth
Dataset Mirror of BEDLAM Dataset (Depth Data Subset)
Project site: https://bedlam.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM
Dataset Information
Depth maps (EXR, 32-bit, 3.8TB)
Camera ground truth information is not included but can be found in the BEDLAM dataset mirror
Image/video data with motion blur is not included but can be found in the BEDLAM dataset… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM-depth.BEDLAM
Dataset Mirror of BEDLAM Dataset
Project site: https://bedlam.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM-depth
Dataset Information
Synthetic video data (15h)
10450 image sequences, 30fps, 1280x720
1.6 million images (PNG, 2.2TB)
movies (MP4/H.264, 20GB)
camera/scene ground truth for all sequences (CSV+JSON, 100MB)
segmentation masks (PNG, 30GB)
Depth data is… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM.Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.MammaSyn
MammaSyn
MAMMA training dataset
Project site: https://mamma.is.tue.mpg.de/
Synthetic training data, each scene rendered from 8 views, in WebDataset format.
Dataset Groups
MammaSyn-Interactions
Includes Harmony4D, Inter-X, InteractionCouple, LatinDance10
Hi4D is currently not included but will be added when we receive permission
MammaSyn-Singles
Includes BEDLAM, MOYO
MammaSyn-Hands
Includes InterHand
SignAvatars is currently not included but will be added… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/MammaSyn.ict_s2s_refactoredWaste-Dumpsites-DroneImagery
Dataset for Waste/Dumpsite Detection using drone imagery
Contains 2115 drone images of illegal waste dumpsites
1280 x 1280 px resolution
Nadir perspective (camera pointing straight down at a 90-degree angle to the ground)
Annotations and Images
train | valid | test
actual images
COCO - annotations_coco.json files in each split directory
.parquet files in data directory with embeded images
The dataset was collected as part of the [ Raven Scan ] project, more… See the full description on the dataset page: https://huggingface.co/datasets/INS-IntelligentNetworkSolutions/Waste-Dumpsites-DroneImagery.IntelliSA-dataset
IntelliSA Dataset
Infrastructure as Code security vulnerability dataset with ground truth labels and pseudo-labeled training data across Chef, Ansible, and Puppet.
Dataset Overview
Component
Size
Purpose
Oracle
241 scripts, 213 smells
Ground truth evaluation set
Training
2,300 instances + 6,070 raw scripts
Model training data
Oracle Dataset (Ground Truth)
Ansible: 81 scripts, 44 smells
Chef: 80 scripts, 104 smells
Puppet: 80 scripts, 65… See the full description on the dataset page: https://huggingface.co/datasets/colemei/IntelliSA-dataset.LHTB-leaderboard
LHTB Leaderboard — Long-Horizon Terminal-Bench
This repository hosts submitted runs for
Long-Horizon Terminal-Bench (LHTB),
a 46-task benchmark measuring how well LLM agents sustain useful work in a
containerized terminal over hundreds of steps.
Every entry below ships its complete run artifacts — per-trial configs, results,
verifier outputs and terminal recordings — so any score on this board can be audited
without rerunning the suite.
📊 Benchmark dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard.BEDLAM2
Dataset Mirror of BEDLAM2.0 Dataset
Project site: https://bedlam2.is.tuebingen.mpg.de/
Please register at project site for additional information and data (Download section)
Related Hugging Face dataset mirror: BEDLAM2-depth
Dataset Information
Synthetic video data (75h)
27480 image sequences, 30fps, 1280x720
8 million images (PNG, 11TB)
movies (MP4/H.264, 160GB)
camera/scene ground truth for all sequences (CSV+JSON, 4GB)
overview images and plots (6GB)
Depth data is… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2.snapshotsBEDLAM2-depth
Dataset Mirror of BEDLAM2.0 Dataset (Depth Data Subset)
Project site: https://bedlam2.is.tuebingen.mpg.de/
Please register at project site for additional information and data in its Download section.
Related Hugging Face dataset mirror: BEDLAM2
Dataset Information
Depth maps (Multilayer EXR, 16-bit, available for 44% of images, 15TB)
Multilayer EXR details
16-bit float depth in red channel (FinalImageMovieRenderQueue_WorldDepth.R)
Color image without motion blur
Body… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Systems/BEDLAM2-depth.INTELLECT-3-RLagi-structural-intelligence-protocols
AGI Structural Intelligence Protocols
Current positioning: SI-Core specifications, evaluation materials, implementation scaffolds, and historical LLM protocol experiments
Status note
The repository name reflects the project's early history. It is not a claim that AGI, machine consciousness, persistent selfhood, or permanent model transformation has been achieved.
This repository now contains two distinct generations of work:
Historical prompt-level experiments that explored… See the full description on the dataset page: https://huggingface.co/datasets/kanaria007/agi-structural-intelligence-protocols.II-Medical-Reasoning-SFT
II-Medical-Reasoning-SFT
II-Medical SFT is a curated dataset designed to support the supervised fine-tuning of large language models (LLMs) for medical reasoning tasks. It comprises multi-turn dialogues, clinical case scenarios, and question-answer pairs that reflect the complex reasoning processes encountered in real-world clinical practice.
The dataset is intended to help models develop key competencies such as differential diagnosis, evidence-based decision-making, patient… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/II-Medical-Reasoning-SFT.sim-datasets
SIM-Datasets: A Unified Symbolic Regression Benchmark
A standardized benchmark collection designed for the Scientific Intelligent Modelling (SIM) toolkit, providing comprehensive datasets for symbolic regression research and applications.
Overview
SIM-Datasets serves as a unified benchmark for symbolic regression tasks, offering standardized datasets with consistent formatting and evaluation protocols. This collection is specifically curated to support the Scientific… See the full description on the dataset page: https://huggingface.co/datasets/scientific-intelligent-modelling/sim-datasets.nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
pd12m
PD12M
This is a curated PD12M dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Public Domain 12M image collection, refined by filtering for active image links. EXIF data was extracted, and images underwent preprocessing and feature extraction using SigLIP 2. All vector embeddings are normalized 16-bit half-precision vectors optimized for L2 indexing with vectorchord.… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/pd12m.orca_dpo_pairsThe dataset contains 12k examples from Orca style dataset Open-Orca/OpenOrca.
ogbench
OgBench: Benchmarking Graph Neural Networks on Omics Data
OgBench is the first benchmark suite for graph-level prediction in the
n ≪ p regime characteristic of omics data, where the number of
patient samples n is much smaller than the number of nodes (genes or
proteins) p per graph.
Datasets
This repository contains four preprocessed omics graph classification
datasets:
Dataset
Modality
n
p
Task
HERITAGE
Proteomics
654
4,977
Exercise responder… See the full description on the dataset page: https://huggingface.co/datasets/geometric-intelligence/ogbench.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
clawskills-intelligence-corpus
Clawskills: The Complete OpenClaw Skill Collection
The definitive archive of 5,200+ community-built skills for autonomous AI agents.
Discover and deploy the most comprehensive database of OpenClaw skills. High-fidelity local manifests for deep indexing and research.
🔍 Overview
The Clawskills Collection is a high-fidelity local archive of the entire community skill registry. This repository hosts all 5,155 skill manifests locally within the /skills directory, providing a… See the full description on the dataset page: https://huggingface.co/datasets/amoghacloud/clawskills-intelligence-corpus.INTELLECT-3-SFTwikipedia_en
wikipedia_en
This is a curated Wikipedia English dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Wikipedia English pages. Data sourced directly from the official English Wikipedia database dump. We extract the pages, chunk them into smaller pieces, and embed them using Snowflake/snowflake-arctic-embed-m-v2.0. All vector embeddings are 16-bit half-precision vectors optimized for cosine indexing… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/wikipedia_en.intel-image-classification
Intel Image Classification
The Intel Image Classification dataset contains images of natural scenes categorized into six classes:
Buildings
Forest
Glacier
Mountain
Sea
Street
📆 Content
The dataset contains ~25,000 images of size 150x150 pixels.
Images are evenly distributed across 6 categories:
{'buildings' -> 0,
'forest' -> 1,
'glacier' -> 2,
'mountain' -> 3,
'sea' -> 4,
'street' -> 5 }
It is divided into three parts:
Training set: ~14… See the full description on the dataset page: https://huggingface.co/datasets/sfarrukhm/intel-image-classification.OmniEgo
D1 Headset Egocentric Whole-body Dataset
D1 is a headset multi-camera human motion dataset for humanoid intelligence, embodied AI, whole-body motion understanding, and imitation learning.
Overview
The D1 dataset is exported from the D1 headset multi-camera human motion capture system developed by Delta Intelligence. Each recorded episode contains synchronized multi-view video streams and whole-body skeleton and headset pose data.
The dataset supports research… See the full description on the dataset page: https://huggingface.co/datasets/Delta-Intelligence/OmniEgo.
