datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
liberoThis dataset was created using LeRobot.
Dataset Description
This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object, Libero-Goal and Libero-10.
All datasets were taken from here and converted into LeRobot format.
Homepage: https://libero-project.github.io
Paper: https://arxiv.org/abs/2306.03310
License: CC-BY 4.0
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 1693… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/libero.assetsLong-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.LHTB-leaderboard
LHTB Leaderboard — Long-Horizon Terminal-Bench
This repository hosts submitted runs for
Long-Horizon Terminal-Bench (LHTB),
a 46-task benchmark measuring how well LLM agents sustain useful work in a
containerized terminal over hundreds of steps.
Every entry below ships its complete run artifacts — per-trial configs, results,
verifier outputs and terminal recordings — so any score on this board can be audited
without rerunning the suite.
📊 Benchmark dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard.snapshotsagi-structural-intelligence-protocols
AGI Structural Intelligence Protocols
Current positioning: SI-Core specifications, evaluation materials, implementation scaffolds, and historical LLM protocol experiments
Status note
The repository name reflects the project's early history. It is not a claim that AGI, machine consciousness, persistent selfhood, or permanent model transformation has been achieved.
This repository now contains two distinct generations of work:
Historical prompt-level experiments that explored… See the full description on the dataset page: https://huggingface.co/datasets/kanaria007/agi-structural-intelligence-protocols.nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
ogbench
OgBench: Benchmarking Graph Neural Networks on Omics Data
OgBench is the first benchmark suite for graph-level prediction in the
n ≪ p regime characteristic of omics data, where the number of
patient samples n is much smaller than the number of nodes (genes or
proteins) p per graph.
Datasets
This repository contains four preprocessed omics graph classification
datasets:
Dataset
Modality
n
p
Task
HERITAGE
Proteomics
654
4,977
Exercise responder… See the full description on the dataset page: https://huggingface.co/datasets/geometric-intelligence/ogbench.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
clawskills-intelligence-corpus
Clawskills: The Complete OpenClaw Skill Collection
The definitive archive of 5,200+ community-built skills for autonomous AI agents.
Discover and deploy the most comprehensive database of OpenClaw skills. High-fidelity local manifests for deep indexing and research.
🔍 Overview
The Clawskills Collection is a high-fidelity local archive of the entire community skill registry. This repository hosts all 5,155 skill manifests locally within the /skills directory, providing a… See the full description on the dataset page: https://huggingface.co/datasets/amoghacloud/clawskills-intelligence-corpus.OmniEgo
D1 Headset Egocentric Whole-body Dataset
D1 is a headset multi-camera human motion dataset for humanoid intelligence, embodied AI, whole-body motion understanding, and imitation learning.
Overview
The D1 dataset is exported from the D1 headset multi-camera human motion capture system developed by Delta Intelligence. Each recorded episode contains synchronized multi-view video streams and whole-body skeleton and headset pose data.
The dataset supports research… See the full description on the dataset page: https://huggingface.co/datasets/Delta-Intelligence/OmniEgo.Hausa
Hausa Ajami OCR Dataset
Ce dataset contient des paires image/transcription de manuscrits haoussa en écriture ajami (écriture arabe adaptée au haoussa).
Contenu
Chaque ligne du fichier data/train/metadata.jsonl correspond à une ligne de texte ajami segmentée, avec :
file_name : nom du fichier image correspondant (image de la ligne, recadrée)
transcript : translittération en écriture latine de la ligne
source : identifiant du manuscrit d'origine (voir tableau… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceResearchLab/Hausa.syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
aloha_pen_uncap_diverseThis dataset was created using LeRobot.
Dataset Description
This dataset is a lerobot conversion of the aloha_pen_uncap_diverse subset of BiPlay.
BiPlay contains 9.7 hours of bimanual data collected with an aloha robot at the RAIL lab @ UC Berkeley, USA. It contains 7023 clips, 2000 language annotations and 326 unique scenes.
Paper: https://huggingface.co/papers/2410.10088 Code: https://github.com/sudeepdasari/dit-policy If you use the dataset please cite:… See the full description on the dataset page: https://huggingface.co/datasets/physical-intelligence/aloha_pen_uncap_diverse.COMPASS-Policy-Alignment-Testbed-Dataset
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
This dataset evaluates how well Large Language Models (LLMs) follow organization-specific policies in realistic enterprise-style settings.
What is COMPASS?
COMPASS is a framework for evaluating policy alignment: given only an organization’s policy (e.g., allow/deny rules), it enables you to benchmark whether an LLM’s responses comply with that policy in structured, enterprise-like… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/COMPASS-Policy-Alignment-Testbed-Dataset.GDPval-CN-Seed-Set
GDPval-CN Seed Set
中文详细说明 · English documentation · 样本说明
GDPval-CN 种子集包含 11 个中文任务,取材自日常知识工作场景。每个任务包括一份任务说明和一组办公材料,例如表格、PDF、文档和结构化数据文件。
我们同时公开了与任务配套的专家工作流,用于设计评分标准和辅助人工复核。
这 11 个任务来自 11 个选定的专业领域,适合用于了解任务形式、测试文件处理能力和搭建评测流程。
GDPval-CN Seed Set contains 11 Chinese-language tasks drawn from everyday knowledge work. Each task includes a task brief, a set of office files, and a separately published expert workflow for rubric design and review.
数据概览
项目
内容
任务数… See the full description on the dataset page: https://huggingface.co/datasets/human-intelligence-ai/GDPval-CN-Seed-Set.docker_to_podmanXL-SafetyBench
XL-SafetyBench
A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
⚠️ Content Warning: This dataset contains adversarial prompts and
culturally sensitive content for safety and cultural-evaluation research.
By using this dataset, you agree to use it solely for research purposes
and not for malicious applications.
Paper: https://arxiv.org/abs/2605.05662
Eval Code: github.com/AIM-Intelligence/XL-SafetyBench
Overview… See the full description on the dataset page: https://huggingface.co/datasets/AIM-Intelligence/XL-SafetyBench.KSAFE-MM
KSAFE-MM
📑 Paper |
🛠️ Technical Blog
📢 News
⚡️ 2026/06/11: Released on Hugging Face 🤗
📑 2026/05/29: arXiv preprint released
📕 2026/05/20: Technical blog article published
⚠️ CONTENT WARNING
This dataset contains potentially harmful and sensitive visual and textual content across the following 11 safety risk categories:
Risk Domain
Categories
Content Safety Risks
Hate and Unfairness, Violence, Sexual, Self-harm
Socio-economic Risks
Political and… See the full description on the dataset page: https://huggingface.co/datasets/K-intelligence/KSAFE-MM.VideoHallu
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos
Zongxia Li*, Xiyang Wu*, Guangyao Shi, Yubin Qin, Hongyang Du, Tianyi Zhou, Dinesh Manocha, Jordan Lee Boyd-Graber
[📖 Paper] [🤗 Dataset] [🌍Website]
👀 About VideoHallu
Synthetic video generation has gained significant attention for its realism and broad applications, but remains prone to violations of common sense and physical laws. This highlights the need for reliable abnormality… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/VideoHallu.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.domain-intelligence-dataset
Domain Intelligence Dataset
A large-scale, derived snapshot of the public internet's domain graph: who links to whom, where domains resolve, which nameservers host them, how their DNS records change over time, and computed authority/spam signals on top.
Built from three public sources:
ICANN CZDS zone files — daily TLD zone snapshots (.com, .net, .org, …) giving the authoritative set of registered domains and their nameserver delegations.
CommonCrawl WARC archives — parsed… See the full description on the dataset page: https://huggingface.co/datasets/sskapci/domain-intelligence-dataset.ridgeline-terrain
ridgeline terrain
Baked global elevation heightfields for ridgeline,
a 3D explorer of the solar system's solid worlds (Rust→WASM + WebGPU) that draws each body as a
globe of stacked latitude ridgelines.
Try the demo → · it streams these files
straight from this dataset.
Derived from public-domain government datasets (U.S. Government works — no copyright); this repo
redistributes resampled, reformatted versions.
Files
File
Body
Grid
Source… See the full description on the dataset page: https://huggingface.co/datasets/idle-intelligence/ridgeline-terrain.syntheticDocQA_artificial_intelligence_test
Dataset Description
This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications.
It includes documents about the Artificial Intelligence.
Data Collection
Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('artificial intelligence'). From these documents, we randomly sampled 1000 pages.
We associated these with 100 questions and answers… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_artificial_intelligence_test.solar-MID-descriptorsFineWeb2-MSA
FineWeb2 MSA Arabic
This is the MSA Arabic Portion of The FineWeb2 Dataset.
This dataset contains a rich collection of text in MSA Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family.
With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic.
Purpose of This Repository
This repository provides easy access to the Arabic portion - MSA… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-MSA.german-news-intelligencechinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.SenseXperience_WristCam_SampleData
SenseXperience Raw MCAP Sample Data
Raw capture episodes from SenseXperience (IO-AI): 12 human motion episodes in ROS 2 MCAP, with 4× compressed video + head IMU. Format details: data format reference.
Item
Value
Episodes
12
Date
2026-07-13
Format
ROS 2 MCAP
Duration
~63–100 s / episode
Modalities
4 cameras (MJPEG) + head IMU (~120 Hz)
Layout
episode_<id>_yyyy_mm_dd_hh_mm_ss/
├── *_mcap_0.mcap # ROS 2 MCAP bag
├── metadata.yaml… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/SenseXperience_WristCam_SampleData.realman_aidal_desktop_cleanupThe dataset was collected and open-sourced by IO Intelligence, and exported in the LeRobot format provided by the IO Data Platform.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_arm",
"total_episodes": 1099,
"total_frames": 246816,
"total_tasks": 323,
"total_videos": 4396,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1099"},
"data_path":… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/realman_aidal_desktop_cleanup.
