datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Long-Horizon-Terminal-Bench
Long-Horizon Terminal-Bench (LHTB)
LHTB is a 46-task benchmark for measuring how well LLM agents sustain useful
work in a containerized terminal over hundreds of steps. Unlike short-horizon
coding benchmarks where an agent writes one artifact and stops, LHTB drops the agent
into a stateful environment and grades it with hidden, rebuild-from-artifact
verifiers — self-reported progress does not count.
📝 Blog: https://zli12321.github.io/LHTB/
🏆 Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Long-Horizon-Terminal-Bench.LHTB-leaderboard
LHTB Leaderboard — Long-Horizon Terminal-Bench
This repository hosts submitted runs for
Long-Horizon Terminal-Bench (LHTB),
a 46-task benchmark measuring how well LLM agents sustain useful work in a
containerized terminal over hundreds of steps.
Every entry below ships its complete run artifacts — per-trial configs, results,
verifier outputs and terminal recordings — so any score on this board can be audited
without rerunning the suite.
📊 Benchmark dataset:… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/LHTB-leaderboard.nuclear-intelligence-dataset
Nuclear Intelligence Dataset
Public, auto-generated dataset of validated nuclear-energy research cycles.
Latest stats (auto-updated):
🪙 NES tokens minted: 0
⛓️ Blockchain length: 1 blocks
🕸️ Knowledge entities: 2
Source
GitHub: https://github.com/QalamHipHop/nuclear-intelligence
HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence
License
MIT
syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
syntheticDocQA_artificial_intelligence_test_beirBEIR version of vidore/syntheticDocQA_artificial_intelligence_test.
domain-intelligence-dataset
Domain Intelligence Dataset
A large-scale, derived snapshot of the public internet's domain graph: who links to whom, where domains resolve, which nameservers host them, how their DNS records change over time, and computed authority/spam signals on top.
Built from three public sources:
ICANN CZDS zone files — daily TLD zone snapshots (.com, .net, .org, …) giving the authoritative set of registered domains and their nameserver delegations.
CommonCrawl WARC archives — parsed… See the full description on the dataset page: https://huggingface.co/datasets/sskapci/domain-intelligence-dataset.chinese-materials-science-open-intelligence
🔬 Chinese Materials Science & Metallurgy Open Intelligence Dataset
Curated open intelligence dataset providing English research briefs, authoritative DOIs, executive summaries, and high-resolution micrographs of breakthrough Chinese scientific research in Materials Science, Metallurgy, Advanced Alloys, and Mining Engineering.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-materials-science-open-intelligence.realman_aidal_desktop_cleanupThe dataset was collected and open-sourced by IO Intelligence, and exported in the LeRobot format provided by the IO Data Platform.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "custom_arm",
"total_episodes": 1099,
"total_frames": 246816,
"total_tasks": 323,
"total_videos": 4396,
"total_chunks": 2,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1099"},
"data_path":… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/realman_aidal_desktop_cleanup.docqa_artificial_intelligence_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_artificial_intelligence_beir.chinese-clean-energy-battery-open-intelligence
🔬 Chinese Clean Energy, Battery Chemistry & Smart Grid Open Intelligence Dataset
Curated open intelligence dataset tracking authentic Chinese scientific breakthroughs in Solid-State Battery chemistry, Perovskite Solar cells, Ultra-High Voltage (UHV) power grids, and industrial decarbonization.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-clean-energy-battery-open-intelligence.objaverse.data.intelligence
Objaverse Data Intelligence: Scene-Level Structural Analysis and Decomposition Metadata
Large-scale scene-level geometry, structure, and material analysis of ~680K Objaverse assets, designed for ML-ready filtering, structural decomposition, and dataset curation.
🎥 Watch Reel
Quick Navigation
Overview
Key Features
Dataset Structure
Additional Files
Dataset Statistics
Classification & Detection Quality
Attribution
Citation
License
Acknowledgements… See the full description on the dataset page: https://huggingface.co/datasets/milimlee-synth3d/objaverse.data.intelligence.chinese-ai-and-robotics-open-intelligence
🔬 Chinese AI, Humanoid Robotics & Neural Systems Open Intelligence Dataset
Curated open intelligence dataset tracking Chinese frontier developments in Large Language Models (LLMs), Humanoid Dynamic Locomotion, 3D Computer Vision, and Neuromorphic edge processors.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author institutional affiliations, and… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-ai-and-robotics-open-intelligence.chinese-biomedicine-and-genomics-open-intelligence
🔬 Chinese Biomedicine, Cell Therapy & Genomics Open Intelligence Dataset
Curated open intelligence dataset providing English briefs, clinical trial benchmarks, verified abstracts, and DOIs of frontier Chinese research in Cellular Therapeutics, Gene Editing, ADCs, and NMPA Clinical Approvals.
[!IMPORTANT]
Data Completeness & Research Authenticity Notice:
Included in this Hugging Face Open Dataset: English structured abstracts, core quantitative takeaways, author… See the full description on the dataset page: https://huggingface.co/datasets/simpleG2023/chinese-biomedicine-and-genomics-open-intelligence.so101_stack_cupsThis dataset was created using LeRobot.
Dataset Description
SO-101 (so_follower) teleoperation dataset for stacking cups. Contains 660 episodes / 188035 frames at 30 FPS, with three RGB cameras (observation.images.front, observation.images.top, observation.images.wrist) and 6-DoF joint state/action. Task prompt: "Stack the cups". Stored in LeRobot v3.0 format.
Homepage: https://huggingface.co/datasets/io-intelligence/so101_stack_cups
Paper: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/so101_stack_cups.IndustryInstruction_Artificial-Intelligence
IndustryInstruction: Artificial Intelligence
This repository contains the IndustryInstruction: Artificial Intelligence domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Artificial-Intelligence.WipeTable_DualArxR5a_TeleXperienceThis dataset was created using LeRobot.
Dataset Description
73 real-robot teleoperation episodes for “Wipe the table.” on a DualArxR5a dual-arm robot. Format: LeRobot v3.0 (30 Hz parquet + H.264 videos).
Collected with TeleXperience, IO-AI’s product for real-robot teleoperation and data collection.
Task / language prompt: Wipe the table.
Robot: DualArxR5a (bimanual, parallel-jaw grippers)
Frames: 629523 at 30 Hz
Cameras: camera_high (overhead), camera_low (lower scene)… See the full description on the dataset page: https://huggingface.co/datasets/io-intelligence/WipeTable_DualArxR5a_TeleXperience.details_Omartificial-Intelligence-Space__al-baka-llama3-8b-experimental
Dataset Card for Evaluation run of Omartificial-Intelligence-Space/al-baka-llama3-8b-experimental
Dataset automatically created during the evaluation run of model Omartificial-Intelligence-Space/al-baka-llama3-8b-experimental.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Omartificial-Intelligence-Space__al-baka-llama3-8b-experimental.Task_cyclo_intelligence_pickplacetip_lerobot_v21
Task_0001_pickplacetip_MCAP
Created with Cyclo Intelligence by ROBOTIS.
automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.egolongqa-synth-annotations
EgoLongQA synthetic MCQs, teacher traces and annotation outputs
Everything produced by the annotation and synthesis pipelines for the AI Wearables Challenge 2026
EgoLongQA ≤2B track, other than the distillation set (which lives in
infinitylogesh/egolongqa-junior-distill).
⚠️ Read this before counting rows
The synthetic set is 943 questions over 408 videos, and it is stored two ways:
file
rows
shape
training_sets/train_synth_v3.jsonl
943
flat — one row… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-synth-annotations.robot-intelligence-dataset
🤖 Robot Intelligence Dataset
A collection of three real-world robotics sensor datasets used to train and evaluate machine learning pipelines across four intelligence challenges:
Perception · Navigation · Failure Detection · Autonomous Decision-Making (RL).
Companion GitHub repo: KushSaraf/Robot_Intelligence
Made by: Kush Saraf & Yash Chavda
Datasets
🧠 1. Perception Dataset — Human Activity Recognition (HAR)
Source: UCI HAR Dataset
Property… See the full description on the dataset page: https://huggingface.co/datasets/kushsaraf/robot-intelligence-dataset.Cos-Play-Cold-Start
COS-PLAY Cold-Start Data
Pre-generated cold-start data for COS-PLAY (COLM 2026): Co-Evolving LLM Decision and Skill Bank Agents for Long-Horizon Game Play.
📄 Paper: arXiv:2604.20987 · HuggingFace Paper Page
💻 Code: github.com/wuxiyang1996/cos-play
🌐 Project page: wuxiyang1996.github.io/COSPLAY_page
🤖 Models: IntelligenceLab/COS-PLAY
Dataset Summary
This dataset contains GPT-5.4-generated seed trajectories and skill-labeled episodes for 8 games, used to bootstrap… See the full description on the dataset page: https://huggingface.co/datasets/IntelligenceLab/Cos-Play-Cold-Start.egolongqa-junior-distill
EgoLongQA junior-perceiver distillation data
Supervision for training a small VLM as the junior perceiver of an egocentric long-video
QA pipeline. Input protocol: ~100 UNIFORM frames of the whole video (max dim 768) + an MCQ
question. Output: a timestamped <video_description> followed by
<answer><choice>X</choice><reason>..</reason><citation>..</citation></answer>.
Configs
sft — teacher traces harvested from large-model agentic runs (122B / 27B / gemma juniors).… See the full description on the dataset page: https://huggingface.co/datasets/ambient-intelligence-labs/egolongqa-junior-distill.avian-climate-intelligencemasmils-global-mineral-resource-intelligence
Quick Links
Website
GitHub
Hugging Face
LinkedIn
X
MAS/MILS Global Mineral-Resource Intelligence, Modernized
1,846 mineral-property and processing-facility point features derived from
MAS/MILS (Minerals Availability System / Mineral Industry Location System),
originally built by the U.S. Bureau of Mines and transferred to USGS in
1996. No geological reinterpretation or machine-learning modelling is
performed — attributes and geometries… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/masmils-global-mineral-resource-intelligence.global-asset-market-cap-intelligence
Global Asset Market Capitalization Intelligence Dataset (GAMCID)
What is GAMCID?
GAMCID is a research-grade, machine-learning-ready dataset capturing the historical evolution of global assets ranked by market capitalization. It covers public companies, precious metals, cryptocurrencies, ETFs, and commodities — sourced from CompaniesMarketCap.com.
Why does it exist?
Existing financial datasets typically focus on single asset classes (stocks OR crypto… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/global-asset-market-cap-intelligence.ChenLong_Embodied_Intelligence_Dataset
ChenLong Embodied Intelligence Dataset
本仓库用于统一管理辰龙机器人实习中的数据集、模型权重、训练结果和说明文档。后续新增不同任务、采集批次、模型版本或实验资源时,都放在这里统一维护。
当前目录
embodied_dataset/:具身智能采集数据集,采用 LeRobot v3.0 结构,包含 data/、meta/、videos/。
yolo_dataset/:YOLO 目标检测数据、模型权重、训练参数和评估结果,当前包含 blue_bucket_yolov8/。
待新增新的数据集或模型。
新增数据集要求
具身数据优先采用 LeRobot v3.0 格式:meta/info.json、meta/stats.json、tasks、episodes、逐帧 Parquet 数据和按相机划分的视频。新增数据集至少写清:
任务:任务文本、目标物、成功标准、失败标准。
硬件:机器人型号、自由度、夹爪、相机位置、分辨率、FPS。… See the full description on the dataset page: https://huggingface.co/datasets/vvzc/ChenLong_Embodied_Intelligence_Dataset.liquidity-intelligence-benchmarks
VOIDTRACE AI Liquidity Intelligence Benchmarks
Benchmark dataset of 20 crypto liquidity intelligence cases with individual scores for liquidity flow, stablecoin intelligence, capital rotation, DEX activity, bridge activity, and ecosystem momentum across 8 blockchain networks.
Built by VOIDTRACE AI.
Dataset Description
This dataset contains benchmark data for the VOIDTRACE AI Crypto Liquidity Intelligence Engine — a blockchain intelligence software concept… See the full description on the dataset page: https://huggingface.co/datasets/voidtrace-ai/liquidity-intelligence-benchmarks.cyclo_intelligence_subtask_failed_dataset_test_lerobot_v21
Task_1_subtask_failed_dataset_test_MCAP
Created with Cyclo Intelligence by ROBOTIS.
subscriber-retention-intelligence
Subscriber Retention Intelligence
Complete transformed analytical corpus behind the Subscriber Retention Intelligence product. Release public-m12 contains 442,211,685 accepted rows in 32 Zstandard-compressed Parquet files.
Contents
Configuration
Grain
Rows
member
One row per subscriber
6,769,473
subscription_transaction
One row per accepted transaction
22,975,416
churn_label
One subscriber per observed label window
1,963,891
listening_day
One… See the full description on the dataset page: https://huggingface.co/datasets/vaibhavkhurana/subscriber-retention-intelligence.
