datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PhysicalAI-Autonomous-Vehicles
PHYSICAL AI AUTONOMOUS VEHICLES
The PhysicalAI-Autonomous-Vehicles dataset provides one of the largest, geographically diverse collections of multi-sensor data empowering AV researchers to build the next generation of Physical AI based end-to-end driving systems. This dataset is ready for commercial/non-commercial AV use per the license agreement.
Data Collection Method
Automatic/Sensor
Labeling Method
Automatic/Sensor
This dataset has a total of 1700 hours of driving… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles.resultsPhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios
Dataset Description:
PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.fev_datasets
Forecast evaluation datasets
This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models.
The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities.
The datasets follow a format that is compatible with the fev package.
Data format and usage
Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.AutoMathText-V2
🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset
🎉 AutoMathText-v2 has surpassed 1.5 million downloads! We'd love to know how you're using it. Please take 1 minute to fill out our use case survey. Your feedback will directly shape the future roadmap of this dataset.👉 Share your use case here
📊 AutoMathText-V2 consists of 2.46 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSQZ/AutoMathText-V2.chronos_datasets
Chronos datasets
Time series datasets used for training and evaluation of the Chronos forecasting models.
Note that some Chronos datasets (ETTh, ETTm, brazilian_cities_temperature and spanish_energy_and_weather) that rely on a custom builder script are available in the companion repo autogluon/chronos_datasets_extra.
See the paper for more information.
Data format and usage
The recommended way to use these datasets is via https://github.com/autogluon/fev.
All datasets… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets.PhysicalAI-Autonomous-Vehicles-NuRec
task_categories:
- robotics
tags:
- physicalAI
Find the 1500+ scenes in the sample_set/26.04_release folder.
Dataset Description:
Neural reconstructed dataset that carries 3D reconstructed driving scenes. The scenes are about 20 second long and stored in form of usdz files, along with respective xodr map files, surface mesh. The reconstructions were generated using 6 camera views (front-wide 120 deg, front-tele 30 deg, cross right/left 120 deg and rear right/left… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles-NuRec.AutoMathText-2.5
AutoMathText-2.5
🚀 AutoMathText-2.5: A Foundational High-Quality STEM Training Dataset
📊 AutoMathText-2.5 consists of over 2 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and bilingual data. This dataset was meticulously curated using a three-tier deduplication pipeline and AI-powered quality assessment to provide superior training data for large language models.
Our dataset combines 50+… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/AutoMathText-2.5.auto_evalAutoMathText🎉 This work, introducing the AutoMathText dataset and the AutoDS method, has been accepted to The 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Findings)! 🎉
AutoMathText
AutoMathText is an extensive and carefully curated dataset encompassing around 200 GB of mathematical texts. It's a compilation sourced from a diverse range of platforms including various websites, arXiv, and GitHub (OpenWebMath, RedPajama, Algebraic Stack). This rich repository… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/AutoMathText.PhysicalAI-Autonomous-Vehicles-NCore
PhysicalAI Autonomous Vehicles - NCore
A subset of ~1.1k clips from the PhysicalAI-Autonomous-Vehicles (PAI-AV)
dataset, converted to NCore format
using the PAI data converter.
The subset contains clips that expose accurate offline calibration,
egomotion, and cuboid labels.
Source Dataset
PhysicalAI-Autonomous-Vehicles
contains 306,152 clips (1,700 hours) of multi-sensor driving data
collected across 25 countries. See the source dataset card for full details
on… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles-NCore.PhysicalAI-Autonomous-Vehicle-Cosmos-Drive-Dreams
PhysicalAI-Autonomous-Vehicle-Cosmos-Drive-Dreams
Paper | Paper Website | GitHub
Download
We provide a download script to download our dataset. If you have enough space, you can use git to download a dataset from huggingface.
usage: download.py [-h] --odir ODIR
[--file_types {hdmap,lidar,synthetic}[,…]]
[--workers N] [--clean_cache]
required arguments:
--odir ODIR Output… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicle-Cosmos-Drive-Dreams.autoresearch-solo-vs-forum
Solo vs forum: long-horizon coding-agent runs on 12 research-engineering tasks
1334 runs (186 solo, 1148 forum), 9841 agent trials, 8975 transcripts, 209088 forum posts; 29772 files, 36.7 GB. Models: deepseek-v4.1-flash, glm-5.2, gpt-5.6-sol, qwen3.8-27b. Tasks: actlearn, adplace, borden, carleson, dabic, exploit, graph, kda, mega, moe, swinmlp, topopt.
What the experiment is
Each agent is a coding-agent CLI (Claude Code for Qwen3.8-27B / DeepSeek-V4.1-Flash /… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/autoresearch-solo-vs-forum.requestswiki_auto_asset_turk
Dataset Card for GEM/wiki_auto_asset_turk
Link to Main Data Card
You can find the main data card on the GEM Website.
Dataset Summary
WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.autonomous-vehicle-sensor-fusion
Autonomous Vehicle Raw Sensor Telemetry & Simulation Dataset
This repository contains raw, uncompressed sensor buffers, massive neural network weights, federated learning node dumps, and VRAM memory snapshots collected from autonomous vehicle test fleets. Data is provided "as is" for offline perception model training, system debugging, and Hardware-in-the-Loop (HIL) simulations.
Dataset Structure (Full Schema)
Due to horizontal scaling and massive daily ingestions… See the full description on the dataset page: https://huggingface.co/datasets/born5149/autonomous-vehicle-sensor-fusion.ModelNet40_Auto_aligned
ModelNet40 Auto Aligned
Auto-aligned version of the ModelNet40 3D CAD dataset. Each sample is an OFF mesh file organized by class and train/test split.
This dataset mirrors the layout of naderalfares/ModelNet40, but uses the auto-aligned meshes from the Princeton ModelNet release.
Dataset structure
modelnet40_auto_aligned/
{class}/
train/{class}_{id}.off
test/{class}_{id}.off
40 classes (airplane, bathtub, bed, …, xbox)
9,843 training meshes
2,468 test… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/ModelNet40_Auto_aligned.autoscirub-rcb-main-exp
AutoSciRub Main Experiments on ResearchClawBench
This repository is an artifact archive for the main ResearchClawBench experiments in:
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research AgentsXuehai Wang, Haowei Qin, Tongxin Liu, Junkai Li, Buqiang Xu, Jintian Zhang, Yijun Chen, Zirui Xue, and Shumin Deng.arXiv:2608.31076
It contains generated scientific reports, figures, code, supporting outputs, run metadata, and evaluator records for… See the full description on the dataset page: https://huggingface.co/datasets/huminclu/autoscirub-rcb-main-exp.Timeseries-PILE
Time Series PILE
The Time-series Pile is a large collection of publicly available data from diverse domains, ranging from healthcare to engineering and finance. It comprises of over 5
public time-series databases, from several diverse domains for time series foundation model pre-training and evaluation.
Time Series PILE Description
We compiled a large collection of publicly available datasets from diverse domains into the Time Series Pile. It has 13 unique domains of data… See the full description on the dataset page: https://huggingface.co/datasets/AutonLab/Timeseries-PILE.fluidgym-dataauto-video-public-media-relayminuszero-indian-autonomous-driving-dataset-v2
INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving
Overview
INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving.
This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.formal-math-autoformalization
Formal Math Autoformalization Dataset
A growing, CC0 public-domain corpus of ⟨natural-language statement ↔ Lean 4 statement + proof⟩ pairs, contributed through the Agentic Commons network.
Why this is scarce data. Mathlib already contains millions of proven Lean theorems — but as bare Lean, with no paired natural language:
theorem add_comm (a b : ℕ) : a + b = b + a := ... -- no "addition on naturals is commutative" attached
The scarce, valuable artifact is the pairing of the… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/formal-math-autoformalization.AutoGaze-Training-Dataarena-hard-auto
Arena-Hard-Auto
Repo for storing pre-generated model answers and judgment for
Arena-Hard-v0.1
Arena-Hard-v2.0-Preview
Repo -> https://github.com/lmarena/arena-hard-auto
Paper -> https://arxiv.org/abs/2406.11939
Citation
The code in this repository is developed from the papers below. Please cite it if you find the repository helpful.
@article{li2024crowdsourced,
title={From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline}… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-hard-auto.autofishThe AUTOFISH dataset comprises 1500 high-quality images of fish on a conveyor belt. It features 454 unique fish with class labels, IDs, manual length measurements,
and a total of 18,160 instance segmentation masks.
The fish are partitioned into 25 groups, with 14 to 24 fish in each group. Each fish only appears in one group, making it easy to create training splits. The
number of fish and distribution of species in each group were pseudo-randomly selected to mimic real-world scenarios.
Every… See the full description on the dataset page: https://huggingface.co/datasets/vapaau/autofish.Auto-ClawEval
Auto-ClawEval
Auto-generated agent evaluation benchmark with 1,040 tasks across 104 unique scenarios created by ClawEnvKit.
Statistics
Tasks
1,040
Categories
24
Mock services
20
Task types
API-based (77%) + file-dependent (23%)
Quick Start
# Download
huggingface-cli download AIcell/Auto-ClawEval --repo-type dataset --local-dir Auto-ClawEval
# Evaluate with ClawEnvKit (Docker harness)
bash run_harnesses.sh --harness claudecode… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/Auto-ClawEval.automoma-500kautoresearch-crypto-data
Binance public crypto market data
This public dataset contains typed, compressed, and audited copies of market data
published for free by Binance at https://data.binance.vision/. It is organized
for reproducible point-in-time research across the markets represented in the
coverage report. The production backfill stores the complete compact aggregate
layer for a pinned liquid USD-M and COIN-M futures universe, plus every option
index and option-surface underlying published in the… See the full description on the dataset page: https://huggingface.co/datasets/tmmycruise/autoresearch-crypto-data.chronos_datasets_extra
Chronos datasets
Time series datasets used for training and evaluation of the Chronos forecasting models.
This repository contains scripts for constructing datasets that cannot be hosted in the main Chronos datasets repository due to license restrictions.
Usage
Datasets can be loaded using the 🤗 datasets library
import datasets
ds = datasets.load_dataset("autogluon/chronos_datasets_extra", "ETTh", split="train", trust_remote_code=True)
ds.set_format("numpy") #… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets_extra.
