CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /PhysicalAI-Autonomous-Vehiclesgated PHYSICAL AI AUTONOMOUS VEHICLES The PhysicalAI-Autonomous-Vehicles dataset provides one of the largest, geographically diverse collections of multi-sensor data empowering AV researchers to build the next generation of Physical AI based end-to-end driving systems. This dataset is ready for commercial/non-commercial AV use per the license agreement. Data Collection Method Automatic/Sensor Labeling Method Automatic/Sensor This dataset has a total of 1700 hours of driving… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles.1.1k likes211k downloads11d agoHugging Face02Autonomous-Scientific-Agents /results1 likes55k downloads1mo agoHugging Face03nvidia /PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios Dataset Description: PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios is a large-scale synthetic video dataset of autonomous-driving scenes generated with NVIDIA's internal Omniverse simulation platform. Each clip is a temporally consistent multi-camera surround capture of one ego vehicle and surrounding traffic participants, paired with per-camera VLM captions. The dataset is designed to fill gaps in real-world driving data along two axes: (1) targeted long-tail… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-WorldModel-Synthetic-Autonomous-Driving-Scenarios.video100K<n<1M26 likes48k downloads4mo agoHugging Face04autogluon /fev_datasets Forecast evaluation datasets This repository contains time series datasets that can be used for evaluation of univariate & multivariate forecasting models. The main focus of this repository is on datasets that reflect real-world forecasting scenarios, such as those involving covariates, missing values, and other practical complexities. The datasets follow a format that is compatible with the fev package. Data format and usage Each dataset satisfies the following… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/fev_datasets.tabulartime-series-forecasting100K<n<1M13 likes43k downloads8mo agoHugging Face05OpenSQZ /AutoMathText-V2 🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset &nbsp; 🎉 AutoMathText-v2 has surpassed 1.5 million downloads! We'd love to know how you're using it. Please take 1 minute to fill out our use case survey. Your feedback will directly shape the future roadmap of this dataset.👉 Share your use case here 📊 AutoMathText-V2 consists of 2.46 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSQZ/AutoMathText-V2.tabulartext-generation1B<n<10B78 likes42k downloads4mo agoHugging Face06autogluon /chronos_datasets Chronos datasets Time series datasets used for training and evaluation of the Chronos forecasting models. Note that some Chronos datasets (ETTh, ETTm, brazilian_cities_temperature and spanish_energy_and_weather) that rely on a custom builder script are available in the companion repo autogluon/chronos_datasets_extra. See the paper for more information. Data format and usage The recommended way to use these datasets is via https://github.com/autogluon/fev. All datasets… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets.tabulartime-series-forecasting10M<n<100M76 likes33k downloads2y agoHugging Face07nvidia /PhysicalAI-Autonomous-Vehicles-NuRecgated task_categories: - robotics tags: - physicalAI Find the 1500+ scenes in the sample_set/26.04_release folder. Dataset Description: Neural reconstructed dataset that carries 3D reconstructed driving scenes. The scenes are about 20 second long and stored in form of usdz files, along with respective xodr map files, surface mesh. The reconstructions were generated using 6 camera views (front-wide 120 deg, front-tele 30 deg, cross right/left 120 deg and rear right/left… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles-NuRec.239 likes27k downloads2mo agoHugging Face08math-ai /AutoMathText-2.5 AutoMathText-2.5 🚀 AutoMathText-2.5: A Foundational High-Quality STEM Training Dataset &nbsp; 📊 AutoMathText-2.5 consists of over 2 trillion tokens of high-quality, deduplicated text spanning web content, mathematics, code, reasoning, and bilingual data. This dataset was meticulously curated using a three-tier deduplication pipeline and AI-powered quality assessment to provide superior training data for large language models. Our dataset combines 50+… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/AutoMathText-2.5.text-generation10B<n<100B2 likes26k downloads2mo agoHugging Face09zhouzypaul /auto_evaltext1K<n<10K0 likes17k downloads3mo agoHugging Face10math-ai /AutoMathText🎉 This work, introducing the AutoMathText dataset and the AutoDS method, has been accepted to The 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025 Findings)! 🎉 AutoMathText AutoMathText is an extensive and carefully curated dataset encompassing around 200 GB of mathematical texts. It's a compilation sourced from a diverse range of platforms including various websites, arXiv, and GitHub (OpenWebMath, RedPajama, Algebraic Stack). This rich repository… See the full description on the dataset page: https://huggingface.co/datasets/math-ai/AutoMathText.texttext-generation1M<n<10M186 likes13k downloads1y agoHugging Face11nvidia /PhysicalAI-Autonomous-Vehicles-NCoregated PhysicalAI Autonomous Vehicles - NCore A subset of ~1.1k clips from the PhysicalAI-Autonomous-Vehicles (PAI-AV) dataset, converted to NCore format using the PAI data converter. The subset contains clips that expose accurate offline calibration, egomotion, and cuboid labels. Source Dataset PhysicalAI-Autonomous-Vehicles contains 306,152 clips (1,700 hours) of multi-sensor driving data collected across 25 countries. See the source dataset card for full details on… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicles-NCore.28 likes12k downloads6mo agoHugging Face12nvidia /PhysicalAI-Autonomous-Vehicle-Cosmos-Drive-Dreams PhysicalAI-Autonomous-Vehicle-Cosmos-Drive-Dreams Paper | Paper Website | GitHub Download We provide a download script to download our dataset. If you have enough space, you can use git to download a dataset from huggingface. usage: download.py [-h] --odir ODIR [--file_types {hdmap,lidar,synthetic}[,…]] [--workers N] [--clean_cache] required arguments: --odir ODIR Output… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Autonomous-Vehicle-Cosmos-Drive-Dreams.roboticsn>1T61 likes11k downloads1y agoHugging Face13junlinw /autoresearch-solo-vs-forum Solo vs forum: long-horizon coding-agent runs on 12 research-engineering tasks 1334 runs (186 solo, 1148 forum), 9841 agent trials, 8975 transcripts, 209088 forum posts; 29772 files, 36.7 GB. Models: deepseek-v4.1-flash, glm-5.2, gpt-5.6-sol, qwen3.8-27b. Tasks: actlearn, adplace, borden, carleson, dabic, exploit, graph, kda, mega, moe, swinmlp, topopt. What the experiment is Each agent is a coding-agent CLI (Claude Code for Qwen3.8-27B / DeepSeek-V4.1-Flash /… See the full description on the dataset page: https://huggingface.co/datasets/junlinw/autoresearch-solo-vs-forum.1K<n<10K0 likes9.6k downloads2d agoHugging Face14Autonomous-Scientific-Agents /requests1 likes9k downloads1mo agoHugging Face15GEM /wiki_auto_asset_turk Dataset Card for GEM/wiki_auto_asset_turk Link to Main Data Card You can find the main data card on the GEM Website. Dataset Summary WikiAuto is an English simplification dataset that we paired with ASSET and TURK, two very high-quality evaluation datasets, as test sets. The input is an English sentence taken from Wikipedia and the target a simplified sentence. ASSET and TURK contain the same test examples but have references that are simplified in different… See the full description on the dataset page: https://huggingface.co/datasets/GEM/wiki_auto_asset_turk.text100K<n<1M8 likes8.8k downloads2y agoHugging Face16born5149 /autonomous-vehicle-sensor-fusion Autonomous Vehicle Raw Sensor Telemetry & Simulation Dataset This repository contains raw, uncompressed sensor buffers, massive neural network weights, federated learning node dumps, and VRAM memory snapshots collected from autonomous vehicle test fleets. Data is provided "as is" for offline perception model training, system debugging, and Hardware-in-the-Loop (HIL) simulations. Dataset Structure (Full Schema) Due to horizontal scaling and massive daily ingestions… See the full description on the dataset page: https://huggingface.co/datasets/born5149/autonomous-vehicle-sensor-fusion.2 likes7.6k downloads15d agoHugging Face17naderalfares /ModelNet40_Auto_aligned ModelNet40 Auto Aligned Auto-aligned version of the ModelNet40 3D CAD dataset. Each sample is an OFF mesh file organized by class and train/test split. This dataset mirrors the layout of naderalfares/ModelNet40, but uses the auto-aligned meshes from the Princeton ModelNet release. Dataset structure modelnet40_auto_aligned/ {class}/ train/{class}_{id}.off test/{class}_{id}.off 40 classes (airplane, bathtub, bed, …, xbox) 9,843 training meshes 2,468 test… See the full description on the dataset page: https://huggingface.co/datasets/naderalfares/ModelNet40_Auto_aligned.3dimage-classification10K<n<100K0 likes7k downloads3mo agoHugging Face18huminclu /autoscirub-rcb-main-exp AutoSciRub Main Experiments on ResearchClawBench This repository is an artifact archive for the main ResearchClawBench experiments in: Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research AgentsXuehai Wang, Haowei Qin, Tongxin Liu, Junkai Li, Buqiang Xu, Jintian Zhang, Yijun Chen, Zirui Xue, and Shumin Deng.arXiv:2608.31076 It contains generated scientific reports, figures, code, supporting outputs, run metadata, and evaluator records for… See the full description on the dataset page: https://huggingface.co/datasets/huminclu/autoscirub-rcb-main-exp.imagen<1K1 likes6.4k downloads20d agoHugging Face19AutonLab /Timeseries-PILE Time Series PILE The Time-series Pile is a large collection of publicly available data from diverse domains, ranging from healthcare to engineering and finance. It comprises of over 5 public time-series databases, from several diverse domains for time series foundation model pre-training and evaluation. Time Series PILE Description We compiled a large collection of publicly available datasets from diverse domains into the Time Series Pile. It has 13 unique domains of data… See the full description on the dataset page: https://huggingface.co/datasets/AutonLab/Timeseries-PILE.time-series-forecasting43 likes5.7k downloads2y agoHugging Face20safe-autonomous-systems /fluidgym-data0 likes5.5k downloads6mo agoHugging Face21churchill91123 /auto-video-public-media-relayimage100K<n<1M1 likes5.3k downloads7h agoHugging Face22gagandeepreehal /minuszero-indian-autonomous-driving-dataset-v2gated INDUS-AD: Indian Dataset of Unstructured Urban Scenes for Autonomous Driving Overview INDUS-AD is the largest publicly released Indian autonomous-driving dataset for end-to-end autonomous-driving research. Its name expands to Indian Dataset of Unstructured Urban Scenes for Autonomous Driving. This gated dataset is the decoded companion to the Minus Zero Indian Urban Autonomous Driving Dataset. It provides directly usable camera MP4s, normalized sensor tables… See the full description on the dataset page: https://huggingface.co/datasets/gagandeepreehal/minuszero-indian-autonomous-driving-dataset-v2.imagerobotics10M<n<100M1 likes5.1k downloads4d agoHugging Face23AgenticCommons /formal-math-autoformalization Formal Math Autoformalization Dataset A growing, CC0 public-domain corpus of ⟨natural-language statement ↔ Lean 4 statement + proof⟩ pairs, contributed through the Agentic Commons network. Why this is scarce data. Mathlib already contains millions of proven Lean theorems — but as bare Lean, with no paired natural language: theorem add_comm (a b : ℕ) : a + b = b + a := ... -- no "addition on naturals is commutative" attached The scarce, valuable artifact is the pairing of the… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/formal-math-autoformalization.texttext-generation1K<n<10K3 likes4.4k downloads2h agoHugging Face24bfshi /AutoGaze-Training-Data0 likes4.4k downloads7mo agoHugging Face25lmarena-ai /arena-hard-auto Arena-Hard-Auto Repo for storing pre-generated model answers and judgment for Arena-Hard-v0.1 Arena-Hard-v2.0-Preview Repo -> https://github.com/lmarena/arena-hard-auto Paper -> https://arxiv.org/abs/2406.11939 Citation The code in this repository is developed from the papers below. Please cite it if you find the repository helpful. @article{li2024crowdsourced, title={From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline}… See the full description on the dataset page: https://huggingface.co/datasets/lmarena-ai/arena-hard-auto.8 likes3.8k downloads1y agoHugging Face26vapaau /autofishThe AUTOFISH dataset comprises 1500 high-quality images of fish on a conveyor belt. It features 454 unique fish with class labels, IDs, manual length measurements, and a total of 18,160 instance segmentation masks. The fish are partitioned into 25 groups, with 14 to 24 fish in each group. Each fish only appears in one group, making it easy to create training splits. The number of fish and distribution of species in each group were pseudo-randomly selected to mimic real-world scenarios. Every… See the full description on the dataset page: https://huggingface.co/datasets/vapaau/autofish.imageimage-segmentationn<1K3 likes3.4k downloads2y agoHugging Face27AIcell /Auto-ClawEval Auto-ClawEval Auto-generated agent evaluation benchmark with 1,040 tasks across 104 unique scenarios created by ClawEnvKit. Statistics Tasks 1,040 Categories 24 Mock services 20 Task types API-based (77%) + file-dependent (23%) Quick Start # Download huggingface-cli download AIcell/Auto-ClawEval --repo-type dataset --local-dir Auto-ClawEval # Evaluate with ClawEnvKit (Docker harness) bash run_harnesses.sh --harness claudecode… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/Auto-ClawEval.imageother1K<n<10K2 likes3.1k downloads5mo agoHugging Face28automoma /automoma-500k0 likes2.8k downloads5mo agoHugging Face29tmmycruise /autoresearch-crypto-data Binance public crypto market data This public dataset contains typed, compressed, and audited copies of market data published for free by Binance at https://data.binance.vision/. It is organized for reproducible point-in-time research across the markets represented in the coverage report. The production backfill stores the complete compact aggregate layer for a pinned liquid USD-M and COIN-M futures universe, plus every option index and option-surface underlying published in the… See the full description on the dataset page: https://huggingface.co/datasets/tmmycruise/autoresearch-crypto-data.0 likes2.7k downloads26d agoHugging Face30autogluon /chronos_datasets_extra Chronos datasets Time series datasets used for training and evaluation of the Chronos forecasting models. This repository contains scripts for constructing datasets that cannot be hosted in the main Chronos datasets repository due to license restrictions. Usage Datasets can be loaded using the 🤗 datasets library import datasets ds = datasets.load_dataset("autogluon/chronos_datasets_extra", "ETTh", split="train", trust_remote_code=True) ds.set_format("numpy") #… See the full description on the dataset page: https://huggingface.co/datasets/autogluon/chronos_datasets_extra.time-series-forecasting10 likes2.3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.