datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
glue
Dataset Card for GLUE
Dataset Summary
GLUE, the General Language Understanding Evaluation benchmark (https://gluebenchmark.com/) is a collection of resources for training, evaluating, and analyzing natural language understanding systems.
Supported Tasks and Leaderboards
The leaderboard for the GLUE benchmark can be found at this address. It comprises the following tasks:
ax
A manually-curated evaluation dataset for fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/nyu-mll/glue.HelpSteer2
HelpSteer2: Open-source dataset for training top-performing reward models
HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses.
This dataset has been created in partnership with Scale AI.
When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.china-a-share-1min-ohlcv
China A-Share Equities 1-Minute OHLCV
Minute-level OHLCV bars for exchange-listed Chinese A-share equities. The release uses a stable Parquet schema, one canonical file per instrument, and machine-readable coverage reports.
Dataset summary
This snapshot contains 3,475,824,481 rows for 5,795 instruments across China A-share equities on the Shanghai, Shenzhen, and Beijing exchanges. It covers 2010-01-04 09:30:00 through 2026-08-07 10:21:00. Prices are unadjusted.… See the full description on the dataset page: https://huggingface.co/datasets/neigezhu/china-a-share-1min-ohlcv.infini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.llm-network-study-data
LLM-Network-Study-Data
Per-request network captures (.pcapng) collected by the
LLM-Network-Study benchmark harness (benchmark.py and the
per-workload test scripts). Each directory holds one capture file per request,
named request_<id>_run<n>_<timestamp>.pcapng.
A directory name encodes four dimensions:
<capture-env>_<provider/model>_<workload>[_<dataset/variant>]_results
Dimension legend
Dimension
Values
Meaning
Capture env
ethernet
Wired connection to… See the full description on the dataset page: https://huggingface.co/datasets/wayslab/llm-network-study-data.PhysicalAI-Robotics-GR00T-Teleop-Sim
Simulation GR1 Tabletop Task 1K Dataset
Dataset Description:
The PhysicalAI-Robotics-GR00T-Teleop-GR1 dataset consists of 1000 teleoperation trajectories in simulation using the GR1 robot with upper body control. The simulation setup mimics tabletop manipulation tasks and uses RGB observations with a virtual camera. The robot is equipped with simulated Fourier hands.
This dataset is ready for non-commercial use.
Dataset Owner(s):
NVIDIA GEAR… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-Sim.PhysicalAI-Robotics-GR00T-Teleop-GR1
Introduction
TL;DR: DreamDojo is a generalist robot world model pretrained on 44k hours of human egocentric data, showing unprecedented generalization to diverse objects and environments.
Project page: https://dreamdojo-world.github.io/
Paper: https://arxiv.org/abs/2602.06949
Code: https://github.com/NVIDIA/DreamDojo
How to Use
Check out https://github.com/NVIDIA/DreamDojo
Citation
@article{gao2026dreamdojo,
title={DreamDojo: A Generalist Robot… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/PhysicalAI-Robotics-GR00T-Teleop-GR1.osm-polygon-selection
osm-polygon-selection dataset
A curated set of OpenStreetMap polygons from 310
geographic units — sovereign countries plus sub-country regions
like Brazilian states, Chinese provinces, Indian zones, US states,
Canadian provinces, Japanese regions, and Indonesian islands —
classified by size bin (small / medium / large, area in
[0.1, 100] km²) and tagged by continent (Natural Earth admin0 lookup).
Size bins:
small — area in [0.1, 1) km² (10,000 m² to 1 km², roughly
100 m × 100 m… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-selection.Aiice
Dataset
Aiice benchmark dataset for Arctic sea ice concentration (SIC) forecasting,
based on OSI-SAF satellite products (CC BY 4.0).
Coverage
Period: October 1978 – April 2026
Resolution: 25 km spatial, daily temporal
Grid: 432×432 (Lambert Azimuthal Equal Area, EPSG:6931)
Source products
Product
Source
Period
OSI-450-a
SMMR, SSM/I, SSMIS
1978–2020
OSI-430-a
SSMIS
2021–Jul 2025
OSI-438
AMSR2
Jul 2025–present… See the full description on the dataset page: https://huggingface.co/datasets/ITMO-NSS/Aiice.Calc-asdiv_a
Dataset Card for Calc-asdiv_a
Summary
The dataset is a collection of simple math word problems focused on arithmetics. It is derived from the arithmetic subset of ASDiv (original repo).
The main addition in this dataset variant is the chain column. It was created by converting the solution to a simple html-like language that can be easily
parsed (e.g. by BeautifulSoup). The data contains 3 types of tags:
gadget: A tag whose content is intended to be evaluated by calling… See the full description on the dataset page: https://huggingface.co/datasets/MU-NLPC/Calc-asdiv_a.NECL_GPUslegalbench
Dataset Card for Dataset Name
Homepage: https://hazyresearch.stanford.edu/legalbench/
Repository: https://github.com/HazyResearch/legalbench/
Paper: https://arxiv.org/abs/2308.11462
Dataset Description
Dataset Summary
The LegalBench project is an ongoing open science effort to collaboratively curate tasks for evaluating legal reasoning in English large language models (LLMs). The benchmark currently consists of 162 tasks gathered from 40… See the full description on the dataset page: https://huggingface.co/datasets/nguha/legalbench.new-york-smells
New York Smells: A Large Multimodal Dataset for Olfaction
While olfaction is central to how animals perceive the world, this rich chemical
sensory modality remains largely inaccessible to machines. One key bottleneck is the
lack of diverse, multimodal olfactory data collected in natural settings. We present
New York Smells, a large-scale dataset of paired image and olfactory signals
captured in-the-wild. Our dataset contains 7,000 smell-image pairs from 3,500 distinct
objects… See the full description on the dataset page: https://huggingface.co/datasets/cvlab/new-york-smells.CodeTraceBenchCodeTraceBench
A Benchmark for Agent Trajectory Diagnosis
CodeTraceBench is a large-scale benchmark of 4,316 agent trajectories with human-verified step-level annotations for evaluating trajectory diagnosis systems. Each trajectory records the full action-observation sequence of a coding agent, annotated with incorrect and unuseful step labels.
Part of the CodeTracer project — a self-evolving agent trajectory diagnosis system.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/NJU-LINK/CodeTraceBench.nexar_collision_prediction
Nexar Collision Prediction Dataset
This dataset is part of the Nexar Dashcam Crash Prediction Challenge on Kaggle.
Dataset
The Nexar collision prediction dataset comprises videos from Nexar dashcams. Videos have a resolution of 1280x720 at 30 frames per second and typically have about 40 seconds of duration. The dataset contains 1500 videos where half show events where there was a collision or a collision was eminent (positive cases), and the other half shows… See the full description on the dataset page: https://huggingface.co/datasets/nexar-ai/nexar_collision_prediction.nrvbench-review
NR Video Editing Benchmark
This repository contains two non-rigid video editing benchmark subsets for evaluating instruction-driven video editing methods. Each row in metadata.csv corresponds to one editing instruction for a source video, with relative paths to the source video, extracted frames, binary masks, prompts, and evaluation questions.
The dataset card is written without author or institution identifiers so it can be used for anonymous review uploads. Before a non-anonymous… See the full description on the dataset page: https://huggingface.co/datasets/NRVBench/nrvbench-review.osm-polygon-wikidata-only
OSM Polygon Wikidata, Wikipedia and Wikivoyage
OSM polygons carrying wikidata=*, enriched with multilingual Wikipedia and Wikivoyage documents. The published tables preserve regional records and provenance.
Source code: GitHub repository.
Dataset snapshot
Metric
Value
Polygon rows across regional extracts
1,184,110
Unique polygon identities (osm_type, osm_id)
1,157,841
Polygons with successful non-empty text (unique OSM identities)
650,663… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/osm-polygon-wikidata-only.SWE-rebench-leaderboard
Dataset Summary
❗❗❗ Please use Harbour Hub for the July 2026 evaluation split:https://hub.harborframework.com/datasets/ibragim-badertdinov/swe-rebench-07-2026/latest
SWE-rebench-leaderboard is a continuously updated, curated subset of the full SWE-rebench corpus, tailored for benchmarking software engineering agents on real-world tasks.
These tasks are used in the SWE-rebench leaderboard. For more details on the benchmark methodology and data collection process, please refer to… See the full description on the dataset page: https://huggingface.co/datasets/nebius/SWE-rebench-leaderboard.19c_newspapers_images_altohelaxai_data_pluse
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/nhblk123/helaxai_data_pluse.Nemotron-ClimbMix
ClimbMix Dataset
🚀 Creating the highest-quality pre-training datasets for LLMs 🌟
📄 PAPER
🤗 CLIMBLAB
🤗 CLIMBMIX
🏠 HOMEPAGE
Figure 1: Continuously training a 1B model yields a 2.0% improvement over Llama-3.2-1B, demonstrating a more efficient scaling trend compared to prior models.
Figure 2: Pre-training a 1B model from scratch on ClimbMix shows better scaling effects than training on other datasets.… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-ClimbMix.hle-gpt-oss-120b-no-python-260222
hle-gpt-oss-120b-no-python-260222
Deep research agent evaluation on rl-rag/hle_text_only (test split).
Results
Metric
Value
pass@4
47.9%
avg@4
26.6%
Trajectory accuracy
26.6% (2292/8632)
Questions
2158
Trajectories
8632 (4 per question)
Avg tool calls
14.5
Full conversations
❌
Model & Setup
Model
gpt-oss-120b
Judge
gpt-4o
Max tool calls
50
Temperature
0.7
Blocked domains
huggingface.co
Tool Usage… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/hle-gpt-oss-120b-no-python-260222.libero_spatial_no_noops_1.0.0_lerobotFrench-PD-Newspapers
🇫🇷 French Public Domain Newspapers 🇫🇷
French-Public Domain-Newspapers or French-PD-Newpapers is a large collection aiming to agregate all the French newspapers and periodicals in the public domain.
The collection has been originally compiled by Pierre-Carl Langlais, on the basis of a large corpus curated by Benoît de Courson, Benjamin Azoulay for Gallicagram and in cooperation with OpenLLMFrance. Gallicagram is leading cultural analytics project giving access to word and ngram… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-PD-Newspapers.FakeCOCO
FakeCOCO dataset
Using 10 SOTA text-to-image models to generate fake images based on COCO captions
over 1M images
These models include:
SD15
SD21
SDXL
SD3
Playground2.5
PixArt alpha
PixArt sigma
unidiffuser
Flux.1
Stable Cascade
24-game
Math Twenty Four (24s Game) Dataset
A comprehensive dataset for the classic math twenty four game (also known as the 4 numbers game / 24s game / Game of 24). This dataset of mathematical reasoning challenges was collected from 4nums.com, featuring over 1,300 unique puzzles of the Game of 24, with difficulty metrics derived from over 6.4 million human solution attempts since 2012.
In each puzzle, players must use exactly four numbers and basic arithmetic operations (+, -, ×, /) to… See the full description on the dataset page: https://huggingface.co/datasets/nlile/24-game.libero_goal_no_noops_1.0.0_lerobotlibero_object_no_noops_1.0.0_lerobotlibero_10_no_noops_1.0.0_lerobotNetEaseCrowd
🧑🤝🧑 NetEaseCrowd: A Dataset for Long-term and Online Crowdsourcing Truth Inference
View it in GitHub
Introduction
We introduce NetEaseCrowd, a large-scale crowdsourcing annotation dataset based on
a mature Chinese data crowdsourcing platform of NetEase Inc..
NetEaseCrowd dataset contains about 2,400 workers, 1,000,000 tasks, and 6,000,000 annotations between them,
where the annotations are collected in about 6 months.
In this dataset, we provide ground truths for… See the full description on the dataset page: https://huggingface.co/datasets/liuhyuu/NetEaseCrowd.
