datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Polaris
Polaris Dataset
🌟 [CVPR24] Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
Accepted at CVPR 2024
🌐 project page
📄 arXiv
🤗 Dataset
Establishing an automatic evaluation metric that closely aligns with human judgements is essential for the effective development of image captioning models. Data-driven metrics have recently gained prominence in this field, demonstrating a stronger correlation with human judgements than classic metrics such as CIDEr and… See the full description on the dataset page: https://huggingface.co/datasets/yuwd/Polaris.Polaris-Dataset-53K
Overview
Training dataset for Polaris Preview models. The dataset is filtered from DeepScaleR-Preview-Dataset and AReal-boba-Data
Format
Each row in the jsonl file contains:
problem: The input problem.
answer: The answer to the problem
difficulty: The pass rate of the problem estimated by Deepseek-R1-distill-Qwen-7B
Citation
@misc{Polaris2025,
title = {POLARIS: A Post-Training Recipe for Scaling Reinforcement Learning on Advanced Reasoning Models}… See the full description on the dataset page: https://huggingface.co/datasets/POLARIS-Project/Polaris-Dataset-53K.PolaRiS-Hub
PolaRiS Hub
PolaRiS Hub is a lightweight collection of reconstructed real-to-sim manipulation environments used for evaluating generalist robot policies in simulation.
It is designed to be used with the PolaRiS evaluation framework.
Instructions on how to use can be found at the Github PolaRiS Codebase
What’s in the Hub
Each environment in PolaRiS Hub includes:
A reconstructed scene (mesh + Gaussian splats)
Canonical initial conditions
A language instruction… See the full description on the dataset page: https://huggingface.co/datasets/owhan/PolaRiS-Hub.PolaRiS-datasets
Updates
As of 12/19/2025, this cotraining dataset has been updated to contain the OOD dataset as described the paper.
polaris-bench
The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space
Polaris-Bench: Official Evaluation Dataset
Xia Hu1,
Zhenrui Yue1,
Brian Potetz1,
Howard Zhou1,
Leonidas Guibas1,2,
Chun-Ta Lu3,
Zhicheng Wang1
1Google DeepMind 2Stanford University 3Google Research
Overview
As current Multimodal Large Language Models (MLLMs) rapidly saturate canonical visual reasoning benchmarks, a key… See the full description on the dataset page: https://huggingface.co/datasets/google/polaris-bench.bento_ur7e_v1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 60,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_0.pos",
"joint_1.pos",
"joint_2.pos",
"joint_3.pos",
"joint_4.pos",
"joint_5.pos",
"gripper.pos"
],
"shape": [… See the full description on the dataset page: https://huggingface.co/datasets/polarisai-robots/bento_ur7e_v1.maxrl_qwen3_4B_base_polaris_rollouts
MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts)
Every training rollout from an online RL run, with exact token ids, sampling
log-probs, and raw rewards — usable as a replay buffer to study off-policy RL
for LLM reasoning completely offline.
The run: Qwen3-4B-Base trained with the maxRL advantage estimator
(A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt;
maxRL paper) and a pure REINFORCE loss
(L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.PolaRiS-Hub
PolaRiS Hub
PolaRiS Hub is a lightweight collection of reconstructed real-to-sim manipulation environments used for evaluating generalist robot policies in simulation.
It is designed to be used with the PolaRiS evaluation framework.
Instructions on how to use can be found at the Github PolaRiS Codebase
What’s in the Hub
Each environment in PolaRiS Hub includes:
A reconstructed scene (mesh + Gaussian splats)
Canonical initial conditions
A language instruction describing… See the full description on the dataset page: https://huggingface.co/datasets/MaximumY/PolaRiS-Hub.qwen3-1.7b-polaris-fp8-rollouts-20260912
Qwen3-1.7B Polaris FP8 rollouts
Snapshot of three FP8 rollout datasets taken on 2026-09-12. The model is Qwen3-1.7B-Base. Original Parquet files and attempt, file, and checkpoint-lineage metadata are preserved without rewriting.
Configuration
Training batches
Training responses
Validation responses
Total bytes
maxrl_strict
101
827392
144320
6654288117
maxrl_permissive
107
876544
144320
8168064932
dppo
126
1032192
173184
6733682534
Provenance and… See the full description on the dataset page: https://huggingface.co/datasets/steviel/qwen3-1.7b-polaris-fp8-rollouts-20260912.zjumocapPolaris-Qwen3-4B-Instruct-eval-32k-64_partialpolaris_imagereward_v3rloo_qwen3_1p7B_base_polaris_rolloutspolaris_imagerewardpolaris_droid_cotrainThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 315,
"total_frames": 49528,
"total_tasks": 17,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:315"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/imishutk0410/polaris_droid_cotrain.polaris-trajectory-v2-token-entropy
Polaris Trajectory-v2 Token Entropy
Teacher-forced token entropy for all 4,572 complete endpoints represented in the
Polaris trajectory-v2 PRM dataset (4,379 train and 193 validation trajectories).
Scores use the exact Qwen3.5-4B text tower that collected the trajectories.
Fork definition
Canonical forks are reasoning tokens whose raw full-vocabulary entropy is greater
than or equal to a threshold. The four global thresholds (q80, q90, q95, and
q99) are fitted… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/polaris-trajectory-v2-token-entropy.ASAP_Polaris_OpenADMET_challenge
Raw datasets for the ASAP-Polaris-OpenADMET blind challenge.
This repo contains the datasets used in the ASAP-Polaris-OpenADMET challenge, comprising Potency, ADMET and crystallography data from ASAP Discovery's SARS-CoV-2 / MERS-CoV main protease inhibitor program.
More detail on the competition and its design can be found at the relevant preprint.
Data is assigned to a train or test split based on criteria laid out in the above preprint. For the tabular data (Potency and… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/ASAP_Polaris_OpenADMET_challenge.polaris-ecir-v1
ECIR v1
ECIR is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, arctic, lter, wikitables, and
wtr.
It holds 2,100 tables published on the US government open data portal and 12 keyword queries over
them. For each query–table pair, a person decided how well that table answers that query and gave it
a score; those scores are the relevance judgments, and they live in qrels.csv. Given a query, a
system ranks the 2,100… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-ecir-v1.polaris-wikitables-v2
WikiTables v2
WikiTables is one of six datasets in Polaris: Learning to Generate Table Descriptions from
Retrieval Feedback, alongside aw, arctic, lter, ecir,
and wtr.
It holds 3,361 tables scraped from Wikipedia articles and 57 keyword queries over them. For each
query–table pair, a person scored how well that table answers that query; those scores are the
relevance judgments, and they live in qrels.csv.
The tables have no names. What describes a table is its column names and… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-wikitables-v2.polaris_imagereward_v2leaderboard-datamonitorability-as-a-free-gift-data-training-data
Monitorability as a free gift training data reordered, or normalized during packaging.
Configurations
Config
Rows
Purpose
Original file
all
18,591
Main all-domain experiment
combined_dataset.parquet
no_if
13,591
All-domain experiment without instruction following
combined_dataset_noif.parquet
instruction_following
5,000
Instruction-following experiments
instruction_following_ai2_5000.parquet
math
5,000
Main math experiments
skywork_math.parquet… See the full description on the dataset page: https://huggingface.co/datasets/polaris-73/monitorability-as-a-free-gift-data-training-data.Polaris-Hard
Polaris-Hard
A random subset of Polaris-Dataset-53K, with 3,200 distinct math problems in a single train split.
Original difficulty
Source pool
Eligible pool
Selected
Share
0/8
15,368
9,331
2,000
62.5%
1/8
6,956
4,929
1,200
37.5%
Total
22,324
14,260
3,200
100%
Sampling
Source revision: 296f8e34132e63f4a1d70e0dcc8bddebb43f03e4.
Seed: 42. Uniform random sampling without replacement within each group, followed by a deterministic shuffle of the… See the full description on the dataset page: https://huggingface.co/datasets/hi-todayis-jh/Polaris-Hard.polaris-acemath-gemini-rubrics-v2bento_v2_openarmThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"joint_1.pos",
"joint_1.vel",
"joint_1.torque",
"joint_2.pos",
"joint_2.vel",
"joint_2.torque",
"joint_3.pos",
"joint_3.vel"… See the full description on the dataset page: https://huggingface.co/datasets/polarisai-robots/bento_v2_openarm.Polaris-cleanfaithful-thinking-draft
Dataset Card for Thinking Draft Faithfulness Evaluation
This dataset accompanies the paper "Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models".
Dataset Description
The Faithful Thinking Draft dataset is designed to evaluate how faithfully language models follow their own thinking drafts. It contains benchmarks for two key aspects of reasoning faithfulness:
Intra-Draft Faithfulness: Tests how consistently models follow their own reasoning steps when… See the full description on the dataset page: https://huggingface.co/datasets/polaris-73/faithful-thinking-draft.POLARIS
POLARIS
POLARIS is a prompt-only dataset release for long-form story generation. It contains the training prompts used for the POLARIS story-writing models together with the official test prompts used in evaluation.
This release is intentionally narrow: it is designed to support reproducibility for prompt-based evaluation and generation experiments without releasing copyrighted story text or training-time reasoning traces.
What is included
The dataset has two… See the full description on the dataset page: https://huggingface.co/datasets/rishanthrajendhran/POLARIS.polaris-lter-v1
LTER v1
LTER is one of six datasets in Polaris: Learning to Generate Table Descriptions from Retrieval
Feedback, alongside aw, arctic, ecir, wikitables, and
wtr.
It holds 2,015 tables from the top-downloaded collections of the Long-Term Ecological Research sites
in the Environmental Data Initiative (EDI) — bird surveys, forest phenology, reef colonisation,
cattle records — and 15 keyword queries over them. For each query–table pair, a person decided
whether that table answers… See the full description on the dataset page: https://huggingface.co/datasets/anhaidgroup/polaris-lter-v1.Polaris-clean-v2
