datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fixed-tokenizer-segmentsycb-fixed-meshesThis folder has updated versions of the YCB meshes. All updates are in google_16k folders for each object.
The following updated are available:
nontextured_proc.stl: These are simplified meshes with the normals fixed recommended to be used as collision models. (Note: The normal fixes has to be done manually so not all meshes are verfied, feel free to update them using meshlab, blender, etc).
nontextured_binvox.bt: These file are voxelised representation of the meshes (resolution up to 1mm).… See the full description on the dataset page: https://huggingface.co/datasets/ll4ma-lab/ycb-fixed-meshes.labbench2-fixed
LABBench2 PMID-enriched public mirror
This is a public, schema-compatible mirror of EdisonScientific/labbench2, pinned to upstream revision 27d12d72af24e3f70db8a99df63e567366cbdb80. Original columns and source URLs are unchanged.
Two columns are added to every configuration:
pmids: deduplicated PubMed identifiers resolved for the row's sources.
source_pmids: aligned one-to-one with sources; unresolved or non-PubMed sources are null.
LABBench2
LABBench2 is a… See the full description on the dataset page: https://huggingface.co/datasets/haydn-jones/labbench2-fixed.partnetsim-1024-fixed-viewpointsb1k-224x224-gop8-fixed
BEHAVIOR-1K 2026 — 224x224, GOP=8, upstream-matched encoding
A 224x224 re-encode of the 2026 challenge demos (100 tasks, 3 RGB cameras) whose image
statistics match the dataset the widely-used 50-task checkpoint was pretrained on
(IliaLarchenko/behavior_224_rgb), while keeping GOP=8 for fast random-frame access during
training.
Why "fixed"
Earlier 224 re-encodes of this data used bicubic + libx264 CRF 23, which lands 13% softer
(high-frequency content) than the… See the full description on the dataset page: https://huggingface.co/datasets/JackLiu0406/b1k-224x224-gop8-fixed.ViDoSeek-page-fixedMMLongBench-page-fixedSeamless_Dummy_Dataset_Fixed
MMLU-Pro json
This is a reupload of MMLU-Pro in json format. Please, refer to the original dataset for details.
proof-pile-2-fixed
The original EleutherAI/proof-pile-2 dataset uses a custom python script and .jsonl.zst files, which some versions of the datasets library struggle with.
This dataset contains the same data, subsets, and splits as EleutherAI/proof-pile-2, converted into standard parquet format.
Each subset and split was also shuffled so that you can directly train on the data without issue.
Conversion was performed using the following script:
import os
importzstandard as zstd
import json
import pandas as pd… See the full description on the dataset page: https://huggingface.co/datasets/aklein4/proof-pile-2-fixed.gal_yair_166000_1664x832_fixed
Dataset Card for "gal_yair_large"
More Information needed
spare-fixed-corpus-gpt53-6skill
SPARE fixed corpus — gpt-5.3, 6 skills
Balanced 2,400-game validated corpus (400 per skill) generated by gpt-5.3-chat, the external-generator counterpart to the self-generated 30B corpus. Top-level games are the kept set; _unused/ and _invalid/ retain the full validation record.
2400 validated environments.
Skill
Games
Causal Inference
400
Logical Deduction
400
Mathematical Reasoning
400
Optimization
400
Pattern Recognition
400
Spatial Reasoning
400… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/spare-fixed-corpus-gpt53-6skill.dev-set-71-tasks-fixed-nov-7pushtModel-Collected-fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 206,
"total_frames": 25650,
"total_tasks": 1,
"total_videos": 206,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:206"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beable/pushtModel-Collected-fixed.nopm_claude_writing_fixedThis is Nopm/Opus_WritingStruct, reuploaded and properly converted to ShareGPT format.
ctb-m5-cleaned-fixed-from-2026-01-01-to-2026-07-13IISC_parallel_franka_all_fix_tracks_fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "franka",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"total_videos": 0,
"total_chunks": 0,
"chunks_size": 1000,
"fps": 15,
"splits": {},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/hyzhang01/IISC_parallel_franka_all_fix_tracks_fixed.dvla-can2k-fixed-250hz-events-250fps-delta-trimThis dataset was created using LeRobot.
Dataset Description
dvla-can2k-fixed-250hz-events (250 fps native) - corrected grasp alignment
2,005 episodes / 1,805,397 frames at the native 250 Hz rate. Rolling-can place, one object
(can11) into one of 5 receptacles (3 bowls + plate + tray). 20% static, 80% launched at
0.25-1.5 m/s.
What is "fixed" here. The earlier multi-object builds ran the intercept state machine with its
alignment gate OFF (it skips the wait so it… See the full description on the dataset page: https://huggingface.co/datasets/mickeykang/dvla-can2k-fixed-250hz-events-250fps-delta-trim.wiki-fixedThis repository contains the Wiki-Fixed corpus, presented in the paper Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?.
Code: https://github.com/YiboZhao624/SearchAgentReview
Description
The Wiki-Fixed corpus is based on the Wikipedia 2018 (Wiki-18) corpus and supplemented by the HotpotQA, 2WikiMultiHopQA, and Musique datasets. Compared to the original Wiki-18 corpus, this version contains 295,311 new documents which are critical for answering… See the full description on the dataset page: https://huggingface.co/datasets/ybyby624/wiki-fixed.2026-09-14-colosseum-hospital-self-sacrificial-qwen36-unfiltered-no-synthetic-fixed
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-08-qwen36-0-nosynth (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-08-qwen36-0-nosynth (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
date_generated
2026-09-14
constitution
none
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-colosseum-hospital-self-sacrificial-qwen36-unfiltered-no-synthetic-fixed.Polymarket_data_fixed
Polymarket Data
Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze
A comprehensive dataset of 1.9 billion trading records from Polymarket, processed into multiple analysis-ready formats. Features cleaned data, unified token perspectives, and user-level transformations — ready for market research, behavioral studies, and quantitative analysis.
Zhengjie Wang1,2, Leiyu Chao1,3, Yu Bao1,4, Lian Cheng1,3, Jianhan Liao1,5, Yikang Li1,†
1Shanghai Innovation Institute… See the full description on the dataset page: https://huggingface.co/datasets/wilsonwangwang/Polymarket_data_fixed.2026-09-14-colosseum-hospital-self-sacrificial-qwen36-table2-only-9284-fixed
colosseum_hospital self_sacrificial of LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_hospital self_sacrificial of LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
date_generated
2026-09-14
constitution
none… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-colosseum-hospital-self-sacrificial-qwen36-table2-only-9284-fixed.20_Newsgroups_Fixed
Dataset Card for 20_Newsgroups_Fixed
Dataset Summary
This dataset is a version of the 20 Newsgroups dataset fixed with the help of the Galileo ML Data Intelligence Platform. In a matter of minutes, Galileo enabled us to uncover and fix a multitude of errors within the original dataset. In the end, we present this improved dataset as a new standard for natural language experimentation and benchmarking using the Newsgroups dataset.
Curation Rationale
This… See the full description on the dataset page: https://huggingface.co/datasets/galileo-ai/20_Newsgroups_Fixed.fixed-alpha-scale-pilot-v1
Fixed Alpha Score-Scale Pilot v1
Public artifacts for a controlled 125M-parameter, 4K-context language-model study of fixed attention score scaling with full sparsemax (alpha=2), full entmax-1.5, and a softmax reference.
Layout
metadata/: immutable execution records, plans, logs, source/config snapshots, and SHA-256 manifests.
runs/<run-name>/: complete run backups, including configuration, provenance, metrics, final training state, and preregistered analysis… See the full description on the dataset page: https://huggingface.co/datasets/trytryw/fixed-alpha-scale-pilot-v1.himalaya-fixed-line-trainingfixed-tokenizer-morphscore-segmentsdiverse_risk_acts_fixed2026-09-14-colosseum-hospital-self-sacrificial-qwen36-unfiltered-difficult-agentic-task-fixed
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-08-qwen36-0-dat-7 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
field
value
experiment
colosseum_hospital self_sacrificial of dougalldeepmind/2026-09-08-qwen36-0-dat-7 (mode=think), mixed-checkpoint team; peer=LASR-Callum/2026-08-04-qwen36-lora-table2-only-9284-rank-64
date_generated
2026-09-14
constitution
none
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-colosseum-hospital-self-sacrificial-qwen36-unfiltered-difficult-agentic-task-fixed.vggt_dataset_100_fixed_parquetso101_merged_20260701_fixedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/makermods/so101_merged_20260701_fixed.qwen3-30b-0622-fixed-corpus-actor-spare-games-envs
qwen3-30B-A3B-Instruct-0622-fixed-corpus-actor — generated environments
Environments generated by the SPARE proposer during training run
csjxug1o (qwen3-30B-A3B-Instruct-0622-fixed-corpus-actor), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
1295
Steps covered
230 (step 0–351)
With recovered skill
0
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0622-fixed-corpus-actor-spare-games-envs.
