datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-lma-corpus-phase1PCMind-2.1-Kaiyuan-2B-phase1-part1-2
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-2.PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323.aimip-phase1-submissions
AIMIP Phase 1 model submissions
Summary
This dataset contains output from six model submission groups' ~46-year atmospheric simulations produced as contributions to the AI Model Intercomparison Project (AIMIP) Phase 1. AIMIP systematically evaluates AI weather/climate models trained on ERA5 reanalysis by running standardized AMIP-style simulations and comparing their climate statistics against the ERA5 reanalysis and a conventional climate model.
There are… See the full description on the dataset page: https://huggingface.co/datasets/allenai/aimip-phase1-submissions.pretrain_phase1_ja_shufflepretrain_phase1_jakupe-thinkspark-270m-phase1-data
ThinkSpark-v2-350M — Phase-1 free-audio training data
Pre-encoded Mimi cb0 (12.5Hz semantic) tokens +
per-frame energy/f0 prosody, packaged for Phase 1 of ThinkSpark-v2-350M — teaching a
270M-parameter Gemma-3-based full-duplex floor-controller that a stream of Mimi audio
tokens carries language + prosody, before Phase 2 teaches it to referee turn-taking.
Sourced from free/open corpora (LibriSpeech, AI4Bharat Kathbath/Shrutilipi, IndicTTS,
Google FLEURS) via… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-thinkspark-270m-phase1-data.cl3410-phase1
CL3410 Phase 1 — Malayalam and Assamese language-model corpora
Two independently built pretraining corpora with their own tokenizers:
Malayalam as the higher-resource language and Assamese as the
lower-resource one. Nothing is shared between them — separate sources,
separate cleaning thresholds, separate vocabularies, separate models.
Only the language-agnostic pipeline code is common, parameterised per
language.
Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.phase1_pick_place_A1_10fps
Phase1 Pick Place A1 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Pick up the red block and place it on the blue dish."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 100 (28,459 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac Sim 5.1 /… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/phase1_pick_place_A1_10fps.SOLOMON-Phase1-Stocks-Onlypan_safety_phase1_details_teacher_forcingopenai-prm800k-phase1_train-stepwise-bestphase1_pick_place_A1_10fps_via4cm
Phase1 Pick Place A1 10Fps Via4Cm
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Pick up the red block and place it on the blue dish."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 100 (28,530 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac Sim 5.1… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/phase1_pick_place_A1_10fps_via4cm.phase1_pick_place_A2_10fps_via4cm
Phase1 Pick Place A2 10Fps Via4Cm
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Pick up the red block and place it on the blue dish."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 100 (28,755 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac Sim 5.1… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/phase1_pick_place_A2_10fps_via4cm.sdxl-qwen-phase1-cache
SDXL + Qwen3.5 Phase-1 Conditioning Cache
Precomputed, frozen-encoder conditioning for an SDXL + Qwen rectified-flow finetune: SDXL VAE latents, SDXL CLIP-L/CLIP-G text features, Qwen3.5 pooled and full-sequence text features, and geolip aleph addresses — one row per source image/caption pair, sharded so the build survives ephemeral compute and the result is reusable across runs and projects.
It exists so the expensive encode pass is paid once. Every array a downstream trainer… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase1-cache.openai-prm800k-phase1_train-stepwise-critiquepull_cube_ours_phase1_80_10fps
Pull Cube Ours Phase1 80 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Pull the cube to the target marker."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 80 (25,568 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac Sim 5.1 / IsaacLab 2.3.2… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/pull_cube_ours_phase1_80_10fps.extract_cube_ours_phase1_80_10fps
Extract Cube Ours Phase1 80 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Extract the cube from the pocket and place it on the target marker."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 80 (25,138 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/extract_cube_ours_phase1_80_10fps.pick_place_ours_phase1_80_10fps
Pick Place Ours Phase1 80 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Pick up the red block and place it on the blue dish."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 80 (22,998 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac Sim 5.1 /… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/pick_place_ours_phase1_80_10fps.browser-agent-phase1-sft-action-only
Browser Agent Phase 1 SFT Action-Only
What this is
Action-only step-level chat SFT data for browser-agent training.
Each example teaches the model to predict the next BrowserGym action from:
the original generation-time system prompt used for data collection
task goal and URL
short recent history
current observation text and diagnostics
Assistant targets contain only the next action.
Why this format
This is the primary training format for small-model SFT… See the full description on the dataset page: https://huggingface.co/datasets/saital/browser-agent-phase1-sft-action-only.phase1-dataset
Phase 1 Dataset
Data Card
Field
Type
Description
data_id
string
Unique identifier (hash of graph_id + target + scm_type)
graph_id
string
Source graph identifier
X
List[List[float]]
Feature matrix (n_samples × n_features)
y
List[float]
Target variable values (n_samples,)
adjacency_matrix
List[List[int]]
NxN binary matrix where N = (n_feature + 1), A[i,j]=1 means i→j (target as last node)
feature_mask
List[int]
Binary mask for Markov blanket features… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/phase1-dataset.UR7e_phase1_Sort_RGBblock_10fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur7e",
"total_episodes": 20,
"total_frames": 11441,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/UR7e_phase1_Sort_RGBblock_10fps.openai-prm800k-phase1_test-stepwise-bestnext-jev-phase1-accepted
NextJev Phase 1 accepted
This dataset contains all 137,960 independently verified Phase 1 accepted records.
train: 124,341 records compatible with the current three-way NextJev trainer.
validation: 2,377 held-out records, split by evidence group with seed 42.
excluded: 11,242 accepted records retained for audit but excluded from current NextJev training because their labels cannot be converted losslessly.
Images are embedded in the Parquet image column using the Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/next-jev-phase1-accepted.rq2_extract_cube_phase1_20_10fps
Rq2 Extract Cube Phase1 20 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Extract the cube from the pocket and place it on the target marker."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 20 (6,281 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/rq2_extract_cube_phase1_20_10fps.data_ablation_phase1_loss
Tập dữ liệu được chọn từ split default của OpenR1-Math-220k, 10.000 mẫu hợp lệ.
Loại bỏ toàn bộ các bài toán có uuid trùng với tập s1K-1.1.
Với mỗi bài toán, chỉ giữ các generation có is_reasoning_complete = True, generation được chọn phải có correctness_math_verify = True; nếu có kết quả đánh giá từ correctness_llama, giá trị này cũng phải là True.
Mỗi bài toán chỉ giữ một reasoning trajectory: generation hợp lệ đầu tiên theo thứ tự trong trường generations.
Dataset cuối cùng gồm các trường:… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/data_ablation_phase1_loss.openai-prm800k-phase1_test-solutions-onlyepistemic-humility-phase1-evals
Epistemic Humility Phase 1 Evaluation Artifacts
This dataset repository publishes the release-safe Phase 1 evaluation analysis
artifacts for the Epistemic Humility research program.
Source repository:
https://github.com/ProfSynapse/Epistemic-Humility-Research
Local provenance:
Analysis root: experiment/phase1/eval/analysis/
Results provenance inventory: experiment/paper/results-provenance-inventory.md
Public artifact manifest: docs/public-artifacts.md
Contents… See the full description on the dataset page: https://huggingface.co/datasets/professorsynapse/epistemic-humility-phase1-evals.msmarco_passage_gpro_phase1_embedderIsaacLab-SO101-Phase1-sort_by_color-80episode-10fps
Isaaclab So101 Phase1 Sort By Color 80Episode 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Sort the blocks onto the matching colored dishes."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 80 (67,786 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on… See the full description on the dataset page: https://huggingface.co/datasets/Cache-SCA/IsaacLab-SO101-Phase1-sort_by_color-80episode-10fps.
