datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
beetle-phase-transition
Induction-head phase transition in BeetleLM
Per-checkpoint mechanistic metrics for Beetle language models, tracking how the
induction circuit forms during training.
The files here have six different schemas, so they are exposed as separate
configs. Loading the directory as a single table fails with a cast error —
that is why the configs above exist, not a bug.
from datasets import load_dataset
traj = load_dataset("suchirsalhan/beetle-phase-transition", "trajectories")
abl =… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/beetle-phase-transition.slue-phase-2
Licensing Information
SLUE-HVB
SLUE-HVB dataset contains a subset of the Gridspace-Stanford Harper Valley speech dataset and the copyright of this subset remains the same with the original license, CC-BY-4.0. See also original license notice (https://github.com/cricketclub/gridspace-stanford-harper-valley/blob/master/LICENSE)
Additionally, we provide dialog act classification annotation and it is covered with the same license as CC-BY-4.0.
SLUE-SQA-5… See the full description on the dataset page: https://huggingface.co/datasets/asapp/slue-phase-2.indic-lma-corpus-phase1ndlj_tosho_1
国会図書館に収蔵される著作権切れのデータです
sdxl-qwen-phase0
SDXL–Qwen Phase-0 dataset
Purpose-built training set for AbstractPhil/geolip-sdxl-aleph.
Each row pairs a Qwen-Image-Lightning render with the caption that produced it and an
encoder-invariant geometric "aleph" address derived from the caption's bytes. It exists to
retrain SDXL (which stays the base model) around a new text encoder (Qwen in place of
CLIP-G) under a rectified-flow objective: the render is the flow-matching target, and the
student learns to reproduce it from the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/sdxl-qwen-phase0.PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-1-0323.PCMind-2.1-Kaiyuan-2B-phase1-part1-2
This repository contains the complete pretraining dataset for
PCMind-v2.1-Kaiyuan-2B, a leading fully open-source language model.
Overview
The dataset is organized into 5 training phases, with all phase datasets open-sourced in this repository. Our training methodology employs domain-specific mixing strategies across five primary domains:
English: General English text
Chinese: General Chinese text
Code: Programming and code-related content
Math: Mathematical reasoning and… See the full description on the dataset page: https://huggingface.co/datasets/Lxd99/PCMind-2.1-Kaiyuan-2B-phase1-part1-2.aimip-phase1-submissions
AIMIP Phase 1 model submissions
Summary
This dataset contains output from six model submission groups' ~46-year atmospheric simulations produced as contributions to the AI Model Intercomparison Project (AIMIP) Phase 1. AIMIP systematically evaluates AI weather/climate models trained on ERA5 reanalysis by running standardized AMIP-style simulations and comparing their climate statistics against the ERA5 reanalysis and a conventional climate model.
There are… See the full description on the dataset page: https://huggingface.co/datasets/allenai/aimip-phase1-submissions.FineVideo-Phase7-Flattened
FineVideo-Phase7-Flattened
Recaption + grounding augment (v8) release of FineVideo-VLA (window=8)
training text -- 371,892 rows, exact row-count match with the prior v6/v7
release (no videos/activities lost). Pose/cosmos/seed2/snac token payloads
are functionally unchanged; what changed is the caption quality and the
USER instruction text.
What changed and why
Captions replaced: the old caption prompt ("Describe what the person is doing in one short sentence."… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/FineVideo-Phase7-Flattened.pretrain_phase1_ja_shufflepretrain_phase4_curatedclaude-merged-tracesphase3-dataset
Dataset
Generated dataset with 120 configurations.
Configuration
dataset_name: phase3-dataset
graphs_path: hf://CSE472-blanket-challenge/phase3-graphs
output_path: data/datasets/${dataset_name}
n_samples: 1000
n_datasets: 1
scm_type: linear, nonlinear
coeff_range: 1.0
noise_std: 0.5
env_type: iid, covariate, label
projection: pca
shift_mean: 0.8
shift_std: 0.2
train_fraction: 0.8
seed: 42
overwrite: true
Load data
from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/CSE472-blanket-challenge/phase3-dataset.phase2pretrain_phase1_japretrain_phase2_enkupe-thinkspark-270m-phase1-data
ThinkSpark-v2-350M — Phase-1 free-audio training data
Pre-encoded Mimi cb0 (12.5Hz semantic) tokens +
per-frame energy/f0 prosody, packaged for Phase 1 of ThinkSpark-v2-350M — teaching a
270M-parameter Gemma-3-based full-duplex floor-controller that a stream of Mimi audio
tokens carries language + prosody, before Phase 2 teaches it to referee turn-taking.
Sourced from free/open corpora (LibriSpeech, AI4Bharat Kathbath/Shrutilipi, IndicTTS,
Google FLEURS) via… See the full description on the dataset page: https://huggingface.co/datasets/anuj-inavlabs/kupe-thinkspark-270m-phase1-data.pretrain_phase2_en_shufflephase_tree_results
PHASE-Tree Evaluation Results
Full evaluation outputs for the PHASE-Tree paper
(Psychology-grounded Hierarchical Attribute-Structured Evolving Tree),
covering 8 character-dialogue datasets, 4 experimental paradigms, and
2 evaluation splits (random test + OOD test).
Please cite this work if you use these results for analysis, comparison, reproduction, or any other research purpose.
🔗 Resources:
📄 Paper: arXiv:2608.06975
📦 GitHub Repository: MemTensor/PHASE-Tree (code… See the full description on the dataset page: https://huggingface.co/datasets/Mathematics-Yang/phase_tree_results.43-143-phase2-appconv-ime-large-v3-preparedaiws5.3_phase3phase2onevision-phase2-data
OneVision-Encoder Phase 2 — caches (data)
Pre-computed feature caches required to train Phase 2 (video generation) experiments on top of the a8 codec.
Companion repo: code + frozen a8 ckpt live at NNNNNr/onevision-phase2-code (private).
Contents
caches/
fused_tokens_25k/ # 48 GB · 25,188 × .pt (a8 mu tokens)
ff_feats_dinov2/ # 13 GB · 25,387 × .pt (DINOv2-L first-frame feats)… See the full description on the dataset page: https://huggingface.co/datasets/NNNNNr/onevision-phase2-data.phase_miphase1_pick_place_A1_10fps
Phase1 Pick Place A1 10Fps
LeRobot v3.0 dataset collected via
SCRAPE-IsaacLab — a
Code-as-Policies replay pipeline running inside Isaac Sim 5.1 / IsaacLab 2.3.2.
Task instruction: "Pick up the red block and place it on the blue dish."
Robot: so101_follower
Cameras: top + left-wrist RGB @ 10 fps
Episodes: 100 (28,459 frames total)
Labels: per-frame natural-language skill labels in skill.natural_language
and subtask.* columns (labeled by Gemini)
Generated on Isaac Sim 5.1 /… See the full description on the dataset page: https://huggingface.co/datasets/HyeonseokE/phase1_pick_place_A1_10fps.phased-self-discover-mistral-structured-5-shot-bbh-evalcl3410-phase1
CL3410 Phase 1 — Malayalam and Assamese language-model corpora
Two independently built pretraining corpora with their own tokenizers:
Malayalam as the higher-resource language and Assamese as the
lower-resource one. Nothing is shared between them — separate sources,
separate cleaning thresholds, separate vocabularies, separate models.
Only the language-agnostic pipeline code is common, parameterised per
language.
Everything here was collected and cleaned for this project. No… See the full description on the dataset page: https://huggingface.co/datasets/amritha27/cl3410-phase1.FineVideo-Phase2-3DPose
FineVideo-Phase2-3DPose — 3D Human Pose from MotionBERT
Overview
This dataset contains 3D human pose data lifted from 2D detections using MotionBERT, extracted from ~40K YouTube videos in the FineVideo dataset.
This is the output of Phase 2 (+ Phase 2.5 resampling) in the FineVideo-VLA pipeline. It contains raw 3D joint positions as NumPy arrays at 30fps, before any filtering, normalisation, or tokenisation.
Statistics
Metric
Value
Source… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/FineVideo-Phase2-3DPose.SOLOMON-Phase1-Stocks-Onlyphased-self-discover-mistral-unstructured-5-shot-bbh-eval
