datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dclm-baseline-1.0-parquet
DCLM-baseline
Note: this is an identical copy of https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0, where all the files have been mapped to a parquet format.
DCLM-baseline is a 4T token / 3B document pretraining dataset that achieves strong performance on language model benchmarks.
Below are comparisions of model trained on DCLM-baseline with other models in the 7B regime.
Model
Params
Tokens
Open dataset?
CORE
MMLU
EXTENDED
Open weights, closed datasets… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0-parquet.Qwen3.5-4B-Base
juiceb0xc0de/Qwen3.5-4B-Base
A brain atlas for Qwen/Qwen3.5-4B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-4B-Base.msmarco-msmarco-distilbert-base-v3
MS MARCO with hard negatives from msmarco-distilbert-base-v3
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:
msmarco-bm25… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-v3.TTCW-Based-Review
TTCW Creative Writing Evaluation Dataset
If you use this dataset in your research, please cite our paper — it helps support ongoing academic work. Citation details are at the bottom of this page.
Dataset Description
Summary
A supervised fine-tuning (SFT) dataset for training LLMs to act as creative writing evaluators. Each example contains a creative story and four message-format columns representing different evaluation objectives — from… See the full description on the dataset page: https://huggingface.co/datasets/VibrantVista/TTCW-Based-Review.baseball
Baseball data
MLB datasets published as Parquet, one subset per table (select it in the
Data Studio dropdown). Maintained by the etl hf GitHub Actions jobs; each run
merges newly-fetched rows into the existing file (dedup on each table's primary
key). The statcast subset combines the season-partitioned statcast_<year>
files.
msmarco-msmarco-distilbert-base-tas-b
MS MARCO with hard negatives from msmarco-distilbert-base-tas-b
MS MARCO is a large scale information retrieval corpus that was created based on real user search queries using the Bing search engine.
For each query and gold positive passage, the 50 most similar paragraphs were mined using 13 different models. The resulting data can be used to train Sentence Transformer models.
Related Datasets
These are the datasets generated using the 13 different models:… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/msmarco-msmarco-distilbert-base-tas-b.finance-baselatenet-v0-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Full-sequence activations (80 layers, 8192 dim, float16, all tokens) from meta-llama/Llama-3.1-70B (base) on 23724 LateNet v0 statements (affirmative + negated). Extracted via NDIF. Raw statements only (no chat template). Prompts ordered by negated→generator→pair_id for contiguous domain shards.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-70b-base.latenet-v0-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision b906e4dc842aa489c962f9db26554dcfdde901fe).
LateNet v0 activations for Llama 3.1 405B base (all layers, full sequence)
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
20
-
Prompts: 23724
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/latenet-v0-activations-llama3.1-405b-base.hle-context-baseline-deepgot-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown).
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
12
-
Prompts: 7660
Format version: 1.1
Load with lmprobe
from lmprobe import pull_dataset, load_activation_dataset
# Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.adapter-based-multimodal-fusion
Falcon-Audio Training Dataset
Training-ready Parquet shards for Falcon-Audio. Rows contain Gemma-tokenized inputs/labels and fp16 Whisper encoder features encoded as raw bytes.
Qwen3.5-9B-Base
juiceb0xc0de/Qwen3.5-9B-Base
A brain atlas for Qwen/Qwen3.5-9B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
This is a base model, before any instruction tuning. That makes it a useful thing to have a map of: whatever… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-9B-Base.jetson1-062626-grab-and-place-salome-baseline-v1-trimThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos",
"tilt.pos"
]… See the full description on the dataset page: https://huggingface.co/datasets/VibeCuisine/jetson1-062626-grab-and-place-salome-baseline-v1-trim.maxrl_qwen3_4B_base_polaris_rollouts
MaxRL Qwen3-4B-Base training rollouts (POLARIS math prompts)
Every training rollout from an online RL run, with exact token ids, sampling
log-probs, and raw rewards — usable as a replay buffer to study off-policy RL
for LLM reasoning completely offline.
The run: Qwen3-4B-Base trained with the maxRL advantage estimator
(A = (r - mean)/(mean + eps), group mean over 16 rollouts per prompt;
maxRL paper) and a pure REINFORCE loss
(L = -A * log pi; no importance ratio, no clipping, no… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/maxrl_qwen3_4B_base_polaris_rollouts.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13Qwen3.5-2B-Base
juiceb0xc0de/Qwen3.5-2B-Base
A brain atlas for Qwen/Qwen3.5-2B-Base, a 24-layer hybrid that runs linear attention on 18 layers and full attention on the other 6. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing.
This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-2B-Base.boris-open-base-cabinet-sim-v24This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 6310,
"total_frames": 454320,
"total_tasks": 1,
"total_videos": 12620,
"total_chunks": 7,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:6310"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/crislmfroes/boris-open-base-cabinet-sim-v24.real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rolloutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-r2-baseline-uniform-c100000-heval-s2026070802-policy-rollouts.bokeh-eval-lfrepro-baseA15k-fullego-binocular-v1
BaseMatrix EGO Binocular v1
Egocentric bimanual manipulation dataset with 3D hand tracking, EMG muscle signals, and robot-ready action representations. Captured from a first-person perspective using head-mounted stereo cameras and forearm EMG wristbands, processed through a 7-stage automated pipeline.
Key differentiators:
Egocentric + binocular stereo — first-person view matching humanoid robot camera placement
21-joint 3D hand skeleton per hand (MANO topology) — retargetable… See the full description on the dataset page: https://huggingface.co/datasets/basematrix/ego-binocular-v1.base4-clean-table-01-BC-FVThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "u850",
"total_episodes": 50,
"total_frames": 131695,
"total_tasks": 1,
"total_videos": 150,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 50,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/AmmarWaheed/base4-clean-table-01-BC-FV.qwen3-8b-base-atlas-SAE
Qwen3-8B-Base Feature Atlas
A single queryable SQLite database (atlas.sqlite, ~570 MB) that maps the internals of
Qwen/Qwen3-8B-Base — every weight channel and
every sparse-autoencoder feature scored for what it selects for, across a register-diverse
corpus of 4,946 prompts.
It is not a text dataset. There are no training rows. It is an index of model internals
— the kind of thing you query to find "which channels in layer 23 discriminate compliance from
authentic-personality… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/qwen3-8b-base-atlas-SAE.OpenVLA-on-libero-base
OpenVLA on LIBERO-base
OpenVLA rollouts on the standard LIBERO base suites: libero_spatial, libero_object, libero_goal, and libero_10.
The target collection is 40 tasks with 500 episodes per task, for 20,000 episodes total.
Episodes are uploaded incrementally while collection is running. See metadata/upload_state.json for upload progress.
dclm-baseline-1.0_subset_30Mgot-activations-llama3.1-70b-base
meta-llama/Llama-3.1-70B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-70B (revision 349b2ddb53ce8f2849a6c168a81980ab25258dac).
Geometry of Truth curated dataset activations for Llama 3.1 70B base
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-79
8192
-
4
-
Prompts: 7660
Format version: 2.0
Load with lmprobe
from lmprobe import load_activations, Probe
acts =… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/got-activations-llama3.1-70b-base.real01b-square-d2-r3-baseline-nocf-c100k-heval-s2026063004-policy-rolloutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-square-d2-r3-baseline-nocf-c100k-heval-s2026063004-policy-rollouts.real01b-marker-d2-r2-baseline-nocf-c100k-heval-s2026062701-policy-rolloutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-marker-d2-r2-baseline-nocf-c100k-heval-s2026062701-policy-rollouts.real01b-md2-r5-repeat-base-dp-filmtiidk4-c200k-n32-s2026070704This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-md2-r5-repeat-base-dp-filmtiidk4-c200k-n32-s2026070704.real01b-routing-d1-r4-baseline-uniform-c100000-heval-s2026071007-policy-rolloutsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
7
],
"names": [
"cart_pos_x",
"cart_pos_y",
"cart_pos_z",
"cart_rot_x",
"cart_rot_y",
"cart_rot_z"… See the full description on the dataset page: https://huggingface.co/datasets/ankile/real01b-routing-d1-r4-baseline-uniform-c100000-heval-s2026071007-policy-rollouts.
