datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.reflection-50m
SPP Reflection 50M
The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero — the production half-corpus run, and the dataset the
released models were actually trained on.
🔬 Small sample (same format): dlab-spp/reflection-sample-2k
📉 Earlier 10M run: dlab-spp/reflection-10m
🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications
Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.corpus-1T-manifest
SPP Corpus 1T Manifest
The selection manifest for the ~1.0T-token pretraining corpus used in
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
The corpus is a seeded subsample of allenai/dolma3_mix-6T.
Rather than redistribute ~2.6 TB of text that is already public, this dataset
publishes the selection decisions keyed by upstream document id, so the corpus
can be reconstructed exactly by replaying against upstream.
📄 Reflections + text for the annotated half:… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-1T-manifest.safety-classifications
Safety Annotations for dolma3_mix
Safety score annotations for a 20K-shard subset of allenai/dolma3_mix-6T using
locuslab/safety-classifier_gte-large-en-v1.5.
Schema
Column
Type
Description
id
string
Row identifier (matches source dataset)
safety_score
int8
Argmax safety class (0-5)
safety_probs
list[float32]
Full 6-class probability distribution
Safety scale
Score
Label
Count
Percentage
0
safe
302,972,734
77.39%
1… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/safety-classifications.reflection-10m
SPP Reflection 10M
The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection.
Each row pairs a pretraining document with a synthetic, value-laden reflection
generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.SPPsppairisk_dilemmas
AIRiskDilemmas risky_behaviors label audit
A full manual re-audit of every risky_behaviors tag in the full split of
kellycyy/AIRiskDilemmas (Chiu et al. 2025, arXiv:2505.14633), triggered by a
suspicion that the Alignment Faking category specifically was mislabeled.
It was — and so, to varying degrees, are the other seven categories.
Why this exists
Every tag in the dataset's risky_behaviors field was produced by a single
one-shot Claude 3.5 Sonnet call per action… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/airisk_dilemmas.details_v000000__SwallowMaid-8B-L3-SPPO-abliterated
Dataset Card for Evaluation run of v000000/SwallowMaid-8B-L3-SPPO-abliterated
Dataset automatically created during the evaluation run of model v000000/SwallowMaid-8B-L3-SPPO-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_v000000__SwallowMaid-8B-L3-SPPO-abliterated.op-spp-streams-v2
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v2.constitution-eval
ConstitutionEval
A blind, behavioral multiple-choice benchmark that measures whether a language model's
behavior aligns with a value constitution, without the constitution in context. Each item is a
realistic scenario ending at a decision point with four candidate courses of action; exactly one
is fully constitution-consistent and the other three each enact a specific, attractively-packaged
violation. A model scores well only if its internalised values match the constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/constitution-eval.sp-sft-normal-300k
model-raising-pbsft-instruct-300k
A constitution-aware paired SFT dataset of 300,000 general-purpose (WildChat) instruct
prompts. Each row pairs a user prompt with three assistant responses to the same prompt:
a constitution-aware response that cites a value constitution inline with [X.Y] markers,
a constitution-invisible rendering of that same response (no markers, no constitution vocabulary), and
the original response that shipped with the prompt in WildChat-1M.
It is part… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/sp-sft-normal-300k.SPPIDER-seq-datasets
SPPIDER-seq Datasets
Datasets used for training, validation, blind testing, and benchmarking of the
SPPIDER-seq partner-aware protein–protein interaction site prediction models.
For the software, documentation, and usage instructions, see the
SPPIDER-seq GitHub repository.
For the current pretrained models, see the
SPPIDER-seq model repository.
Please cite the SPPIDER-seq publication when using these datasets:
Porollo A, Jadhav O, Alvarez A, Chen J. SPPIDER-seq: sequence-based… See the full description on the dataset page: https://huggingface.co/datasets/aporollo-lab/SPPIDER-seq-datasets.sp-sft-safety-180k
model-raising-pbsft-safety-180k
A constitution-aware paired SFT dataset of 182,688 safety-relevant prompts. Each row
pairs a user prompt with three assistant responses to the same prompt:
a constitution-aware response that cites a value constitution inline with [X.Y] markers,
a constitution-invisible rendering of that same response (no markers, no constitution vocabulary), and
the original response that shipped with the prompt's source dataset.
It is part of the Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/sp-sft-safety-180k.corpus-verification
SPP Corpus Verification
Checksums and document-boundary indices for verifying a rebuilt copy of the
Synthetic Persona Pretraining (SPP) training corpus, byte for byte.
The Megatron token streams themselves are 2.17 TB (annotated.bin 421 GB,
compact.bin 1.75 TB) and are fully derived from the published reflections, the
uid manifest, and the tokenizer recipe — so they are not published. These
.idx sidecars carry per-document boundaries and lengths, which is enough to
prove an… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-verification.SPP_30K_reasoning_tasks
Dataset Card for "SPP_30K_verified_tasks"
Dataset Summary
This is an augmented version of the Synthetic Python Problems(SPP) Dataset.
This dataset has been generated from the subset of the data has been de-duplicated and verified using a Python interpreter. (SPP_30k_verified.jsonl).
The original dataset contains small Python functions that include a docstring with a small description of what the function does and some calling examples
for the function.
The current… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/SPP_30K_reasoning_tasks.reflection-sample-2k
SPP Reflection 2k Sample
A 2,000-row sample (seed 42) of dlab-spp/reflection-10m,
in the identical format, for quick inspection of the data from
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
📦 Full dataset: dlab-spp/reflection-10m (~10M documents).
Each row pairs a pretraining document with a synthetic, value-laden reflection
(first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.op-spp-streams-v1
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v1.spp
Synthetic Python Problems(SPP) Dataset
The dataset includes around 450k synthetic Python programming problems. Each Python problem consists of a task description, 1-3 examples, code solution and 1-3 test cases.
The CodeGeeX-13B model was used to generate this dataset.
A subset of the data has been verified by Python interpreter and de-duplicated. This data is SPP_30k_verified.jsonl.
The dataset is in a .jsonl format (json per line).
Released as part of Self-Learning to Improve Code… See the full description on the dataset page: https://huggingface.co/datasets/wuyetao/spp.SPP_30K_reasoning_tasks
Dataset Card for "SPP_30K_verified_tasks"
Dataset Summary
This is an augmented version of the Synthetic Python Problems(SPP) Dataset.
This dataset has been generated from the subset of the data has been de-duplicated and verified using a Python interpreter. (SPP_30k_verified.jsonl).
The original dataset contains small Python functions that include a docstring with a small description of what the function does and some calling examples
for the function.
The current… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/SPP_30K_reasoning_tasks.intern-industrial
Industrial Image Audio Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for Industrial work with Image Audio inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/sppereira/intern-industrial.sppu-chroma-dbnews-media-text-tabular-clean
News Media Text Tabular Data Notes
Dataset summary
This repository contains a preparation pipeline and a small metadata sample for News Media work with Text Tabular inputs. It does not claim to be a complete benchmark release; the loader documents how source data is normalized and validated.
Included material
prepare.py — loading, cleaning, and split preparation code.
dataset_infos.json — schema and split metadata.
metadata_sample.jsonl — small… See the full description on the dataset page: https://huggingface.co/datasets/sppereira/news-media-text-tabular-clean.data-mistral-7b-instruct-sppo-iter1UCLA-AGI__Mistral7B-PairRM-SPPOdata-mistral-7b-instruct-sppo-iter2
Dataset Card for "data-mistral-7b-instruct-sppo-iter2"
More Information needed
spp-assistant-axis-resultsdata-mistral-7b-instruct-sppo-iter3
Dataset Card for "data-mistral-7b-instruct-sppo-iter3"
More Information needed
ucla-agi-data-mistral-7b-instruct-sppo-iter1details_UCLA-AGI__Gemma-2-9B-It-SPPO-Iter3
Dataset Card for Evaluation run of UCLA-AGI/Gemma-2-9B-It-SPPO-Iter3
Dataset automatically created during the evaluation run of model UCLA-AGI/Gemma-2-9B-It-SPPO-Iter3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_UCLA-AGI__Gemma-2-9B-It-SPPO-Iter3.
