datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.corpus-1T-manifest
SPP Corpus 1T Manifest
The selection manifest for the ~1.0T-token pretraining corpus used in
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
The corpus is a seeded subsample of allenai/dolma3_mix-6T.
Rather than redistribute ~2.6 TB of text that is already public, this dataset
publishes the selection decisions keyed by upstream document id, so the corpus
can be reconstructed exactly by replaying against upstream.
📄 Reflections + text for the annotated half:… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-1T-manifest.reflection-50m
SPP Reflection 50M
The 51.4M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero — the production half-corpus run, and the dataset the
released models were actually trained on.
🔬 Small sample (same format): dlab-spp/reflection-sample-2k
📉 Earlier 10M run: dlab-spp/reflection-10m
🧾 Safety scores for the full 1T corpus: dlab-spp/safety-classifications
Each row pairs a source document with two generated constitution reflections — a… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-50m.safety-classifications
Safety Annotations for dolma3_mix
Safety score annotations for a 20K-shard subset of allenai/dolma3_mix-6T using
locuslab/safety-classifier_gte-large-en-v1.5.
Schema
Column
Type
Description
id
string
Row identifier (matches source dataset)
safety_score
int8
Argmax safety class (0-5)
safety_probs
list[float32]
Full 6-class probability distribution
Safety scale
Score
Label
Count
Percentage
0
safe
302,972,734
77.39%
1… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/safety-classifications.reflection-10m
SPP Reflection 10M
The full ~10M-document reflection set from Synthetic Persona Pretraining (SPP):
Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
🔬 Small sample (same format): dlab-spp/reflection-sample-2k — a 2,000-row sample drawn from this set, for quick inspection.
Each row pairs a pretraining document with a synthetic, value-laden reflection
generated for it: a short first-person (and third-person) moral reflection… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-10m.SPPdetails_v000000__SwallowMaid-8B-L3-SPPO-abliterated
Dataset Card for Evaluation run of v000000/SwallowMaid-8B-L3-SPPO-abliterated
Dataset automatically created during the evaluation run of model v000000/SwallowMaid-8B-L3-SPPO-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_v000000__SwallowMaid-8B-L3-SPPO-abliterated.constitution-eval
ConstitutionEval
A blind, behavioral multiple-choice benchmark that measures whether a language model's
behavior aligns with a value constitution, without the constitution in context. Each item is a
realistic scenario ending at a decision point with four candidate courses of action; exactly one
is fully constitution-consistent and the other three each enact a specific, attractively-packaged
violation. A model scores well only if its internalised values match the constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/constitution-eval.sp-sft-normal-300k
model-raising-pbsft-instruct-300k
A constitution-aware paired SFT dataset of 300,000 general-purpose (WildChat) instruct
prompts. Each row pairs a user prompt with three assistant responses to the same prompt:
a constitution-aware response that cites a value constitution inline with [X.Y] markers,
a constitution-invisible rendering of that same response (no markers, no constitution vocabulary), and
the original response that shipped with the prompt in WildChat-1M.
It is part… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/sp-sft-normal-300k.sp-sft-safety-180k
model-raising-pbsft-safety-180k
A constitution-aware paired SFT dataset of 182,688 safety-relevant prompts. Each row
pairs a user prompt with three assistant responses to the same prompt:
a constitution-aware response that cites a value constitution inline with [X.Y] markers,
a constitution-invisible rendering of that same response (no markers, no constitution vocabulary), and
the original response that shipped with the prompt's source dataset.
It is part of the Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/sp-sft-safety-180k.SPP_30K_reasoning_tasks
Dataset Card for "SPP_30K_verified_tasks"
Dataset Summary
This is an augmented version of the Synthetic Python Problems(SPP) Dataset.
This dataset has been generated from the subset of the data has been de-duplicated and verified using a Python interpreter. (SPP_30k_verified.jsonl).
The original dataset contains small Python functions that include a docstring with a small description of what the function does and some calling examples
for the function.
The current… See the full description on the dataset page: https://huggingface.co/datasets/Nan-Do/SPP_30K_reasoning_tasks.SPP_30K_reasoning_tasks
Dataset Card for "SPP_30K_verified_tasks"
Dataset Summary
This is an augmented version of the Synthetic Python Problems(SPP) Dataset.
This dataset has been generated from the subset of the data has been de-duplicated and verified using a Python interpreter. (SPP_30k_verified.jsonl).
The original dataset contains small Python functions that include a docstring with a small description of what the function does and some calling examples
for the function.
The current… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/SPP_30K_reasoning_tasks.data-mistral-7b-instruct-sppo-iter1data-mistral-7b-instruct-sppo-iter2
Dataset Card for "data-mistral-7b-instruct-sppo-iter2"
More Information needed
data-mistral-7b-instruct-sppo-iter3
Dataset Card for "data-mistral-7b-instruct-sppo-iter3"
More Information needed
ucla-agi-data-mistral-7b-instruct-sppo-iter1UCLA-AGI__Mistral7B-PairRM-SPPOucla-agi-data-mistral-7b-instruct-sppo-iter1_perplexitiesllama3base_bsln_sppo_t1_10kdetails_UCLA-AGI__Gemma-2-9B-It-SPPO-Iter3
Dataset Card for Evaluation run of UCLA-AGI/Gemma-2-9B-It-SPPO-Iter3
Dataset automatically created during the evaluation run of model UCLA-AGI/Gemma-2-9B-It-SPPO-Iter3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_UCLA-AGI__Gemma-2-9B-It-SPPO-Iter3.reflection-sample-2k
SPP Reflection 2k Sample
A 2,000-row sample (seed 42) of dlab-spp/reflection-10m,
in the identical format, for quick inspection of the data from
Synthetic Persona Pretraining (SPP): Alignment from Token Zero.
📝 Read the post: Synthetic Persona Pretraining: Alignment from Token Zero
📦 Full dataset: dlab-spp/reflection-10m (~10M documents).
Each row pairs a pretraining document with a synthetic, value-laden reflection
(first- and third-person) grounded in a value constitution.… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/reflection-sample-2k.mistralbase_bsln_sppo_t3_10kyifAI__Llama-3-8B-Instruct-SPPO-score-Iter3_gp_8b-table-0.002llama3base_bsln_sppo_t0_10kmistralbase_bsln_sppo_t2_10kdetails_grimjim__Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge.details_UCLA-AGI__Mistral7B-PairRM-SPPO
Dataset Card for Evaluation run of UCLA-AGI/Mistral7B-PairRM-SPPO
Dataset automatically created during the evaluation run of model UCLA-AGI/Mistral7B-PairRM-SPPO.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_UCLA-AGI__Mistral7B-PairRM-SPPO.xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_bt_2b-table-0.001yfzp__Llama-3-8B-Instruct-SPPO-score-Iter1_bt_2b-table-0.001mistralbase_bsln_sppo_t0_10k
