datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airisk_dilemmas
AIRiskDilemmas risky_behaviors label audit
A full manual re-audit of every risky_behaviors tag in the full split of
kellycyy/AIRiskDilemmas (Chiu et al. 2025, arXiv:2505.14633), triggered by a
suspicion that the Alignment Faking category specifically was mislabeled.
It was — and so, to varying degrees, are the other seven categories.
Why this exists
Every tag in the dataset's risky_behaviors field was produced by a single
one-shot Claude 3.5 Sonnet call per action… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/airisk_dilemmas.op-spp-streams-v2
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v2.corpus-verification
SPP Corpus Verification
Checksums and document-boundary indices for verifying a rebuilt copy of the
Synthetic Persona Pretraining (SPP) training corpus, byte for byte.
The Megatron token streams themselves are 2.17 TB (annotated.bin 421 GB,
compact.bin 1.75 TB) and are fully derived from the published reflections, the
uid manifest, and the tokenizer recipe — so they are not published. These
.idx sidecars carry per-document boundaries and lengths, which is enough to
prove an… See the full description on the dataset page: https://huggingface.co/datasets/dlab-spp/corpus-verification.op-spp-streams-v1
op-spp-streams-v1 — tokenized Megatron streams (SPP-format pretraining corpus)
Tokenized Dolma v1.7 (ODC-BY)
subsample in Megatron IndexedDataset format (uint16, SmolLM2 tokenizer +
<assistant> extension from epfl-dlab/spp-training):
compact dense-packed 2049-token windows; annotated/canary one document
per window, EOD-padded. Built for a Synthetic-Persona-Pretraining-recipe run
(arXiv:2608.13482) — these files carry ONLY document tokens (raw public-corpus
text); the persona… See the full description on the dataset page: https://huggingface.co/datasets/joshycodes/op-spp-streams-v1.spp
Synthetic Python Problems(SPP) Dataset
The dataset includes around 450k synthetic Python programming problems. Each Python problem consists of a task description, 1-3 examples, code solution and 1-3 test cases.
The CodeGeeX-13B model was used to generate this dataset.
A subset of the data has been verified by Python interpreter and de-duplicated. This data is SPP_30k_verified.jsonl.
The dataset is in a .jsonl format (json per line).
Released as part of Self-Learning to Improve Code… See the full description on the dataset page: https://huggingface.co/datasets/wuyetao/spp.UCLA-AGI__Mistral7B-PairRM-SPPO-Iter3-details
Dataset Card for Evaluation run of UCLA-AGI/Mistral7B-PairRM-SPPO-Iter3
Dataset automatically created during the evaluation run of model UCLA-AGI/Mistral7B-PairRM-SPPO-Iter3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Mistral7B-PairRM-SPPO-Iter3-details.UCLA-AGI__Mistral7B-PairRM-SPPO-Iter2-details
Dataset Card for Evaluation run of UCLA-AGI/Mistral7B-PairRM-SPPO-Iter2
Dataset automatically created during the evaluation run of model UCLA-AGI/Mistral7B-PairRM-SPPO-Iter2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Mistral7B-PairRM-SPPO-Iter2-details.UCLA-AGI__Llama-3-Instruct-8B-SPPO-Iter2-details
Dataset Card for Evaluation run of UCLA-AGI/Llama-3-Instruct-8B-SPPO-Iter2
Dataset automatically created during the evaluation run of model UCLA-AGI/Llama-3-Instruct-8B-SPPO-Iter2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Llama-3-Instruct-8B-SPPO-Iter2-details.yfzp__Llama-3-8B-Instruct-SPPO-score-Iter1_bt_8b-table-0.002-details
Dataset Card for Evaluation run of yfzp/Llama-3-8B-Instruct-SPPO-score-Iter1_bt_8b-table-0.002
Dataset automatically created during the evaluation run of model yfzp/Llama-3-8B-Instruct-SPPO-score-Iter1_bt_8b-table-0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yfzp__Llama-3-8B-Instruct-SPPO-score-Iter1_bt_8b-table-0.002-details.chujiezheng__Mistral7B-PairRM-SPPO-ExPO-details
Dataset Card for Evaluation run of chujiezheng/Mistral7B-PairRM-SPPO-ExPO
Dataset automatically created during the evaluation run of model chujiezheng/Mistral7B-PairRM-SPPO-ExPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/chujiezheng__Mistral7B-PairRM-SPPO-ExPO-details.UCLA-AGI__Gemma-2-9B-It-SPPO-Iter1-details
Dataset Card for Evaluation run of UCLA-AGI/Gemma-2-9B-It-SPPO-Iter1
Dataset automatically created during the evaluation run of model UCLA-AGI/Gemma-2-9B-It-SPPO-Iter1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Gemma-2-9B-It-SPPO-Iter1-details.grimjim__Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge-details
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/grimjim__Llama-3-Instruct-8B-SPPO-Iter3-SimPO-merge-details.grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge-details
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge-details.UCLA-AGI__Llama-3-Instruct-8B-SPPO-Iter1-details
Dataset Card for Evaluation run of UCLA-AGI/Llama-3-Instruct-8B-SPPO-Iter1
Dataset automatically created during the evaluation run of model UCLA-AGI/Llama-3-Instruct-8B-SPPO-Iter1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Llama-3-Instruct-8B-SPPO-Iter1-details.UCLA-AGI__Gemma-2-9B-It-SPPO-Iter3-details
Dataset Card for Evaluation run of UCLA-AGI/Gemma-2-9B-It-SPPO-Iter3
Dataset automatically created during the evaluation run of model UCLA-AGI/Gemma-2-9B-It-SPPO-Iter3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Gemma-2-9B-It-SPPO-Iter3-details.UCLA-AGI__Gemma-2-9B-It-SPPO-Iter2-details
Dataset Card for Evaluation run of UCLA-AGI/Gemma-2-9B-It-SPPO-Iter2
Dataset automatically created during the evaluation run of model UCLA-AGI/Gemma-2-9B-It-SPPO-Iter2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Gemma-2-9B-It-SPPO-Iter2-details.xukp20__llama-3-8b-instruct-sppo-iter1-gp-2b-tau01-table-details
Dataset Card for Evaluation run of xukp20/llama-3-8b-instruct-sppo-iter1-gp-2b-tau01-table
Dataset automatically created during the evaluation run of model xukp20/llama-3-8b-instruct-sppo-iter1-gp-2b-tau01-table
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xukp20__llama-3-8b-instruct-sppo-iter1-gp-2b-tau01-table-details.UCLA-AGI__Mistral7B-PairRM-SPPO-details
Dataset Card for Evaluation run of UCLA-AGI/Mistral7B-PairRM-SPPO
Dataset automatically created during the evaluation run of model UCLA-AGI/Mistral7B-PairRM-SPPO
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Mistral7B-PairRM-SPPO-details.xukp20__Llama-3-8B-Instruct-SPPO-score-Iter3_gp_2b-table-0.001-details
Dataset Card for Evaluation run of xukp20/Llama-3-8B-Instruct-SPPO-score-Iter3_gp_2b-table-0.001
Dataset automatically created during the evaluation run of model xukp20/Llama-3-8B-Instruct-SPPO-score-Iter3_gp_2b-table-0.001
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xukp20__Llama-3-8B-Instruct-SPPO-score-Iter3_gp_2b-table-0.001-details.UCLA-AGI__Mistral7B-PairRM-SPPO-Iter1-details
Dataset Card for Evaluation run of UCLA-AGI/Mistral7B-PairRM-SPPO-Iter1
Dataset automatically created during the evaluation run of model UCLA-AGI/Mistral7B-PairRM-SPPO-Iter1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Mistral7B-PairRM-SPPO-Iter1-details.xukp20__Llama-3-8B-Instruct-SPPO-Iter3_gp_2b-table-details
Dataset Card for Evaluation run of xukp20/Llama-3-8B-Instruct-SPPO-Iter3_gp_2b-table
Dataset automatically created during the evaluation run of model xukp20/Llama-3-8B-Instruct-SPPO-Iter3_gp_2b-table
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xukp20__Llama-3-8B-Instruct-SPPO-Iter3_gp_2b-table-details.xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002-details
Dataset Card for Evaluation run of xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002
Dataset automatically created during the evaluation run of model xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_bt_8b-table-0.002-details.cat-searcher__gemma-2-9b-it-sppo-iter-1-evol-1-details
Dataset Card for Evaluation run of cat-searcher/gemma-2-9b-it-sppo-iter-1-evol-1
Dataset automatically created during the evaluation run of model cat-searcher/gemma-2-9b-it-sppo-iter-1-evol-1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cat-searcher__gemma-2-9b-it-sppo-iter-1-evol-1-details.UCLA-AGI__Llama-3-Instruct-8B-SPPO-Iter3-details
Dataset Card for Evaluation run of UCLA-AGI/Llama-3-Instruct-8B-SPPO-Iter3
Dataset automatically created during the evaluation run of model UCLA-AGI/Llama-3-Instruct-8B-SPPO-Iter3
The dataset is composed of 43 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 9 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/UCLA-AGI__Llama-3-Instruct-8B-SPPO-Iter3-details.xukp20__Llama-3-8B-Instruct-SPPO-Iter3_gp_8b-table-details
Dataset Card for Evaluation run of xukp20/Llama-3-8B-Instruct-SPPO-Iter3_gp_8b-table
Dataset automatically created during the evaluation run of model xukp20/Llama-3-8B-Instruct-SPPO-Iter3_gp_8b-table
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xukp20__Llama-3-8B-Instruct-SPPO-Iter3_gp_8b-table-details.yfzp__Llama-3-8B-Instruct-SPPO-Iter1_gp_8b-table-details
Dataset Card for Evaluation run of yfzp/Llama-3-8B-Instruct-SPPO-Iter1_gp_8b-table
Dataset automatically created during the evaluation run of model yfzp/Llama-3-8B-Instruct-SPPO-Iter1_gp_8b-table
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yfzp__Llama-3-8B-Instruct-SPPO-Iter1_gp_8b-table-details.xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_gp_2b-table-0.001-details
Dataset Card for Evaluation run of xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_gp_2b-table-0.001
Dataset automatically created during the evaluation run of model xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_gp_2b-table-0.001
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_gp_2b-table-0.001-details.xkp24__Llama-3-8B-Instruct-SPPO-Iter2_gp_8b-table-details
Dataset Card for Evaluation run of xkp24/Llama-3-8B-Instruct-SPPO-Iter2_gp_8b-table
Dataset automatically created during the evaluation run of model xkp24/Llama-3-8B-Instruct-SPPO-Iter2_gp_8b-table
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xkp24__Llama-3-8B-Instruct-SPPO-Iter2_gp_8b-table-details.xkp24__Llama-3-8B-Instruct-SPPO-Iter2_bt_2b-table-details
Dataset Card for Evaluation run of xkp24/Llama-3-8B-Instruct-SPPO-Iter2_bt_2b-table
Dataset automatically created during the evaluation run of model xkp24/Llama-3-8B-Instruct-SPPO-Iter2_bt_2b-table
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xkp24__Llama-3-8B-Instruct-SPPO-Iter2_bt_2b-table-details.xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_gp_8b-table-0.002-details
Dataset Card for Evaluation run of xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_gp_8b-table-0.002
Dataset automatically created during the evaluation run of model xkp24/Llama-3-8B-Instruct-SPPO-score-Iter2_gp_8b-table-0.002
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/xkp24__Llama-3-8B-Instruct-SPPO-score-Iter2_gp_8b-table-0.002-details.
