datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alignment_faking_claude_completionsopd-kd-thinky-deepmath-completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
OPD
Model (student)
HuggingFaceH4/KD-Thinky
Model (teacher)
Qwen/Qwen3-8B
Prompt dataset
HuggingFaceH4/DeepMath-103K
Group size
4
Max completion tokens
4096
Temperature
1.0
Learning rate
0.0001
model_revision
v00.08-step-000003125
dataset_configtrl_all
lora_rank
128
opd_kl_coef
1.0… See the full description on the dataset page: https://huggingface.co/datasets/kashif/opd-kd-thinky-deepmath-completions.MultiPL-E-completions
Raw Data from MultiPL-E
This repository is frozen. See https://huggingface.co/datasets/nuprl/MultiPL-E-completions for a more complete version of this repository.
Uploads are a work in progress. If you are interested in a split that is not yet available, please contact a.guha@northeastern.edu.
This repository contains the raw data -- both completions and executions -- from MultiPL-E that was used to generate several experimental results from the
MultiPL-E, SantaCoder, and StarCoder… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/MultiPL-E-completions.CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/wetsoledrysoul/CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions.test-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.judged_science_completionsMultiPL-E-completions
Raw Data from MultiPL-E
This repository contains the raw data -- both completions and executions --
from MultiPL-E that was used to generate several experimental results from the
MultiPL-E, SantaCoder, and StarCoder papers.
The original MultiPL-E completions and executions are stored in JOSN files. We use the following script
to turn each experiment directory into a dataset split and upload to this repository.
Every split is named base_dataset.language.model.temperature.variation… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/MultiPL-E-completions.grammar-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-completions.deepmath-completions-logs
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/qgallouedec/qwen2-0.5b-deepmath-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/deepmath-completions-logs.essay-vocab-range-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-range-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-range-qwen3.5-4b-trl-completions.essay-grammar-range-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-grammar-range-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-grammar-range-qwen3.5-4b-trl-completions.rh_qwen3_8b_prompted_v2_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/sunshineNew/rh_qwen3_8b_prompted_v2_completions.swerebench-filtered-openhands-minimax-m2_5-corrected-completionsrh_qwen3_8b_sdf_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/sunshineNew/rh_qwen3_8b_sdf_completions.essay-vocab-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-trl-completions.toxic-completions
ToxicCompletions
This dataset is a collection of toxic and non-toxic user requests along with appropriate and inappropriate, model-generated completions.
Appropriate completion: Complying with a non-toxic request or refusing a toxic request
Inappropriate completion: Complying with a toxic request or refusing a non-toxic request
Fields
prompt: A real user prompt from the ToxicChat dataset
completion: A model-generated response to the prompt
is_toxic: Whether the… See the full description on the dataset page: https://huggingface.co/datasets/dvruette/toxic-completions.train_rl_agent_completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
GRPO
Model (student)
Qwen/Qwen3-4B-Instruct-2507
Prompt dataset
VerifierEnvDataset
Group size
4
Max completion tokens
512
Temperature
1.0
Learning rate
1e-05
Schema
Each parquet file corresponds to one rollout step and contains the following
columns:
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/kashif/train_rl_agent_completions.train_rl_dpo_completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
Online DPO
Model (student)
Qwen/Qwen3-4B-Instruct-2507
Prompt dataset
openai/gsm8k
Group size
8
Max completion tokens
1024
Temperature
1.0
Learning rate
5e-06
dpo_beta
0.1
dpo_loss_type
sigmoid
Schema
Each parquet file corresponds to one rollout step and contains the following… See the full description on the dataset page: https://huggingface.co/datasets/kashif/train_rl_dpo_completions.Dolci-Think-RL-7B-Completions-SFT
Dolci-Think-Completions-SFT
Dataset Summary
Dolci-Think-Completions-SFT is a set of 5,031,398 completions(!!) from the Olmo-3-7B-Think-SFT model over the prompts considered when making Dolci-Think-RL.
These completions were mainly used to filter easy data, but we believe the completions may be useful in general.
It contains 636,095 high-quality prompts covering:
Math
Code
Precise Instruction Following
General Chat
Puzzles
Each split covers one of the above domains, and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-7B-Completions-SFT.grpo-completions-qwen3-0.6b
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo-completions-qwen3-0.6b.grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.llm-filtered-completionscontext-aware-fim-code-completionsDolci-Instruct-RL-Completions
Dolci Instruct RL Completions
Instruction-following completions sampled from OLMo-3-7B-Instruct with teacher logits for knowledge distillation.
Description
This dataset contains instruction-completion pairs with pre-computed teacher logits from OLMo-3-7B-Instruct. Designed for training smaller student models via KL-divergence distillation.
Generation
Completions were sampled from allenai/OLMo-3-7B-Instruct on instruction prompts. For each token position, we… See the full description on the dataset page: https://huggingface.co/datasets/hbfreed/Dolci-Instruct-RL-Completions.Llama-3.2-1B-Instruct-best-of-N-completionsjudged_logic_completionsUniCo-Completions-SFT
UniCo-Completions-SFT
This repository contains 66,603 SFT training examples generated by the UniCo framework, as introduced in our paper, "Towards a Universal Causal Reasoner". It spans 44 (representation form, query type) configurations, and is stored in the ShareGPT format.
For complete dataset construction details and metadata, please refer to another repo in this collection: ChicagoHAI/UniCo.
Response Curation
As described under in Appendix E of the original paper… See the full description on the dataset page: https://huggingface.co/datasets/ChicagoHAI/UniCo-Completions-SFT.Qwen2.5-7B-Instruct-uPRM-T80-adapters-best_of_n-completionsrh_qwen3_8b_sdf_68k_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/sunshineNew/rh_qwen3_8b_sdf_68k_completions.kw-filtered-completions
