datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opd-kd-thinky-deepmath-completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
OPD
Model (student)
HuggingFaceH4/KD-Thinky
Model (teacher)
Qwen/Qwen3-8B
Prompt dataset
HuggingFaceH4/DeepMath-103K
Group size
4
Max completion tokens
4096
Temperature
1.0
Learning rate
0.0001
model_revision
v00.08-step-000003125
dataset_configtrl_all
lora_rank
128
opd_kl_coef
1.0… See the full description on the dataset page: https://huggingface.co/datasets/kashif/opd-kd-thinky-deepmath-completions.MultiPL-E-completions
Raw Data from MultiPL-E
This repository is frozen. See https://huggingface.co/datasets/nuprl/MultiPL-E-completions for a more complete version of this repository.
Uploads are a work in progress. If you are interested in a split that is not yet available, please contact a.guha@northeastern.edu.
This repository contains the raw data -- both completions and executions -- from MultiPL-E that was used to generate several experimental results from the
MultiPL-E, SantaCoder, and StarCoder… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/MultiPL-E-completions.test-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.judged_science_completionsMultiPL-E-completions
Raw Data from MultiPL-E
This repository contains the raw data -- both completions and executions --
from MultiPL-E that was used to generate several experimental results from the
MultiPL-E, SantaCoder, and StarCoder papers.
The original MultiPL-E completions and executions are stored in JOSN files. We use the following script
to turn each experiment directory into a dataset split and upload to this repository.
Every split is named base_dataset.language.model.temperature.variation… See the full description on the dataset page: https://huggingface.co/datasets/nuprl/MultiPL-E-completions.grammar-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-completions.deepmath-completions-logs
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/qgallouedec/qwen2-0.5b-deepmath-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/deepmath-completions-logs.essay-vocab-range-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-range-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-range-qwen3.5-4b-trl-completions.essay-grammar-range-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-grammar-range-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-grammar-range-qwen3.5-4b-trl-completions.rh_qwen3_8b_prompted_v2_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/sunshineNew/rh_qwen3_8b_prompted_v2_completions.rh_qwen3_8b_sdf_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/sunshineNew/rh_qwen3_8b_sdf_completions.essay-vocab-accuracy-qwen3.5-4b-trl-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion:… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/essay-vocab-accuracy-qwen3.5-4b-trl-completions.pypi-clean-derived-v4-completion
PyPI clean — v4 completion partitions
This public repository stores only the previously missing partitions of a
scope-aware Python derivative of vikp/pypi_clean.
It is not independently a complete copy of that derivative. A completion plan
lists the 1,695 existing local partitions and 801 requested completion partitions.
Outputs under v4-completion/<transform fingerprint>/ contain linked files,
units, and edges Parquet tables, per-partition state and checksummed receipts.
Each… See the full description on the dataset page: https://huggingface.co/datasets/IntellAgents/pypi-clean-derived-v4-completion.toxic-completions
ToxicCompletions
This dataset is a collection of toxic and non-toxic user requests along with appropriate and inappropriate, model-generated completions.
Appropriate completion: Complying with a non-toxic request or refusing a toxic request
Inappropriate completion: Complying with a toxic request or refusing a non-toxic request
Fields
prompt: A real user prompt from the ToxicChat dataset
completion: A model-generated response to the prompt
is_toxic: Whether the… See the full description on the dataset page: https://huggingface.co/datasets/dvruette/toxic-completions.train_rl_agent_completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
GRPO
Model (student)
Qwen/Qwen3-4B-Instruct-2507
Prompt dataset
VerifierEnvDataset
Group size
4
Max completion tokens
512
Temperature
1.0
Learning rate
1e-05
Schema
Each parquet file corresponds to one rollout step and contains the following
columns:
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/kashif/train_rl_agent_completions.train_rl_dpo_completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
Online DPO
Model (student)
Qwen/Qwen3-4B-Instruct-2507
Prompt dataset
openai/gsm8k
Group size
8
Max completion tokens
1024
Temperature
1.0
Learning rate
5e-06
dpo_beta
0.1
dpo_loss_type
sigmoid
Schema
Each parquet file corresponds to one rollout step and contains the following… See the full description on the dataset page: https://huggingface.co/datasets/kashif/train_rl_dpo_completions.Dolci-Think-RL-7B-Completions-SFT
Dolci-Think-Completions-SFT
Dataset Summary
Dolci-Think-Completions-SFT is a set of 5,031,398 completions(!!) from the Olmo-3-7B-Think-SFT model over the prompts considered when making Dolci-Think-RL.
These completions were mainly used to filter easy data, but we believe the completions may be useful in general.
It contains 636,095 high-quality prompts covering:
Math
Code
Precise Instruction Following
General Chat
Puzzles
Each split covers one of the above domains, and… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Think-RL-7B-Completions-SFT.grpo-completions-qwen3-0.6b
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/essobi/grpo-completions-qwen3-0.6b.grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/grammar-accuracy-qwen3.5-4b-trl-grpo-vllm-colocate-completions.Llama-3.2-1B-Instruct-best-of-N-completionsjudged_logic_completionsQwen2.5-7B-Instruct-uPRM-T80-adapters-best_of_n-completionsrh_qwen3_8b_sdf_68k_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/sunshineNew/rh_qwen3_8b_sdf_68k_completions.Llama-3.1-8B-Instruct-uPRM-T80-adapters-best_of_n-completionsLlama-3.2-1B-Instruct-beam-search-completionsdeepmath-completions-logs2
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/deepmath-completions-logs2.Llama-3.2-1B-Instruct-uPRM-T80-adapters-dvts-completionspron-acc-qwen3.5-4b-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/bihungba1101/pron-acc-qwen3.5-4b-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/pron-acc-qwen3.5-4b-completions.Qwen2.5-14B-Instruct-uPRM-T80-adapters-dvts-completionstrain_rl_opd_completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
OPD
Model (student)
Qwen/Qwen3-8B-Base
Model (teacher)
Qwen/Qwen3-8B
Prompt dataset
openai/gsm8k
Group size
4
Max completion tokens
512
Temperature
1.0
Learning rate
1e-05
opd_kl_coef
1.0
Schema
Each parquet file corresponds to one rollout step and contains the following
columns:… See the full description on the dataset page: https://huggingface.co/datasets/kashif/train_rl_opd_completions.
