completions
unsloth-train-completions-testmulti-user-chat-open-llama-7b-v2-open-instruct-completions-onlymulti-user-chat-zephyr-7b-beta-completions-onlyQwen-1.5B-SFT-FinTabNet-ShortPrompt-CompletionsRoblox-completions-BASE-Qwen3-4B-MERGEDmulti-user-chat-openchat-3.5-0106-completions-onlymulti-user-chat-llama-2-7b-chat-completions-onlymulti-user-chat-openchat-3.5-0106-completions-only
alignment_faking_claude_completionsopd-kd-thinky-deepmath-completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
OPD
Model (student)
HuggingFaceH4/KD-Thinky
Model (teacher)
Qwen/Qwen3-8B
Prompt dataset
HuggingFaceH4/DeepMath-103K
Group size
4
Max completion tokens
4096
Temperature
1.0
Learning rate
0.0001
model_revision
v00.08-step-000003125
dataset_configtrl_all
lora_rank
128
opd_kl_coef
1.0… See the full description on the dataset page: https://huggingface.co/datasets/kashif/opd-kd-thinky-deepmath-completions.MultiPL-E-completions
Raw Data from MultiPL-E
This repository is frozen. See https://huggingface.co/datasets/nuprl/MultiPL-E-completions for a more complete version of this repository.
Uploads are a work in progress. If you are interested in a split that is not yet available, please contact a.guha@northeastern.edu.
This repository contains the raw data -- both completions and executions -- from MultiPL-E that was used to generate several experimental results from the
MultiPL-E, SantaCoder, and StarCoder… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/MultiPL-E-completions.CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the… See the full description on the dataset page: https://huggingface.co/datasets/wetsoledrysoul/CerebRM-olmo-3-7b-instruct-sft-list_em-so1_completions.test-grpo-vlm-log-completions
TRL Completion logs
This dataset contains the completions generated during training using trl.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion generated by the model
<reward_function_name>: the reward(s) assigned to the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/test-grpo-vlm-log-completions.judged_science_completions
