datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lilm1-230m-posttraining
LiLM1-230M post-training data
This dataset contains the selected post-training data for LiLM1-230M.
Method
The records combine general assistant text with structured tool-use examples.
The configurations preserve the binding stage, the ratio study, and the
selected 4:8 continuation.
Configurations
Configuration
Content
binding-repair
Tool binding data
ratio-study
Three training splits used for ratio selection
ratio-evaluation
Shared… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-230m-posttraining.post-training-takehome-math500-bon16
MATH-500 Best-of-16 Post-Training Take-Home Results
A 50-problem study of test-time compute, based on the Hugging Face post-training take-home challenge. Nothing here trains or modifies a model: both the generator and the reward model stay frozen, and the only variable is how a final answer is chosen from 16 sampled candidates.
Construction
Filtered MATH-500 to levels 1-3, shuffled with seed 1, and selected 50 rows.
Generated one greedy solution per problem with… See the full description on the dataset page: https://huggingface.co/datasets/augustoFranke/post-training-takehome-math500-bon16.
