datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.post-training-benchmarks-viewerlilm1-230m-posttraining
LiLM1-230M post-training data
This dataset contains the selected post-training data for LiLM1-230M.
Method
The records combine general assistant text with structured tool-use examples.
The configurations preserve the binding stage, the ratio study, and the
selected 4:8 continuation.
Configurations
Configuration
Content
binding-repair
Tool binding data
ratio-study
Three training splits used for ratio selection
ratio-evaluation
Shared… See the full description on the dataset page: https://huggingface.co/datasets/glouriousgautam/lilm1-230m-posttraining.dqs-post-training
DQS Post-Training Preference Data
Strict English-to-Korean preference data for three post-training objectives.
All three configurations contain the same ordered set of 5,200 preference
examples after source-quality review and exclusion of one Teacher/Student pair
with no response-level preference.
Run-prefixed layout
The original root-level mpo/, cpo/, dpo/, and manifest.json are retained as the legacy Gemma release for compatibility with existing download… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/dqs-post-training.post-training-trackio-datasetpost-training-takehome-math500-bon16
MATH-500 Best-of-16 Post-Training Take-Home Results
A 50-problem study of test-time compute, based on the Hugging Face post-training take-home challenge. Nothing here trains or modifies a model: both the generator and the reward model stay frozen, and the only variable is how a final answer is chosen from 16 sampled candidates.
Construction
Filtered MATH-500 to levels 1-3, shuffled with seed 1, and selected 50 rows.
Generated one greedy solution per problem with… See the full description on the dataset page: https://huggingface.co/datasets/augustoFranke/post-training-takehome-math500-bon16.posttraining-eval-results-extensionposttraining-eval-results-extensionlucky-initialization-posttraining-100m-v3Bllossom-r1-Post-Training-Dataset_0721Bllossom-r1-Post-Training-Dataset_0721_filteredBllossom-r1-Post-Training-Dataset_0721_filteredNemotron-Post-Training-Dataset-Safety
