datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LLM_Post_Training_BestOfNvsGreedy
Best-of-N vs Greedy Model Dataset
This dataset contains problems with their corresponding solutions from two models using Qwen (Qwen2.5-1.5B-Instruct):
Greedy Model: A model that picks the most optimal solution at each step.
Best-of-N Model: A model that generates multiple solutions and picks the best one from a set of N using Skywork (Skywork-o1-Open-PRM-Qwen-2.5-1.5B) reward model.
Dataset Details
Number of Problems: 20 (for testing purposes).
Columns:
id: Unique… See the full description on the dataset page: https://huggingface.co/datasets/QPM777/LLM_Post_Training_BestOfNvsGreedy.Hindi-Gemma-Post-Training
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/1024m/Hindi-Gemma-Post-Training.
