datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenScienceReasoning-2-50K
OpenScienceReasoning-2-50K
This repository contains a random 50,000 sample subset of the NVIDIA OpenScienceReasoning-2 dataset.
OpenScience-Chinese-Reasoning-SFT
OpenScience-Chinese
A Chinese multiple-choice science QA dataset with chain-of-thought reasoning, derived from nvidia/OpenScience through translation and rejection sampling.
Key Features:
Scale: 50,000 high-quality instances.
Reasoning: Built-in Chain-of-Thought (<think> tags) for interpretable AI.
Reliability: Rejection sampling ensures 100% alignment with ground truth.
Data Source
The questions originate from nvidia/OpenScience, a large-scale science QA dataset… See the full description on the dataset page: https://huggingface.co/datasets/zake7749/OpenScience-Chinese-Reasoning-SFT.
