datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
orca-dpo-deepseek-v3.2
Orca DPO - DeepSeek V3.2
An updated successor to argilla/distilabel-intel-orca-dpo-pairs, bringing the preference pairs from the GPT-4 era into 2026.
We took the original Intel/orca_dpo_pairs prompts, generated fresh responses with DeepSeek V3.2, and scored all pairs with Skywork Reward V2 to determine which response is preferred.
What's in the dataset
Each row contains a prompt with two responses — a winner (chosen) and a loser (rejected) — determined by reward model… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/orca-dpo-deepseek-v3.2.DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.
