datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cvalues_samplesThis dataset contains samples of the cvalues english dataset used for training domain invariant reward models by few-shot generalization.
Each just_sample contains the sample (10 examples) by itself, while the sampler files contain the sample repeated 1000 times for a total of 10000 examples.
References:
Cvalues: https://github.com/X-PLUG/CValues
Cvalues english dataset: https://huggingface.co/datasets/david9dragon9/cvalues-english
DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.qwen-reasoning-samples-20260421_221240
Frontier-Class Synthetic Reasoning Samples
Dataset Description
10 synthetic reasoning examples generated with Qwen/Qwen3.6-35B-A3B via vLLM, using a structured prompt designed to elicit frontier-level (Opus 4.7 class) multi-phase reasoning.
Each example is gated through a quality filter that requires the reasoning trace to follow an explicit 6-phase structure (Understand → Decompose → Explore → Execute → Verify → Reflect) and to include an independent verification step.… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/qwen-reasoning-samples-20260421_221240.AI_healthcare_QA_samples_Sonnet3.5
Note
To create data [ChuGyouk/AI_healthcare_QA], I extracted a few samples and get answer from claude-3-5-sonnet-20240620 (by using welcome credit).
(I did this briefly for debugging purposes.)
You are an AI assistant acting in the role of a professional doctor. Your task is to provide reliable and helpful answers to health-related questions posed by patients.
You will be presented with a question from a patient. Your goal is to answer this question professionally, accurately… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI_healthcare_QA_samples_Sonnet3.5.science-qa-samples
Science Q&A Samples
This sample shows structured science question-answer pairs for reviewing subject coverage, difficulty labeling, and answer format before scoping a larger educational dataset.
What This Shows
Q&A examples across science and math subjects
Metadata for topic, difficulty, curriculum alignment, and question type
A view of how text and asset-backed questions are represented
Dataset Specifications
Field
Value
Modality
Text… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/science-qa-samples.industry-intelligence-graph-samples
Fodda Industry Intelligence — Graph Samples
Expert-curated knowledge graph slices for AI agents and LLM fine-tuning.
This dataset contains JSON-LD samples from Fodda's five core domain knowledge graphs — showing the top trending topics across Retail, Beauty, Sports, Fashion, and Culture.
These are slices of a much larger interconnected intelligence system.
What Fodda Is
Fodda is an AI context layer built on PSFK's 20+ years of editorial expertise. It structures… See the full description on the dataset page: https://huggingface.co/datasets/Fodda-ai/industry-intelligence-graph-samples.qwen-reasoning-samples-20260421_215219
Frontier-Class Synthetic Reasoning Samples
Dataset Description
1 synthetic reasoning examples generated with Qwen/Qwen3.6-35B-A3B via vLLM, using a structured prompt designed to elicit frontier-level (Opus 4.7 class) multi-phase reasoning.
Each example is gated through a quality filter that requires the reasoning trace to follow an explicit 6-phase structure (Understand → Decompose → Explore → Execute → Verify → Reflect) and to include an independent verification step.… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/qwen-reasoning-samples-20260421_215219.qwen-reasoning-samples-20260421_220553
Frontier-Class Synthetic Reasoning Samples
Dataset Description
3 synthetic reasoning examples generated with Qwen/Qwen3.6-35B-A3B via vLLM, using a structured prompt designed to elicit frontier-level (Opus 4.7 class) multi-phase reasoning.
Each example is gated through a quality filter that requires the reasoning trace to follow an explicit 6-phase structure (Understand → Decompose → Explore → Execute → Verify → Reflect) and to include an independent verification step.… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/qwen-reasoning-samples-20260421_220553.
