datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cvalues_samplesThis dataset contains samples of the cvalues english dataset used for training domain invariant reward models by few-shot generalization.
Each just_sample contains the sample (10 examples) by itself, while the sampler files contain the sample repeated 1000 times for a total of 10000 examples.
References:
Cvalues: https://github.com/X-PLUG/CValues
Cvalues english dataset: https://huggingface.co/datasets/david9dragon9/cvalues-english
math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on MATH500
This dataset contains 99 self-consistency generations per question for the
MATH500 benchmark, produced with allenai/OLMo-3-7B-Instruct at temperature
0.9, together with token-level log probabilities for each completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: HuggingFaceH4/MATH-500
Model:… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/math500-olmo-3-7b-instruct-temp0.9-samples99-logprobs.DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.qwen-reasoning-samples-20260421_221240
Frontier-Class Synthetic Reasoning Samples
Dataset Description
10 synthetic reasoning examples generated with Qwen/Qwen3.6-35B-A3B via vLLM, using a structured prompt designed to elicit frontier-level (Opus 4.7 class) multi-phase reasoning.
Each example is gated through a quality filter that requires the reasoning trace to follow an explicit 6-phase structure (Understand → Decompose → Explore → Execute → Verify → Reflect) and to include an independent verification step.… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/qwen-reasoning-samples-20260421_221240.AI_healthcare_QA_samples_Sonnet3.5
Note
To create data [ChuGyouk/AI_healthcare_QA], I extracted a few samples and get answer from claude-3-5-sonnet-20240620 (by using welcome credit).
(I did this briefly for debugging purposes.)
You are an AI assistant acting in the role of a professional doctor. Your task is to provide reliable and helpful answers to health-related questions posed by patients.
You will be presented with a question from a patient. Your goal is to answer this question professionally, accurately… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/AI_healthcare_QA_samples_Sonnet3.5.mmlu-pro-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on MMLU-Pro
This dataset contains 99 self-consistency generations per question for the
MMLU-Pro test split, produced with allenai/OLMo-3-7B-Instruct at
temperature 0.9, together with token-level log probabilities for each
completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: TIGER-Lab/MMLU-Pro
Model:… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/mmlu-pro-olmo-3-7b-instruct-temp0.9-samples99-logprobs.popqa-olmo-3-7b-instruct-temp0.9-samples99-logprobs
OLMo-3-7B-Instruct self-consistency generations with logprobs on PopQA
This dataset contains 99 self-consistency generations per question for the
PopQA benchmark, produced with allenai/OLMo-3-7B-Instruct at temperature
0.9, together with token-level log probabilities for each completion.
The file is intended for post-hoc analysis, self-consistency curves, adaptive
stopping, and related aggregation methods.
Source
Base benchmark: PopQA
Model: allenai/OLMo-3-7B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/popqa-olmo-3-7b-instruct-temp0.9-samples99-logprobs.science-qa-samples
Science Q&A Samples
This sample shows structured science question-answer pairs for reviewing subject coverage, difficulty labeling, and answer format before scoping a larger educational dataset.
What This Shows
Q&A examples across science and math subjects
Metadata for topic, difficulty, curriculum alignment, and question type
A view of how text and asset-backed questions are represented
Dataset Specifications
Field
Value
Modality
Text… See the full description on the dataset page: https://huggingface.co/datasets/psdn-ai/science-qa-samples.industry-intelligence-graph-samples
Fodda Industry Intelligence — Graph Samples
Expert-curated knowledge graph slices for AI agents and LLM fine-tuning.
This dataset contains JSON-LD samples from Fodda's five core domain knowledge graphs — showing the top trending topics across Retail, Beauty, Sports, Fashion, and Culture.
These are slices of a much larger interconnected intelligence system.
What Fodda Is
Fodda is an AI context layer built on PSFK's 20+ years of editorial expertise. It structures… See the full description on the dataset page: https://huggingface.co/datasets/Fodda-ai/industry-intelligence-graph-samples.openspaces-depth-aware-32-samples
OpenSpaces Depth-Aware Visual QA Dataset
This is a 32-sample visual question answering (VQA) dataset that includes:
RGB images from the OpenSpaces dataset
Predicted depth maps generated using Depth Anything
3 depth-aware QA pairs per image:
Yes/No question (e.g., “Is there a person near the door?”)
Short answer question (e.g., “What color is the man’s coat?”)
Spatial sorting question (e.g., “Sort the objects from closest to farthest”)
Intended Use
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/srimoyee12/openspaces-depth-aware-32-samples.qwen-reasoning-samples-20260421_215219
Frontier-Class Synthetic Reasoning Samples
Dataset Description
1 synthetic reasoning examples generated with Qwen/Qwen3.6-35B-A3B via vLLM, using a structured prompt designed to elicit frontier-level (Opus 4.7 class) multi-phase reasoning.
Each example is gated through a quality filter that requires the reasoning trace to follow an explicit 6-phase structure (Understand → Decompose → Explore → Execute → Verify → Reflect) and to include an independent verification step.… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/qwen-reasoning-samples-20260421_215219.qwen-reasoning-samples-20260421_220553
Frontier-Class Synthetic Reasoning Samples
Dataset Description
3 synthetic reasoning examples generated with Qwen/Qwen3.6-35B-A3B via vLLM, using a structured prompt designed to elicit frontier-level (Opus 4.7 class) multi-phase reasoning.
Each example is gated through a quality filter that requires the reasoning trace to follow an explicit 6-phase structure (Understand → Decompose → Explore → Execute → Verify → Reflect) and to include an independent verification step.… See the full description on the dataset page: https://huggingface.co/datasets/himanshunakrani9/qwen-reasoning-samples-20260421_220553.QA_hukum_samplesThis dataset sample was constructed by generating QA pairs from Law No. 17 of 2008 on Shipping (Undang-Undang No 17 Tahun 2008 Tentang Pelayaran). It is then manually verified by human validators and reviewers, resulting in QA pairs with fine-grained label categories.
