datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mycotoxin-chemical-research-sythetic-reasoning
mycotoxin-chemical-research-sythetic-reasoning
Synthetic Q&A dataset on Mycotoxin Chemical Research, generated with SDGS (Synthetic Dataset Generation Suite).
Dataset Details
Metric
Value
Topic
Mycotoxin Chemical Research
Total Q&A Pairs
4416
Valid Pairs
4416
Provider/Model
ollama/gpt-oss:120b
Generation Cost
Metric
Value
Prompt Tokens
4,579,253
Completion Tokens
5,326,284
Total Tokens
9,905,537
GPU Energy
3.5712 kWh… See the full description on the dataset page: https://huggingface.co/datasets/Kylan12/mycotoxin-chemical-research-sythetic-reasoning.verified-research-reasoning-trajectories
Verified Research Reasoning Trajectories for RLVR
This repository is the public sample and schema repository for Ulam's research-level mathematical reasoning trajectories for reinforcement learning with verifiable rewards (RLVR), process supervision, judge training, proof criticism, and private evaluations.
Ulam Verified Research Reasoning Trajectories are proof-process data for RLVR. Each record contains a normalized research problem, a golden or partial-golden proof graph… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/verified-research-reasoning-trajectories.kimi-k3-cyber-reasoning-distill
Kimi Cyber Reasoning
997 chain-of-thought records covering 13 cybersecurity disciplines and 4 systems engineering domains, distilled from the Kimi K3 reasoning model via API. Every record provides an explicit step-by-step <think> reasoning trace followed by a technical resolution, unified code diff fix, or structured tool invocation.
The dataset was curated as an anchor set for training, healing, and specializing compact reasoning models on systems security and tool calling… See the full description on the dataset page: https://huggingface.co/datasets/p-research/kimi-k3-cyber-reasoning-distill.Opus-4.6-Reasoning-2160x
Opus-4.6-Reasoning-2160x
2,160 high-quality reasoning traces generated by Claude Opus 4.6 via OpenRouter, covering mathematics, competitive programming, logic, science, and language tasks. Each example includes the full problem, an extended chain-of-thought, and a final solution — making the dataset suitable for supervised fine-tuning, chain-of-thought distillation, and reasoning-capability transfer to smaller models.
Originally generated as a batch of 3,305 examples; 1,145 were… See the full description on the dataset page: https://huggingface.co/datasets/Glint-Research/Opus-4.6-Reasoning-2160x.MathChatSync-reasoning
MathChatSync-reasoning
The dataset was presented in the paper One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning.
The code for the paper can be found at: https://github.com/devrev/One-Pass-to-Reason
Dataset Description
MathChatSync-reasoning is a synthetic dialogue-based mathematics tutoring dataset designed to enable supervised training with explicit step-by-step reasoning. This dataset augments the… See the full description on the dataset page: https://huggingface.co/datasets/devrev-research/MathChatSync-reasoning.research-paper-agent-reasoning-traces-unverifiedcleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research元データ: https://huggingface.co/datasets/moremilk/CoT_Reasoning_Scientific_Discovery_and_Research
使用したコード: https://github.com/LLMTeamAkiyama/0-data_prepare/tree/master/src/CoT_Reasoning_Scientific_Discovery_and_Research
データ件数: 3,733
平均トークン数: 1,193
最大トークン数: 2,489
合計トークン数: 4,453,517
ファイル形式: JSONL
ファイル分割数: 1
合計ファイルサイズ: 23.2 MB
加工内容:
メタデータ列の解析と新列生成: metadata列(辞書型)を解析し、その中のreasoningをthought列に、difficultyをdifficulty列に展開しました。解析に失敗した行は除外されました。また、元のmetadata列は削除されました。
難易度によるフィルタリング:… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_moremilk_CoT_Reasoning_Scientific_Discovery_and_Research.
