datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flashmini-data-v1
FlashMini data v4 (card)
Deterministic FlashMini training corpus. Canonical documents live in
Parquet+ZSTD shards under shards/; each shard carries a manifest with
sha256, counts, and distributions; the frozen corpus identity is
corpus_fingerprint_sha256.
Sources and redistribution: each source carries one of mirror_allowed,
recipe_only, gated_recipe_only, review_required, generated_owned
(fail-closed; see registry/sources.yaml + source_snapshot.lock.json).
Content shards are… See the full description on the dataset page: https://huggingface.co/datasets/mjaso/flashmini-data-v1.glm53-flash-harvest
GLM-5.3-Flash On-Policy Harvest
86,006 responses / 246,034,910 generated tokens written by
zai-org/GLM-5.3-Flash from its reference FP8 weights,
across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the
model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model
actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to
learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.glm-5.3-flash-distillation-chat
Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI).
Split
train — successful generations only.
field
description
instruction
user prompt from FineToMe
response
GLM final answer (message.content)
reasoning
GLM chain-of-thought (reasoning_content), empty if not captured
finish
stop or length
prompt_tokens / completion_tokens / reasoning_tokens
usage
latency_s
request latency
source_index
original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x
DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows
Teacher-distillation corpus generated with
deepseek-ai/DeepSeek-V4-Flash-0731.
The original manifest contained 45,000 unique seeds.
Following generation, QC, retry-based repair, quarantine auditing,
and recovery adjudication, 40,513 rows were retained.
Composition
Bucket
Rows
Coding
5,601
Agentic
9,982
Cyber blue
13,000
Controlled cyber red
6,999
Tool use
4,931
Total
40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928
Scale-SWE DeepSeek V4 Flash 0731 Think Rollouts
Successful AweAgent trajectories generated with deepseek-v4-flash-0731 in think mode.
Dataset summary
Source task instances: 3,393
Rollouts per source instance: 4
Total attempted rollouts: 13,572
Successful exported trajectories: 7,928
Unique instances represented by successful trajectories: 2,250
Scaffold: aweagent
Tool-call format: openai_function
The export retains assistant reasoning_content, function tool… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.Finch-Collection-Gemini-3-Flash
Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks
A mid-training "practice phase" that teaches small open-source LLMs how to evolve solutions.
👋 This is the Gemini-3-Flash teacher variant of the Finch Collection — evolutionary search trajectories from the paper Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks, but with Gemini-3-Flash as the teacher mutation… See the full description on the dataset page: https://huggingface.co/datasets/minnesotanlp/Finch-Collection-Gemini-3-Flash.RAGPulse
RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems
🌐 Github Link |
🤗 Workload Trace |
📑 Arxiv Paper |
🤖 How to use?
RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/flashserve/RAGPulse.bird-train-gemini3-flash
Dataset Card for Think2SQL-SFT
This dataset is a distilled Supervised Fine-Tuning (SFT) dataset designed to improve the reasoning capabilities of models in Text-to-SQL tasks.
It contains high-quality reasoning traces and SQL queries generated by Gemini 3 Flash.
Paper: Think2SQL: Blueprinting Reward Density and Advantage Scaling for Effective Text-To-SQL Reasoning
Base Benchmark: BIRD-Train
Dataset Description
The dataset consists of 9,428 high-quality traces, of… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-2321/bird-train-gemini3-flash.llm_timeline_deepseek_v4_flash-pi
Coding agent session traces
This dataset contains coding agent session traces collected while working on LLM Timeline web app using the prompt from coding-agent-bench-prompts
deepseek-v4-flash-swe-cot
DeepSeek-V4-Flash SWE Agent Trajectories (with raw chain-of-thought)
795 multi-turn software-engineering agent trajectories generated by
DeepSeek-V4-Flash-0731 at reasoning_effort=max, each one executed in a real
repository inside an isolated container and verified by running the repository's own
tests. 469 are verified-correct.
Every assistant turn preserves reasoning_content — the model's raw chain-of-thought,
not a summary. That is the point of this dataset: the DeepSeek API… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-flash-swe-cot.med-synth-questions-gemma-3-27b-deepseek-v4-flash
Med Synth Questions (Gemma-3 + DeepSeek V4 Flash)
Synthetic reasoning traces and answers for medical questions from openmed-community/med-synth-questions-gemma-3-27b-it. Each record contains a medical question with SYNTH-style reasoning and a generated answer by DeepSeek V4 Flash.
Dataset Summary
29,148 records (2 dupes + 3,410 incomplete/truncated removed from 32,560 source)
29,148 reasoning turns (99.2% format compliance)
Average 1,591 chars per reasoning trace… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/med-synth-questions-gemma-3-27b-deepseek-v4-flash.swe2-gpt56-luna-glm53-flash-distillation
SWE 2, GPT 5.6 LUNA, GLM 5.3 FLASH DISTILLATION
Devin CLI reasoning distillation — v2.
Distillation dataset built from Devin CLI session traces, containing internal
reasoning traces (chain-of-thought / thinking), user prompts, assistant answers,
system prompts, and tool calls. Format is identical to
Akahsizrr/devin-cli-reasoning-distillation (OpenAI-style message lists with a
reasoning_content field).
Dataset Summary
Total rows
160 (152 train / 8… See the full description on the dataset page: https://huggingface.co/datasets/Akahsizrr/swe2-gpt56-luna-glm53-flash-distillation.denovoswe-distill2767-deepseek-v4-flash-0731-rollout8-instance674-trajectories3985
DeNovoSWE Distill 2767 — DeepSeek V4 Flash NL2Repo Trajectories
This dataset contains 3,985 difficulty-filtered NL2Repo SFT trajectories from 674 repository
instances. Each instance was sampled with eight rollouts using deepseek-v4-flash-0731.
Selection
Tasks receive a static difficulty score derived only from Stage 1–4 artifacts. The 1,141
successful tasks are split into five equal-count difficulty bins. A rollout is retained when its
evaluator score is greater… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/denovoswe-distill2767-deepseek-v4-flash-0731-rollout8-instance674-trajectories3985.dsv4-flash-tmax-git-pager-recovery-23
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use.
Composition
6 original parser-clean full last-episode exports.
17 additional manually confirmed… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery-23.Step-3.5-Flash-Instruct-EmMcts
Step-3.5-Flash-Instruct-EmMcts
Preference (chosen / rejected) dataset generated with an Empirical-MCTS (Em-Mcts) rollout pipeline
and scored by a reward model. Every sample contains a higher-quality chosen response and a
lower-quality rejected response for the same prompt, making it suitable for DPO / preference
optimization and reward-model training.
Overview
Records: 4,959
Format: JSON Lines (one JSON object per line)
Language: English
Generation model:… See the full description on the dataset page: https://huggingface.co/datasets/Minami-su/Step-3.5-Flash-Instruct-EmMcts.dsv4-flash-tmax-git-pager-recovery
DeepSeek V4 Flash TMax Git Pager Recovery
This dataset contains 6 manually audited, SFT-ready terminal-agent trajectories generated by DeepSeek-V4-Flash-0731 in public TMax environments. The primary subset is deliberately narrow: the agent must actually enter a Git pager or foreground TUI, execute a useful recovery action, return to a shell prompt, and finish the task with reward 1.0.
Source and collection
Environment/task source: TMaxxx/TMax-15K-Harbor, pinned… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery.
