datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FlashRAG_datasets
⚡FlashRAG: A Python Toolkit for Efficient RAG Research
FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms.
With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components.
For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.medical_meadow_medical_flashcards
Dataset Card for Medical Flashcards
Dataset Summary
Medicine as a whole encompasses a wide range of subjects that medical students and graduates must master
in order to practice effectively. This includes a deep understanding of basic medical sciences, clinical knowledge,
and clinical skills. The Anki Medical Curriculum flashcards are created and updated by medical students and cover the
entirety of this curriculum, addressing subjects such as anatomy, physiology… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_medical_flashcards.ioi-eval-openrouter_google_gemini-2_0-flash-thinking-exp-prompt-mem-limitgemini-flash-2.0-speech
🎙️ Gemini Flash 2.0 Speech Dataset
This is a high quality synthetic speech dataset generated by Gemini Flash 2.0 via the Multimodal Live API. It contains speech from 2 speakers - Puck (Male) and Kore (Female) in English.
🏅 #1 Trending Audio Dataset in Feb 2025
🏅 Used in training of Kokoro TTS and LLaSA 1B
〽️ Stats
Total number of audio files: 47,256*2 = 94512Total duration: 1023527.20seconds (284.31 hours)
Average duration: 10.83 seconds
Shortest file: 0.6… See the full description on the dataset page: https://huggingface.co/datasets/shb777/gemini-flash-2.0-speech.tb21-dsv4-flash-0731-dsh
Terminal-Bench 2.1 trajectories: DeepSeek-V4-Flash-0731 + dsh sdk-minimal
Every trial of this one line, in one place: the 89-task main run, both re-run passes, and
the scoring scripts. The trajectories are raw and unedited — each step's reasoning, each
tool call, and the verifier's own stdout.
This is a re-packaging, not a new measurement. The same files were published before,
split across two releases, which made the line look incomplete in both: the first release
carried the… See the full description on the dataset page: https://huggingface.co/datasets/openguardrails/tb21-dsv4-flash-0731-dsh.GLM-5.3-Flash-calibration-activations-v1
GLM-5.3-Flash calibration activations v1 (BF16, natural routing)
Per-layer block-input activations of zai-org/GLM-5.3-Flash-BF16 @ b1967181 over 92x2048
tokens of the exllamav3 standard_cal_data corpus (pinned): per context, layer_NNN.attn_in
and layer_NNN.mlp_in (bf16, post-norm linear inputs; mlp_in is the router + expert gate/up
input) and layer_NNN.router_logits (fp32, natural top-8 routing ground truth).
Per-expert Hessians E[xx^T], routing statistics and down-proj inputs… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/GLM-5.3-Flash-calibration-activations-v1.Taur_CoT_Analysis_Project___google__gemini-1.5-flash-001glm53-flash-harvest
GLM-5.3-Flash On-Policy Harvest
86,006 responses / 246,034,910 generated tokens written by
zai-org/GLM-5.3-Flash from its reference FP8 weights,
across four harvest rounds, 15 registers and both serving modes (22,016 rows carry the
model's inline <think>…</think> chain). It is on-policy text: the corpus records what the target model
actually generates, which is what a speculative-decoding drafter (EAGLE-3 / DFlash / DSpark family) has to
learn to predict. Everything here is MIT.… See the full description on the dataset page: https://huggingface.co/datasets/Zek-Takai/glm53-flash-harvest.bird-critic-1.0-flash-exp
BIRD-CRITIC-1.0-Flash
BIRD-Critic is the first SQL debugging benchmark designed to answer a critical question:
Can large language models (LLMs) fix user issues in real-world database applications? Each task in BIRD-CRITIC has been verified by human experts on the following dimensions:
Reproduction of errors on BIRD env to prevent data leakage.
Carefully curate test case functions for each task specifically.
Soft EX: This metric can evaluate SELECT-ONLY tasks.
Soft EX + Parsing:… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/bird-critic-1.0-flash-exp.wavenet_flashback
Dataset Card for "wavenet_flashback"
https://cloud.google.com/text-to-speech/docs/reference/rest/v1/text/synthesize#AudioConfig
sv-SE-Wavenet-{voice}
https://spraakbanken.gu.se/resurser/flashback-dator
ox-alpha-glm-5.3-flash-distillation-coding-17k-raw
Ox Alpha GLM-5.3-Flash Distillation Coding 17K Raw
A raw collection of 17,138 synthetic coding samples generated with GLM-5.3-Flash, previously exposed through OpenCode under the stealth-model alias Ox Alpha.
The dataset is intended for experimentation with LLM distillation, code-generation models, instruction tuning, supervised fine-tuning, evaluation, and agentic coding systems.
flash-flood-benchmark-data
TORRENT — CONUS Flash-Flood Benchmark (L1–L3), agent-friendly
Traceable, Observation-constrained, Rapid-Response, Episode–gauge–watershed
Network of Testbeds: reproducible flash-flood testbeds for hydrological-response
analysis and model intercomparison. This mirror carries the paper-matched
v1.0 release (companion paper: TORRENT, Earth System Science Data).
Archive of record (v1.0): https://doi.org/10.5281/zenodo.22118051
(concept DOI, always the latest version:… See the full description on the dataset page: https://huggingface.co/datasets/skyan1002/flash-flood-benchmark-data.glm-5.3-flash-distillation-chat
Private distill of domofon/finetome-cot-100k instructions through GLM-5.3-Flash (AutoClaw / Z.AI).
Split
train — successful generations only.
field
description
instruction
user prompt from FineToMe
response
GLM final answer (message.content)
reasoning
GLM chain-of-thought (reasoning_content), empty if not captured
finish
stop or length
prompt_tokens / completion_tokens / reasoning_tokens
usage
latency_s
request latency
source_index
original FineToMe… See the full description on the dataset page: https://huggingface.co/datasets/best-distill/glm-5.3-flash-distillation-chat.Step-3.5-Flash-SFT-No-Tools
Step-3.5-Flash-SFT No-Tools
Filtered subset of stepfun-ai/Step-3.5-Flash-SFT containing only plain chat rows from the raw JSON shards.
Final kept rows: 1493471
No-tool rows before secret filtering: 1495099
Rows removed by accepted secret scan findings: 1628
Primary data files are Parquet shards under data/train-*.parquet.
Filter predicate:
conversations must be a list,
every message must be an object,
message roles must be limited to system, user, and assistant,
no message may… See the full description on the dataset page: https://huggingface.co/datasets/MetonymousAI/Step-3.5-Flash-SFT-No-Tools.Qwen3.8-Flash-Next-GGUF-metricsFlashST-DATAdpsk-v4-flash-data
dpsk-v4-flash-data
Training data built by the AgentPTB arm for cell dpsk-v4-flash — pi / DeepSeek v4-flash @ effort thinking.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/dpsk-v4-flash.h*, and the companion to the run record in agentic-ptb/dpsk-v4-flash-record.
field
value
plot cell
dpsk-v4-flash
driver
pi / DeepSeek v4-flash… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/dpsk-v4-flash-data.gemini31-flash-lite-train20260729_mini-v2.2.8_gemini-3-5-flashglm53-flash-fidelity-root-v1
fidelity--glm53flash.malaiwah.root.bf16
A root fidelity dataset in hidden form, produced by engines/tools/hf_capture.py from zai-org/GLM-5.3-Flash-BF16.
The cut
the final hidden state handed to lm_head -- after the text model's final norm and immediately before the head matmul -- captured as the head module's input via torch.nn.Module.register_forward_pre_hook; replay applies the head ONLY (no final norm at replay time: the capture already sits after it). Same cut… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/glm53-flash-fidelity-root-v1.SWE-smith-rs-gemini-3-flash-trajectories
Trajectories Dataset
Top-level fields:
messages
instance_id
resolved
model
traj_id
patch
Generated at: 2026-02-27 00:16:39Z
Rows: 1449
Shards: 6
Skipped runs (missing/corrupt trajectory): 1
DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x
DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows
Teacher-distillation corpus generated with
deepseek-ai/DeepSeek-V4-Flash-0731.
The original manifest contained 45,000 unique seeds.
Following generation, QC, retry-based repair, quarantine auditing,
and recovery adjudication, 40,513 rows were retained.
Composition
Bucket
Rows
Coding
5,601
Agentic
9,982
Cyber blue
13,000
Controlled cyber red
6,999
Tool use
4,931
Total
40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.20260731_mini-v2.4.2_gemini-3-6-flashWikipedia-FA-EN-DeepSeek-V4-Flash-0731
Wikipedia Persian to English — DeepSeek V4 Flash 0731
Rolling, machine-generated English translations of Persian Wikipedia articles
from Reza2kn/Wikipedia-EN-FA-Accessibility-Bridge, configuration
full_articles_fa_without_en. 129,816 translations are
currently published in 26 immutable Parquet shards.
The target release contains 129,816 translations;
five source rows have empty plain_text and are not translated. Shards are
published only after 5,000 complete, validated records… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/Wikipedia-FA-EN-DeepSeek-V4-Flash-0731.DeepSeek-v4-Flash-ChatThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
Teich Test
This directory contains newline-delimited JSON training examples generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-flash.
Rows: 6313
Format
Each file is newline-delimited JSON where every line is already a training example.
Chat-only datasets include messages… See the full description on the dataset page: https://huggingface.co/datasets/TeichAI/DeepSeek-v4-Flash-Chat.gemini-3.7-flash-ocr-26-aug-2026
gemini-3.7-flash-ocr-26-aug-2026
Page-image → transcription pairs for finetuning a vision-language model to OCR
Devanagari and Tamil printed books.
These labels are not human ground truth. They are the output of a teacher
model, so its accuracy is the ceiling for anything trained on them.
Provenance
Teacher model
google/gemini-3.7-flash (via OpenRouter, reasoning.effort=low)
Page render
PyMuPDF at 200 DPI, grayscale JPEG q90
Sampling
stratified —… See the full description on the dataset page: https://huggingface.co/datasets/kailasa-ngpt/gemini-3.7-flash-ocr-26-aug-2026.Gemini-2.0-Flash-Aoede-Voicescale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928
Scale-SWE DeepSeek V4 Flash 0731 Think Rollouts
Successful AweAgent trajectories generated with deepseek-v4-flash-0731 in think mode.
Dataset summary
Source task instances: 3,393
Rollouts per source instance: 4
Total attempted rollouts: 13,572
Successful exported trajectories: 7,928
Unique instances represented by successful trajectories: 2,250
Scaffold: aweagent
Tool-call format: openai_function
The export retains assistant reasoning_content, function tool… See the full description on the dataset page: https://huggingface.co/datasets/wjn922-01/scale-swe-distill5000-deepseek-v4-flash-0731-think-rollout4-instance3393-trajectories7928.tb2-k5-n4-flashmedical-meadow-medical-flashcards
Dataset Card for medical-meadow-medical-flashcards
This dataset originates from the medAlpaca repository.
The medical-meadow-medical-flashcards dataset is specifically used for models training of medical question-answering.
Dataset Details
Dataset Description
Each sample is comprised of three columns: instruction, input and output.
Language(s): English
Dataset Sources
The code from the original repository was adopted to post it here.
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/medical-meadow-medical-flashcards.
