datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.HMMT_2025
Dataset Summary
This dataset comprises the questions, answers, and solutions from HMMT February 2025, all of which were extracted by OCR, converted to LaTeX, and manually verified by FlagEval Team.
Data Fields
Below one can find the description of each field in the dataset.
id (str): Index of the problem in the competition
problem (str): Full problem statement
answer (str): Ground-truth answer to the question
solution(str): Ground-truth solution to the question… See the full description on the dataset page: https://huggingface.co/datasets/FlagEval/HMMT_2025.flan2021-full
Task Name
FLAN-2021 -> 70
{
"ag_news_subset": 108497,
"ai2_arc/ARC-Challenge": 829,
"ai2_arc/ARC-Easy": 1927,
"aeslc": 13187,
"anli/r1": 15361,
"anli/r2": 41133,
"anli/r3": 91048,
"bool_q": 8343,
"cnn_dailymail": 259607,
"coqa": 6456,
"cosmos_qa": 22996,
"definite_pronoun_resolution": 1079,
"drop": 70045,
"fix_punct": 25690,
"gem/common_gen": 60936,
"gem/dart": 56724,
"gem/e2e_nlg": 30337,
"gem/web_nlg_en": 31899… See the full description on the dataset page: https://huggingface.co/datasets/aslawliet/flan2021-full.Flames-1k-Chinese
FLAMES: Benchmarking Value Alignment of LLMs in Chinese
Introduction
🏠 Homepage | 👍 Our Official Code Repo
This repository organizes the data from FLAMES: Benchmarking Value Alignment of LLMs in Chinese, facilitating evaluation using align-anything.
Citation
The evaluation script for Flames is released in the align-anything repository.
Please cite the repo if you find the benchmark and code in this repo useful 😊
@inproceedings{ji2024align,
title={Align… See the full description on the dataset page: https://huggingface.co/datasets/PKU-Alignment/Flames-1k-Chinese.DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x
DeepSeek V4 Flash 0731 Teacher Distillation — 40,513 Retained Rows
Teacher-distillation corpus generated with
deepseek-ai/DeepSeek-V4-Flash-0731.
The original manifest contained 45,000 unique seeds.
Following generation, QC, retry-based repair, quarantine auditing,
and recovery adjudication, 40,513 rows were retained.
Composition
Bucket
Rows
Coding
5,601
Agentic
9,982
Cyber blue
13,000
Controlled cyber red
6,999
Tool use
4,931
Total
40,513… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/DeepSeek-V4-Flash-0731-Teacher-Distillation-40513x.constructcie
Dataset Card for ConstructCIE
ConstructCIE is a dataset for extracting causal information from construction accident narratives. Each accident report is annotated with a hierarchy of causal factors.
Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
Dataset Details
Dataset Description
The dataset contains 530 English construction accident narratives drawn from OSHA accident investigation summaries published… See the full description on the dataset page: https://huggingface.co/datasets/lab-flair/constructcie.KrynexAI-Dataset-Flash-Instruction
🧠 KrynexAI Dataset
English | Русский
📌 Overview
KrynexAI Dataset is a high-quality, synthetically expanded collection of 10,000+ instruction-response pairs designed for fine-tuning Large Language Models (LLMs).
The dataset covers a wide range of topics including:
💻 Programming (Python, algorithms, data structures)
🤖 AI & Machine Learning (neural networks, transformers, LLMs)
🔭 Science (physics, cosmology, biology, neuroscience)
🧠 Philosophy & Psychology… See the full description on the dataset page: https://huggingface.co/datasets/KrynexLabs/KrynexAI-Dataset-Flash-Instruction.swebench-verified-deepseek-v4-flash-failure-analysis
SWE-bench Verified runs & failure analysis — DeepSeek-V4-flash (local) × mini-swe-agent
Per-instance analysis of SWE-bench Verified runs of a locally-served DeepSeek-V4-flash model
driven by mini-swe-agent, graded with the official
SWE-bench harness. Each instance carries the full agent trajectory, a readable transcript, the
submitted patch, the harness test output, deterministic metrics, and a hand-verified qualitative
root-cause diagnosis.
Current numbers (resolve rates… See the full description on the dataset page: https://huggingface.co/datasets/daaain/swebench-verified-deepseek-v4-flash-failure-analysis.HALO-Gemini-3-Flash-AppWorld
Dataset Card: Gemini 3 Flash Traces on AppWorld (test-normal)
Dataset Overview
This dataset contains agent execution traces of Gemini 3 Flash running on the AppWorld benchmark, specifically evaluated on the test-normal dataset split. The traces capture the full span-level execution detail of the model interacting with AppWorld's simulated app ecosystem.
Field
Value
Model
Gemini 3 Flash
Benchmark
AppWorld
Split
test-normal
Total Traces
168
Total Spans
3… See the full description on the dataset page: https://huggingface.co/datasets/inference-net/HALO-Gemini-3-Flash-AppWorld.glm-5.3-flash-mathnet-bon
glm-5.3-flash-mathnet-bon
This is the continuation and the final set of ox-alpha-mathnet-bon.
Verified chain-of-thought reasoning traces for competition mathematics, generated with GLM-5.3-Flash via best-of-N rejection sampling against the ShadenA/MathNet dataset (ICLR 2026).
Statistics (this split)
Metric
Value
Records (problem × attempt)
6,181
Distinct problems
848
Attempts per problem
7.29 (mean), 8 (max)
Accepted (answer_correct = true)
3,705… See the full description on the dataset page: https://huggingface.co/datasets/zakoman/glm-5.3-flash-mathnet-bon.RAGPulse
RAGPulse: A Real-World RAG Workload Trace to Optimize RAG Serving Systems
🌐 Github Link |
🤗 Workload Trace |
📑 Arxiv Paper |
🤖 How to use?
RAGPulse is a real-world RAG workload trace collected from an university-wide Q&A service scenario. The system has been serving over 40,000 students and faculties since April 2024, providing intelligent policy Q&A services. The trace contains a total of 7,106 records entries, sampled from one week of our Q&A service.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/flashserve/RAGPulse.agentic-code
Unified Agentic Coding CoT Dataset
This dataset is a curated fusion of high-quality agentic coding trajectories, specifically optimized for fine-tuning small, high-performance models like Qwen2.5-Coder-0.5B-Instruct. It combines systematic reasoning (Chain-of-Thought) with practical tool-use and code editing capabilities.
Dataset Summary
The dataset unifies two primary sources into a single, instruction-following format:… See the full description on the dataset page: https://huggingface.co/datasets/FlameF0X/agentic-code.prompt-policy-memory-v0
Prompt Policy Memory v0
Synthetic profile-memory data: 100 training sessions from10users;20test sessions from2fresh users. Test users were generated after the GRPO checkpoint was frozen and must not be used for training or tuning.
Each row includes cumulative plain-text session input, a canonical plain-text key:value reference, chat messages, and evaluator-only target data. messages can be used for supervised fine-tuning. The reference contains all currently revealed facts; it… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/prompt-policy-memory-v0.allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
PJMixers-Dev/allenai_WildChat-1M-prompts with responses generated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.llm_timeline_deepseek_v4_flash-pi
Coding agent session traces
This dataset contains coding agent session traces collected while working on LLM Timeline web app using the prompt from coding-agent-bench-prompts
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bghira/pseudo-camera-10k with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/bghira_pseudo-camera-10k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Deepseek-V4-Flash-11000x
Sherlock Thinking Alpha DeepSeek V4 Flash Distillation
Seed Prompt Dataset
Prompts are sourced from TeichAI/sherlock-thinking-alpha-11000x.
Model
Solutions and reasoning traces were generated with deepseek-ai/DeepSeek-V4-Flash.
DeepSeek-V4-Flash is part of the DeepSeek-V4 preview series. Its model card describes it as a Mixture-of-Experts language model with 284B total parameters, 13B activated parameters, and a 1M-token context length. The model repository is… See the full description on the dataset page: https://huggingface.co/datasets/SLoonker/Deepseek-V4-Flash-11000x.Phoenix-SFT-v1
Phoenix v1 — Memory Extraction Dataset
Synthetic SFT training data for Phoenix, a small (Qwen2.5-class) model
that reads a window of chat messages between a user and an AI character
("Flame") and emits a structured JSON list of memorable facts the Flame
should remember about the user going forward.
This is the v1 dataset shipped by flammen.ai for
the memory layer that drives long-term continuity in async character
conversations. Released as part of flammen.ai's open-source data… See the full description on the dataset page: https://huggingface.co/datasets/flammenai/Phoenix-SFT-v1.deepseek-v4-flash-swe-cot
DeepSeek-V4-Flash SWE Agent Trajectories (with raw chain-of-thought)
795 multi-turn software-engineering agent trajectories generated by
DeepSeek-V4-Flash-0731 at reasoning_effort=max, each one executed in a real
repository inside an isolated container and verified by running the repository's own
tests. 469 are verified-correct.
Every assistant turn preserves reasoning_content — the model's raw chain-of-thought,
not a summary. That is the point of this dataset: the DeepSeek API… See the full description on the dataset page: https://huggingface.co/datasets/blythet/deepseek-v4-flash-swe-cot.Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Weyaxi/HelpSteer-filtered with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
Some random images with responses/captions generated with gemini-2.0-flash-thinking-exp-1219.
The format should be similar to that of liuhaotian/LLaVA-Instruct-150K.
Images can be found in the images.zip folder. The zip also contains .txt captions for ease of use in non-VQA tasks.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Images/Handpicked-Images-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.flame-kindling-v1
flame-kindling-v1
A small, opinionated SFT dataset for finetuning a 3B-class instruct model into a character designer that emits a strict JSON schema from a free-text seed. Built to replace a general RP model (Mistral-Nemo-12B Mahou finetune) being shoehorned into JSON output for flammen.ai's Create-a-Flame pipeline.
400 (seed → DesignedFlame) pairs distilled from Claude Sonnet 4.5 with tool-forcing, validated against a strict pydantic schema, deduplicated by name and… See the full description on the dataset page: https://huggingface.co/datasets/flammenai/flame-kindling-v1.flatbot-mini-35M-dataset
FlatBuild Demo Chat 10K Dataset
The FlatBuild Demo Chat 10K Dataset is the official conversational training dataset for FlatBuild and is used to train Flatbot-Mini-35M, the flagship demonstration language model of the Flatseek ecosystem.
The dataset showcases the complete workflow of building a conversational language model entirely from scratch, including:
dataset preparation
tokenizer training
chat data preprocessing
Transformer training
checkpoint export
GGUF conversion… See the full description on the dataset page: https://huggingface.co/datasets/flatseek/flatbot-mini-35M-dataset.Hermes-Coworker-Flash
⚡ Hermes Coworker Flash – Fast, No‑Code Agent Conversations
Hermes Coworker Flash is a curated instruction‑style dataset built from the best “everyday assistant” traces of lambda/hermes-agent-reasoning-traces.It combines GLM‑5.1 and kimi-2.5 splits, keeping only the non‑programming, quick‑turnaround co‑worker tasks — the ones a user would ask a fast AI assistant, not a full‑fledged software engineer.
🧹 All chain‑of‑thought (<think>…</think>) has been removed so the model learns to… See the full description on the dataset page: https://huggingface.co/datasets/LiteMind/Hermes-Coworker-Flash.allenai_WildChat-1M-gemini-2.0-flash-exp-ShareGPT
allenai_WildChat-1M-gemini-2.0-flash-exp-ShareGPT
PJMixers-Dev/allenai_WildChat-1M-prompts with responses generated with gemini-2.0-flash-exp.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-exp-ShareGPT.grimulkan_theory-of-mind-gemini-2.0-flash-exp-ShareGPT
grimulkan_theory-of-mind-gemini-2.0-flash-exp-ShareGPT
grimulkan/theory-of-mind with responses generated with gemini-2.0-flash-exp.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was skipped.
model = genai.GenerativeModel(
model_name,
safety_settings=[… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/grimulkan_theory-of-mind-gemini-2.0-flash-exp-ShareGPT.bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT
bigcode/self-oss-instruct-sc2-exec-filter-50k with responses regenerated with gemini-2.0-flash-thinking-exp-1219.
Generation Details
If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped.
If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped.
If ["candidates"][0]["finish_reason"] != 1 the sample was… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/bigcode_self-oss-instruct-sc2-exec-filter-50k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.deepshopper-reward-pairs
DeepShopper Reward pairwise-preference data
(need, chosen=gold outfit, rejected=corrupted outfit, neg_type) pairs for Bradley-Terry
reward training. 93,020 train / 63,584 test, balanced over 5 corruption types:
gender_flip, item_swap, duplicate_role, count_drop, cross_need. Built (scripts/build_reward_pairs.py)
from the gold AMZ bundles via the frozen deepshopper-mapper-reward-splits.
Trains flavianv/qwen4b-reward-pairwise-v1. Code: https://github.com/clijo/reco-rl (branch… See the full description on the dataset page: https://huggingface.co/datasets/flavianv/deepshopper-reward-pairs.Step-3.5-Flash-Instruct-EmMcts
Step-3.5-Flash-Instruct-EmMcts
Preference (chosen / rejected) dataset generated with an Empirical-MCTS (Em-Mcts) rollout pipeline
and scored by a reward model. Every sample contains a higher-quality chosen response and a
lower-quality rejected response for the same prompt, making it suitable for DPO / preference
optimization and reward-model training.
Overview
Records: 4,959
Format: JSON Lines (one JSON object per line)
Language: English
Generation model:… See the full description on the dataset page: https://huggingface.co/datasets/Minami-su/Step-3.5-Flash-Instruct-EmMcts.UltraSteer-v0-flatNote 0: This is an flattened version of the dataset where multi-turn samples were split into multiple dataset lines. If you would like the original unflattened version before splitting samples into multiple lines please use UltraSteer-v0.
UltraSteer: A Massive Collection of Multi-Turn Dialogue with Fine-Grained Labels
UltraSteer is a large-scale dataset of single- and multi-turn dialogue with fine-grained labels produced by Nvidia's Llama2-13B-SteerLM-RM reward model using the NeMo… See the full description on the dataset page: https://huggingface.co/datasets/Avelina/UltraSteer-v0-flat.
