datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
simverse2026
SimVerse
⚠️ Anonymized for double-blind review. This dataset is currently undergoing peer review. It is hosted under an anonymous account dedicated to the review process; the author and citation fields are deliberately unfilled. Permanent ownership and citation information will be added after the review concludes. Please do not attempt to deanonymize the maintainers of this dataset during review.
A multi-task benchmark for evaluating multimodal LLMs on interactive simulation… See the full description on the dataset page: https://huggingface.co/datasets/SimVer-ano/simverse2026.agent-simulations
Agent Simulations
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
53,971 synthetic agent trajectories generated by simulations
across 34 agent types. The rows include successful and failed
trajectories for supervised fine-tuning, preference work, reinforcement learning, and
evaluation.
NOTE: This is generated test and training data, not curated ground truth. Review and
filter it for your application before training or… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/agent-simulations.simutrade-rag-sft-28k
📢 Domain & Email Migration Notice
From May 6th, 2026, Simutrade will transition to new domains as simutrade.app will not be renewed:
🌐 Website: simutrade.faizath.com (formerly simutrade.app)
⚙️ API: simutrade-api.faizath.com (formerly api.simutrade.app)
📧 Email: contact@simutrade.faizath.com (formerly contact@simutrade.app)
🛰️ CDN: simutrade-cdn.faizath.com (formerly cdn.simutrade.app)
📈 Status Pages:… See the full description on the dataset page: https://huggingface.co/datasets/simutrade/simutrade-rag-sft-28k.tau2-simulated
tau2 Simulated Training Set
Made with the whileai SDK · Collections: Simulation, Start here: foundational post-training datasets
The training set that took a base model from 5% to 30% on tau2-bench
telecom, made from nothing but the agent's tool list and policy.
If you build a customer-facing agent, you already have the two files this
dataset was made from: the tools it can call and the policy it follows.
The whileai SDK turned those into 1,057 graded conversations across the… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/tau2-simulated.VeriReason-RTL-Coder_7b_reasoning_tb_simple
Verireason-RTL-Coder_7b_reasoning_tb_simple
For implementation details, visit our GitHub repository: VeriReason and our page
Check out our paper: VeriReason: Reinforcement Learning with Testbench Feedback for Reasoning-Enhanced Verilog Generation
Update Log
2025.05.17: Initial release of Nellyw888/Verireason-RTL-Coder_7b_reasoning_tb_simple
Project Description
This study introduces VeriReason, a novel approach utilizing reinforcement learning with… See the full description on the dataset page: https://huggingface.co/datasets/Nellyw888/VeriReason-RTL-Coder_7b_reasoning_tb_simple.age-specific-text-simplification
Age-Specific Text Simplification Dataset
Dataset Description
This dataset contains complex texts simplified into age-appropriate versions for children aged 3, 4, and 5 years old. Each original text has been professionally adapted to match the cognitive development, vocabulary, and comprehension abilities of each specific age group.
Dataset Summary
Total Examples: 17,177
Training Split: 15,459 examples
Validation Split: 1,718 examples
Languages:… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/age-specific-text-simplification.simplemath-cot
🧮 SimpleMath-100k CoT
A chain-of-thought (CoT) extension of the
ProCreations/SimpleMath
dataset. Every one of the 100 000 algebra / arithmetic problems is paired with a
short, numbered reasoning trace (Step 1: … Step 2: …) that walks a language
model from the problem statement to the known-correct answer.
The traces in the Jupyter notebook are generated by
Qwen3.8-27B and then post-processed to strip formatting noise,
enforce sequential step numbering, and cap output at 1 000… See the full description on the dataset page: https://huggingface.co/datasets/alexfromapex/simplemath-cot.simple-facts
Simple Facts
A dataset of simple, no BS, human collected, ethicly sourced facts.
About 1000 examples.
This dataset is growing, and every day I plan to add a few more facts.
tw-finance-reasoning-instruct
tw-finance-reasoning-instruct
台灣金融知識的繁體中文推理指令資料集。每一題都有完整的思考過程(think)與可查證的答案(output)。
欄位規格對齊 twinkle-ai/tw-reasoning-instruct-50k。
⚠️ 使用限制:僅供研究,不得商業使用
本資料集以 CC BY-NC 4.0 授權釋出,僅供學術研究、模型能力探索與方法驗證之用。
不得作商業用途。 包含訓練用於對外營利的模型、包裝為付費產品或服務、
或作為商業交付物的一部分。若有商業需求,請自行重新建置資料並取得合規來源。
這不是財務、稅務、法律或投資建議。 資料中的稅率、費率、法規門檻依 2026 年
(民國 115 年)台灣公開資訊整理,但可能已經過時。任何實際決策前,
請以主管機關公告為準(財政部、金管會、勞動部、衛福部、中央銀行、全國法規資料庫)。
內容為程式化合成,非真實考題逐字收錄。 題目由計算器與模板生成,
並非任何證照考試的原始試題。… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/tw-finance-reasoning-instruct.nemotron-cc-atomic-simplification-gemma4-31b
nemotron-cc atomic-statement simplification (Gemma 4 31B-it)
2,000,000 records: source text from
nvidia/nemotron-cc-v2.1 (High-Quality-Synthetic
split) rewritten by google/gemma-4-31B-it into a sequence of atomic, Subject-Verb-Object
statements.
Generated with vLLM 0.22.1 in-process batch inference (see src/generate/run.py in the
producing repo), TP=4, max_model_len=16384, max_tokens=8192, prompts filtered to
<=8192 templated tokens.
Fields
id: original… See the full description on the dataset page: https://huggingface.co/datasets/rpisano/nemotron-cc-atomic-simplification-gemma4-31b.tw-finance-function-call-reasoning
tw-finance-function-call-reasoning
台灣金融場景的繁體中文 function-calling + 推理鏈微調資料集。
欄位規格對齊 twinkle-ai/tw-function-call-reasoning-10k。
⚠️ 使用限制:僅供研究,不得商業使用
本資料集以 CC BY-NC 4.0 授權釋出,僅供學術研究、模型能力探索與方法驗證之用。
請務必理解以下事項後再使用:
不得作商業用途。 包含但不限於:訓練用於對外營利的模型、包裝為付費產品或服務、
作為商業交付物的一部分。若有商業需求,請自行重新建置資料並取得合規來源。
這不是財務、稅務、法律或投資建議。 資料中的稅率、費率、法規門檻雖依 2026 年
(民國 115 年)台灣公開資訊整理,但可能已經過時或有誤。任何實際決策前,
請以主管機關公告為準(財政部、金管會、勞動部、衛福部、中央銀行、全國法規資料庫)。
內容為程式化合成,非真實考題逐字收錄。 題目由模板與參數取樣組合而成,… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/tw-finance-function-call-reasoning.ACG-SimpleQA
ACG-SimpleQA
🌐 Website •
🤗 Hugging Face
中文 | English
ACG-SimpleQA is an objective knowledge question-answering dataset focused on the Chinese ACG (Animation, Comic, Game) domain, containing 4242 auto-generated carefully designed QA samples. This benchmark aims to evaluate large language models' factual capabilities in the ACG culture domain, featuring Chinese language, diversity, high quality, static answers, and easy evaluation.
📢 Latest Updates… See the full description on the dataset page: https://huggingface.co/datasets/Papersnake/ACG-SimpleQA.kubectl-mcp-server-tool-call-reasoning-6k
kubectl-mcp-server-tool-call-reasoning-6k
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):install_helm_chart, upgrade_helm_chart, uninstall_helm_chart, helm_list, helm_status, helm_history, helm_get_values, helm_get_manifest, helm_get_notes, helm_get_hooks, helm_get_all, helm_show_chart, helm_show_values, helm_show_readme, helm_show_crds, helm_show_all, helm_search_repo, helm_search_hub, helm_repo_list… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/kubectl-mcp-server-tool-call-reasoning-6k.simple-llm-sft
Simple LLM SFT Dataset
This synthetic dataset contains 1,000 English prompt-response pairs for
supervised fine-tuning. It was created to fine-tune
Qwen/Qwen3.5-4B to give clear,
direct, and technically correct answers in simple English.
The writing guidance is inspired by ASD-STE100 Simplified Technical English.
The dataset does not claim official ASD-STE100 compliance or certification.
Dataset structure
The default configuration contains:
Split
Examples… See the full description on the dataset page: https://huggingface.co/datasets/thisisandreeeee/simple-llm-sft.Leesplank_NL_wikipedia_simplificationsThe set contains 2.87M pragraphs of prompt/result combinations, where the prompt is a paragraph from Dutch Wikipedia and the result is a simplified text, which could include more than one paragraph.
This dataset was created by UWV, as a part of project "Leesplank", an effort to generate datasets that are ethically and legally sound.
The basis of this dataset was the wikipedia extract as a part of Gigacorpus (http://gigacorpus.nl/). The lines were fed one by one into GPT 4 1106 preview, where… See the full description on the dataset page: https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications.simple_bench
📊 Simple Bench Dataset
A Compact Benchmark for Structured Reasoning and Multiple-Choice Evaluation in Large Language Models
Simple Bench Dataset is a structured evaluation collection derived from the Simple Bench benchmark, designed to assess reasoning, comprehension, and multiple-choice question-answering capabilities of large language models through concise yet non-trivial problems that require logical inference rather than simple retrieval; each sample consists of a natural… See the full description on the dataset page: https://huggingface.co/datasets/sapiens-technology/simple_bench.simple-wikipedia
English Simple Wikipedia
This is just a copy of english simple Wikipedia dataset
that I converted to jsonl format for testing purpose when jsonl format is needed. Here is the
link to download the jsonl file.
visual_genome-simple-en
Dataset Card for Visual Genome Annotations in Simple English
This dataset contains captions that were rephrased into simple english so that a young child would understand it.
Dataset Details
Dataset Sources
The processed Visual Genome captions in this repo are based on the following sources:
941425b651f50cdb1a6f0673eaab6260 vg_caption.json (https://storage.googleapis.com/sfr-vision-language-research/LAVIS/datasets/visual_genome/vg_caption.json)
Visual… See the full description on the dataset page: https://huggingface.co/datasets/Jotschi/visual_genome-simple-en.Cygnis-Identity-SFT
Cygnis Identity Dataset
Présentation
Ce dépôt contient le jeu de données d'entraînement initial pour l'identité de l'intelligence artificielle Cygnis. Ce dataset est conçu pour le réglage fin (fine-tuning) supervisé afin d'établir les fondements comportementaux et l'identité de l'agent.
Spécifications du Dataset
Le jeu de données est composé de paires d'instructions visant à définir l'origine, le concepteur et la nature du système.
Format : JSONL / Hugging… See the full description on the dataset page: https://huggingface.co/datasets/Simonc-44/Cygnis-Identity-SFT.simple-python-grpo
simple-python-grpo
A curated set of simple Python function problems for GRPO / RLVR fine-tuning.
Each row has a natural-language description, a function signature, and 3
auto-verified test assertions (generated by running a reference implementation,
so every test is correct by construction). The reference is NOT included — the
model must generate the body and is rewarded when the tests pass.
Fields: name, prompt, func_prompt, tests (newline-separated asserts),
setup_code.
Built… See the full description on the dataset page: https://huggingface.co/datasets/sagecodes/simple-python-grpo.Simple-agent-traces
📱 Simple Agent Traces – Tiny Tool‑Calling Conversations for Small Models
Simple Agent Traces is a compact, hand‑picked dataset of 605 real‑world tool‑calling conversations, each carefully truncated to ≤8,192 tokens (using the SmolLM2‑360M tokenizer).It is purpose‑built for training and fine‑tuning tiny language models (≤500M) that must run on‑device – smartphones, edge devices, or any environment with strict memory and latency constraints.
🧹 No chain‑of‑thought, no fluff.Every… See the full description on the dataset page: https://huggingface.co/datasets/LiteMind/Simple-agent-traces.Semantic_similarity_deduplicated_reasoning_data_english
Semantic_similarity_deduplicated_reasoning_data_english
数据集描述
Semantic similarity deduplicated reasoning data filtered from OpenThoughts2-1M, 77662 examples in total, 10000 examples for each category
文件结构
semantic_similarity_deduplicated_reasoning_data_english.jsonl: 主数据文件(JSONL格式)
数据格式
数据集包含以下字段:
question: str
quality: int
difficulty: int
topic: str
validity: int
使用方法
方法1: 使用datasets库
from datasets import load_dataset
#… See the full description on the dataset page: https://huggingface.co/datasets/Ibisbill/Semantic_similarity_deduplicated_reasoning_data_english.SIMXP-26052026-METASYN001
SIMXP-26052026-METASYN001
Multi-Omics Agent Memory Simulation — Metabolic Syndrome TCA Cycle Biomarker Study
This dataset supports the experiment described in the article "Does Your Research Agent Remember? Six Months of Multi-Omics Team Knowledge vs. None — A Controlled Comparison" and demonstrates the etchmem memory system for autonomous AI research agents.
It contains the full event log, synthesized knowledge export, and fine-tuning pairs from a simulated six-month plasma… See the full description on the dataset page: https://huggingface.co/datasets/simulatexp/SIMXP-26052026-METASYN001.countdown-qwen3-0.6b
Countdown Qwen3-0.6B Pass@10 Buckets
Countdown arithmetic problems filtered by observed local Qwen/Qwen3-0.6B success rate over 10 rollouts per problem.
Each problem asks for an arithmetic expression that reaches a target using each listed source number at most once. The final answer should be inside \boxed{...}. Canonical solutions are provided, but any verifier-valid expression is accepted.
Subsets
subset
source bucket
count
observed successes out of 10… See the full description on the dataset page: https://huggingface.co/datasets/simpissa/countdown-qwen3-0.6b.twinkle_hub_finetune_dataset
twinkle_hub_finetune_dataset
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):search_datasets, get_dataset, query_rows, materialize_dataset, search_patents, get_patent_body, search_exam, search_exam_questions, get_exam_paper, search_teacher_exam, search_teacher_exam_questions, get_teacher_exam_paper, search_teacher_recruit, search_teacher_recruit_questions, get_teacher_recruit_paper, search_taiwan_md… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/twinkle_hub_finetune_dataset.haddas-instruct-ti
haddas-instruct-ti
Instruction-tuning pairs synthesized from real Eritrean newspaper articles: summarize, translate, classify topic, extract keywords. Inputs are Tigrinya, outputs are English (or topic label).
Source
Derived from the Haddas Eritrea newspaper archive: 63 PDF issues processed
by the haddas-eritrea pipeline (extract -> clean -> segment -> translate -> label).
Generated: 2026-04-26 12:21 UTC
Row count: 6720
Schema: id, task, instruction, input, output, topic… See the full description on the dataset page: https://huggingface.co/datasets/SIMBA9657/haddas-instruct-ti.medical-reports-simplification-dataset
🏥 Medical Reports Simplification Dataset
📋 Description
Dataset créé avec Gemini 2.5 Pro (Preview) pour entraîner des modèles à simplifier les rapports médicaux complexes en explications compréhensibles pour les patients.
🎯 Objectif : Démocratiser l'accès à l'information médicale en rendant les rapports techniques accessibles au grand public.
🔧 Génération du Dataset
Génération : Gemini 2.5 Pro (Preview)
Validation : Contrôle qualité automatisé… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medical-reports-simplification-dataset.mn_business_benchmark_dataset_simple
mn_business_benchmark_dataset_2000_diverse
Монгол хэл дээрх бизнес, санхүү, борлуулалт, маркетинг, unit economics-ийн 2000 мөртэй синтетик benchmark dataset.
Энэ хувилбар нь блок бүрт нэг тоо л өөрчлөгдөх маягийн жишээнээс зайлсхийж, seed-тэй random generation, олон төрлийн өгүүлбэрийн загвар, олон бизнесийн domain, 25+ topic ашигласан.
Schema
id: 1-ээс 2000 хүртэлх дараалсан дугаар
instruction: бизнесийн бодлогын өгүүлбэр
input: хоосон string
thinking: бодолт… See the full description on the dataset page: https://huggingface.co/datasets/joppari/mn_business_benchmark_dataset_simple.Simon_Trove
SimsonTrove – Racing-Planet Agentic Traces
Synthetischer ReAct-Datensatz (Reason + Act) zur Zweitakt-Tuningberatung auf Basis
des Racing-Planet.de Sortiments. Jeder Trace bildet die ingenieurmäßige Entscheidungs-
kette eines Tuning-Mechanikers ab: Beobachtung → Hypothese → Action (Teilewahl) →
Feedback → Iteration → Final Response.
Schnellüberblick
Property
Value
Anzahl Traces
300
Sprache
Deutsch (technisch)
Format
JSON (Liste von Trace-Objekten)
Avg.… See the full description on the dataset page: https://huggingface.co/datasets/jmp1987/Simon_Trove.simple-text-generation-basic
Simple Text Generation Basic Dataset
This dataset contains very simple text samples designed for testing and basic text generation tasks.
Dataset Structure
Each row contains a single field:
text: a plain English sentence
Example
{"text": "Artificial intelligence is transforming the world."}
