datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dbbench-distilled-qwen3-14b-multiturnopenalex_dbbench_synth_v5
OpenAlex-Inspired Synthetic SQL Agent Dataset
(MySQL/MariaDB, Schema-Aware, Teacher-Guided)
This dataset contains fully synthetic multi-turn SQL agent trajectories
generated over an OpenAlex-inspired relational schema.
It is designed to improve SQL-agent performance in structured,
tool-driven environments such as SQL-agent benchmarks
(e.g., AgentBench-style database tasks).
✅ No real OpenAlex data is included.
All schema definitions and rows are programmatically generated synthetic… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/openalex_dbbench_synth_v5.openalex_dbbench_synth_v6
OpenAlex-Inspired Synthetic SQL Agent Dataset
(MySQL/MariaDB, Schema-Aware, Teacher-Guided)
This dataset contains fully synthetic multi-turn SQL agent trajectories
generated over an OpenAlex-inspired relational schema.
It is designed to improve SQL-agent performance in structured,
tool-driven environments such as SQL-agent benchmarks
(e.g., AgentBench-style database tasks).
✅ No real OpenAlex data is included.
All schema definitions and rows are programmatically generated synthetic… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/openalex_dbbench_synth_v6.dbbench-conversation-datasetopenalex_dbbench_synth_v2
OpenAlex-inspired Synthetic SQL Agent Dataset (SQLite, Teacher-Guided)
This dataset contains synthetic multi-turn agent trajectories
generated over an OpenAlex-inspired SQLite schema.
It is intended to improve model performance on
database reasoning benchmark tasks (e.g., DBBench),
especially aggregation-heavy and tool-driven SQL.
This version uses a teacher LLM to generate natural language
questions and reasoning thoughts, while SQL execution is fully verified.
Key… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/openalex_dbbench_synth_v2.dbbench_and_alfworld_sft_dataset_v4
DBBench + ALFWorld SFT Dataset (Merged)
Overview
This dataset is a simple concatenation (merge) of the following two synthetic SFT datasets:
ALFWorld Trajectory Dataset: moroqq/sft_alfworld_trajectory_dataset_v2
https://huggingface.co/datasets/moroqq/sft_alfworld_trajectory_dataset_v2
DBBench SFT Dataset (ReAct Format): u-10bei/dbbench_sft_dataset_react_v4
https://huggingface.co/datasets/u-10bei/dbbench_sft_dataset_react_v4
The goal is to provide a single dataset… See the full description on the dataset page: https://huggingface.co/datasets/moroqq/dbbench_and_alfworld_sft_dataset_v4.dbbench_v4_plus_alfworld_v5_mixeddbbench_sft_dataset_react_augmented_3dbbench_cleaned_for_agentbench
DBBench Cleaned for AgentBench
u-10bei/dbbench_sft_dataset_react_v4(1,200 件)に対してクレンジング処理を施したデータセット。
AgentBench DBBench 評価用の SFT 訓練データとしてそのまま使用可能。
混合利用を想定: 本データセットは mark-22/dbbench-spider-3500(1,697 件)と混合し、合計 2,897 件 の SFT データとして使用することを想定しています。
Dataset Summary
Metric
Value
Total rows
1,200
Source
u-10bei/dbbench_sft_dataset_react_v4
Avg messages per item
6.7
Items with Final Answer1,200 / 1,200 (100%)
Columns
id, messages, metadata… See the full description on the dataset page: https://huggingface.co/datasets/mark-22/dbbench_cleaned_for_agentbench.dbbench_u-10bei_sft_dataset_modified_v1
DBBench SFT Dataset Modified v1
Overview
This dataset is a modified version of
u-10bei/dbbench_sft_dataset_react_v4.
Modifications
Original dataset: 1,200 samples
Of the ~380 INSERT-type samples, approximately 252 (66%) were replaced with
paraphrased versions using implicit expressions
Total: ~1,200 samples (same size as original)
The replaced samples are paraphrased versions of INSERT-type questions,
generated using a whitelist model (Qwen3-14B) to improve… See the full description on the dataset page: https://huggingface.co/datasets/ShogoMu/dbbench_u-10bei_sft_dataset_modified_v1.openalex_dbbench_synth_v3
OpenAlex-Inspired Synthetic SQL Agent Dataset
(MySQL/MariaDB, Teacher-Guided)
This dataset contains fully synthetic multi-turn SQL agent trajectories
generated over an OpenAlex-inspired relational schema.
It is intended to help models learn tool-driven SQL reasoning skills
that are commonly evaluated in SQL-agent benchmarks (e.g., AgentBench DB tasks),
such as multi-step querying, aggregation, and database modifications.
✅ No real OpenAlex data is included.
All schema and rows are… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/openalex_dbbench_synth_v3.dbbench_and_alfworld_sft_dataset
DBBench + ALFWorld SFT Dataset (Merged)
Overview
This dataset is a simple concatenation (merge) of the following two synthetic SFT datasets:
ALFWorld Trajectory Dataset: u-10bei/sft_alfworld_trajectory_dataset_v5
https://huggingface.co/datasets/u-10bei/sft_alfworld_trajectory_dataset_v5
DBBench SFT Dataset (ReAct Format): u-10bei/dbbench_sft_dataset_react_v4
https://huggingface.co/datasets/u-10bei/dbbench_sft_dataset_react_v4
The goal is to provide a single dataset… See the full description on the dataset page: https://huggingface.co/datasets/moroqq/dbbench_and_alfworld_sft_dataset.openalex_dbbench_synth_v1
OpenAlex-inspired Synthetic SQL Agent Dataset (SQLite)
This dataset contains synthetic multi-turn SQL agent trajectories
generated over an OpenAlex-inspired SQLite schema.
It is intended for training models on database reasoning tasks,
with a focus on aggregation-heavy and tool-augmented SQL execution.
Acknowledgements
This work draws structural inspiration from public academic metadata systems:
OpenAlex (https://openalex.org)The relational schema design is conceptually… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/openalex_dbbench_synth_v1.dbbench_sft_dataset200
DBBench SFT Dataset 200 (Robustness-Heavy Subset)
This dataset is a 200-sample subset of u-10bei/dbbench_sft_dataset_react_v4.
Motivation
You want to mix a small amount of DBBench data into ALFWorld SFT to reduce DBBench regression.
This subset is biased toward retry / error-recovery trajectories to teach robustness.
Selection Policy (Summary)
Total: 200
Seed: 0
Max samples per metadata.table_name: 2
Target allocation (approx.):
{"core_insert": 40… See the full description on the dataset page: https://huggingface.co/datasets/moroqq/dbbench_sft_dataset200.openalex_dbbench_synth_v4
OpenAlex-Inspired Synthetic SQL Agent Dataset
(MySQL/MariaDB, Schema-Aware, Teacher-Guided)
This dataset contains fully synthetic multi-turn SQL agent trajectories
generated over an OpenAlex-inspired relational schema.
It is designed to improve SQL-agent performance in structured,
tool-driven environments such as SQL-agent benchmarks
(e.g., AgentBench-style database tasks).
✅ No real OpenAlex data is included.
All schema definitions and rows are programmatically generated synthetic… See the full description on the dataset page: https://huggingface.co/datasets/tussiiiii/openalex_dbbench_synth_v4.dbbench_sft_dataset_org_v3
dbbench_sft_dataset_org_v3
A DBBench-style SFT dataset combining 295 stratified examples from dbbench_u-10bei_sft_dataset_modified_v1 (100 from INSERT/INSERT_error_recovery + 195 from other types) and 5 handcrafted examples from dbbench_sft_dataset_org_v1.
Overview
Total: 300 conversations
100 randomly sampled from INSERT and INSERT_error_recovery types in modified_v1
195 randomly sampled from other types (UPDATE, aggregation, comparison, etc.) in modified_v1
5… See the full description on the dataset page: https://huggingface.co/datasets/ShogoMu/dbbench_sft_dataset_org_v3.dbbench_and_alfworld_sft_dataset_v3
DBBench + ALFWorld SFT Dataset (Merged)
Overview
This dataset is a simple concatenation (merge) of the following two synthetic SFT datasets:
ALFWorld Trajectory Dataset: moroqq/sft_alfworld_trajectory_dataset_v2
https://huggingface.co/datasets/moroqq/sft_alfworld_trajectory_dataset_v2
DBBench SFT Dataset (ReAct Format): moroqq/dbbench_sft_dataset200
https://huggingface.co/datasets/moroqq/dbbench_sft_dataset200
The goal is to provide a single dataset repo that… See the full description on the dataset page: https://huggingface.co/datasets/moroqq/dbbench_and_alfworld_sft_dataset_v3.dbbench_sft_dataset_react_v4_plus20_fix16_v2distilled_dbbench_dataset_2_cleanedDeepthinking-alfworld_and_dbbench_spider_v2
Deepthinking ALFWorld & DBBench Spider v2
AgentBench 評価の 2 タスク(ALFWorld / DBBench)を統合した マルチタスク SFT 訓練データセット。
フォーマット検査・フィルタリング済みの 7,779 件を、サイズ比率に基づく等間隔インターリーブで結合。
Dataset Summary
Metric
Value
Total rows
7,779
ALFWorld
4,884 (62.8%)
DBBench
2,895 (37.2%)
Avg messages per item
18.3
Columns
messages
Interleave method
比率ベース等間隔マージ
Source Datasets
Source
Rows
Description
mark-22/Deepthinking-sft_alfworld_final1
4,884
ALFWorld… See the full description on the dataset page: https://huggingface.co/datasets/mark-22/Deepthinking-alfworld_and_dbbench_spider_v2.dbbench_distilled_v1
dbbench_distilled_v1
概要
DB Bench(AgentBench)の弱点カテゴリに対して、教師モデルで高品質なエージェントトレースを再生成した蒸留データセット。
生成方法
手法: Oracle注入型マルチターン生成(ハード蒸留)
教師モデル: Qwen/Qwen3-Coder-30B-A3B-Instruct(BF16、A100 80GB)
ソースデータ: u-10bei/dbbench_sft_dataset_react_v4
生成環境: Google Colab A100 80GB(ハイメモリ)
Oracle注入型マルチターン生成とは
dbbench_v4 から弱点カテゴリのサンプルをフィルタ
各サンプルの user質問・tool応答(SQL実行結果)・正解を抽出
教師モデルに「正解ヒント付き」でエージェント応答を1ターンずつ生成させる
教師が SQL を書いた後、元データの tool応答(実際のSQL実行結果)を注入
これを繰り返し、効率的なトレースを再生成… See the full description on the dataset page: https://huggingface.co/datasets/kamaboko2007/dbbench_distilled_v1.dbbench_sft_dataset_react_v3v4_weaktypes
DBBench SFT Trajectories (Weak Types from v3–v4)
This dataset provides a subset of DBBench SFT trajectories focusing on
"weak" categories (i.e., categories where the current agent model shows
lower accuracy).
Source Datasets
The following datasets are used as sources:
u-10bei/dbbench_sft_dataset_react_v3
u-10bei/dbbench_sft_dataset_react_v4
From these datasets, only examples whose metadata.type contains one of
the following keywords are included:
counting
comparison… See the full description on the dataset page: https://huggingface.co/datasets/kuririrn/dbbench_sft_dataset_react_v3v4_weaktypes.dbbench_sft_dataset_react_v4_plus20_fix16_v3dbbench_v3_rlvmr_taggeddbbench_trajectories_teacher_v2
DBBench Trajectories (Teacher Generated v2)
A synthetically generated dataset of multi-turn agent trajectories for database question-answering tasks (DBBench),
intended for Supervised Fine-Tuning (SFT) and Knowledge Distillation of smaller agent models.
Teacher Model: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8
Source Dataset: u-10bei/dbbench_sft_dataset_react_v4
Format: Multi-turn conversational format (OpenAI/ChatML messages list with role and content)
Generation Objective… See the full description on the dataset page: https://huggingface.co/datasets/SELEE/dbbench_trajectories_teacher_v2.dbbench_v4_rlvmr_taggedlogiccat_dbbench_rlvmr_taggeddbbench_sft_dataset_react_v4_augmented_3dbbench_and_alfworld_sft_dataset_v2
DBBench + ALFWorld SFT Dataset (Merged)
Overview
This dataset is a simple concatenation (merge) of the following two synthetic SFT datasets:
ALFWorld Trajectory Dataset: moroqq/sft_alfworld_trajectory_dataset_v5_cleaned
https://huggingface.co/datasets/moroqq/sft_alfworld_trajectory_dataset_v5_cleaned
DBBench SFT Dataset (ReAct Format): moroqq/dbbench_sft_dataset200
https://huggingface.co/datasets/moroqq/dbbench_sft_dataset200
The goal is to provide a single dataset… See the full description on the dataset page: https://huggingface.co/datasets/moroqq/dbbench_and_alfworld_sft_dataset_v2.dbbench_u-10bei_sft_dataset_modified_v2
DBBench SFT Dataset Modified v2
Overview
This dataset is a modified version of
u-10bei/dbbench_sft_dataset_react_v4.
Acknowledgements
We would like to express our sincere gratitude to the creator of
u-10bei/dbbench_sft_dataset_react_v4
for providing a high-quality SFT dataset for DBBench.
This work would not have been possible without their foundational contribution.
Modifications
Base: u-10bei/dbbench_sft_dataset_react_v4 (1,200 samples)… See the full description on the dataset page: https://huggingface.co/datasets/ShogoMu/dbbench_u-10bei_sft_dataset_modified_v2.
