datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/b-mc2/sql-create-context.synthetic_text_to_sql
Image generated by DALL-E. See prompt for more details
synthetic_text_to_sql
gretelai/synthetic_text_to_sql is a rich dataset of high quality synthetic Text-to-SQL samples,
designed and generated using Gretel Navigator, and released under Apache 2.0.
Please see our release blogpost for more details.
The dataset includes:
105,851 records partitioned into 100,000 train and 5,851 test records
~23M total tokens, including ~12M SQL tokens
Coverage across 100 distinct… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_text_to_sql.SQaLe-text-to-SQL-dataset
🧮 SQALE: A Large-Scale Semi-Synthetic Dataset
SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas.
It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions.
The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.gretel-synthetic-text-to-sql
Fork of gretelai/synthetic_text_to_sql
The gretelai/synthetic_text_to_sql dataset is a large, Apache 2.0 licensed, synthetic Text-to-SQL dataset consisting of 105,851 high-quality records across 100 diverse domains, designed for training language models. It includes comprehensive SQL tasks with varying complexities, database contexts, natural language explanations, and contextual tags, outperforming existing datasets in SQL correctness and standards compliance.
spider-text-to-sql
Spider Text-to-SQL with LLM-Judge Labels
This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4.
Files
File
Description
spider_dataset.parquet
Full dataset with predictions and labels
scripts/
Reproduction scripts (see below)
Dataset statistics
Source: Spider 1.0 training split (train_spider.json)
Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.Text-to-sql-v1task076_splash_correcting_sql_mistake
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task076_splash_correcting_sql_mistake
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task076_splash_correcting_sql_mistake.stackoverflow-survey-2023-text-sql
BIQA Text-to-SQL Dataset
The data is from the Stack Overflow Developer Survey 2023.
Created with this Notebook; uses this spreadsheet defining manual adjustments.
data/eval_set_multi_answers_res.json: Question and query pairs as list of SQLSamples with possibly more than one valid SQL for a question. Also results included.
data/survey_results_normalized_v2.db: The main sqlite db file.
The json file contains a list of SQLSample objects as defined:
@dataclass
class SQLQuery:… See the full description on the dataset page: https://huggingface.co/datasets/deepset/stackoverflow-survey-2023-text-sql.omnimcp_sql_bigquery_analytics_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_bigquery_analytics_teaser.omnimcp_python_sqlalchemy_orm_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_python_sqlalchemy_orm_teaser.omnimcp_sql_postgres_pro_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_postgres_pro_teaser.omnimcp_sql_snowflake_warehouse_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_snowflake_warehouse_teaser.omnimcp_sql_dbt_transformations_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_sql_dbt_transformations_teaser.Persian-Business-Text-to-SQL-Gold-1K
Persian Business Text-to-SQL Gold-1K
1,000 Persian-native, execution-verified business Text-to-SQL examples for fine-tuning and benchmarking.
مجموعهای ۱۰۰۰ نمونهای برای تبدیل درخواستهای فارسی کسبوکار به SQL، همراه با دیتابیسهای SQLite اجرایی، schema کامل، متادیتای سختی/مهارت و ارزیابی مبتنی بر Execution Accuracy.
Motivation
BIRD emphasizes database-grounded Text-to-SQL and execution accuracy; Spider 2.0 pushes toward realistic enterprise database workflows.… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/Persian-Business-Text-to-SQL-Gold-1K.mirror-sql
MIRROR-SQL
Provenance-Controlled Database Environments for Text-to-SQL Agents.
13 PostgreSQL environments · 176 tables · 2762 columns · 390 annotated question/SQL pairs.
MIRROR-SQL takes the opposite approach to contamination from every other text-to-SQL corpus.
Spider and BIRD sample public databases. BEAVER uses real private warehouses that cannot be
redistributed. LiveSQLBench out-runs leakage temporally by rebuilding from changing sources.
MIRROR-SQL instead purpose-builds… See the full description on the dataset page: https://huggingface.co/datasets/1digitaldesign/mirror-sql.task077_splash_explanation_to_sql
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task077_splash_explanation_to_sql
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task077_splash_explanation_to_sql.SQLRobustBench
SQLRobustBench
SQLRobustBench is a synthetic benchmark for SQL robustness under explicit schema, parsing, and logical validation constraints.
Dataset Summary
This release focuses on two benchmark families:
SQLCorrupt: invalid SQL detection and repair
SQLNormalize: deterministic SQL canonicalization and normalization
SQLRobustBench evaluates model behavior on SQL tasks that require more than straightforward generation:
generating valid clean SQL under schema constraints… See the full description on the dataset page: https://huggingface.co/datasets/sigdelakshey/SQLRobustBench.text-to-sql-spider-dataset
Text-to-SQL Dataset
A curated dataset for training text-to-SQL models. This dataset contains natural language questions paired with corresponding SQL queries, formatted for instruction fine-tuning.
📊 Dataset Summary
Total Samples: 20000
Format: Chat template (system/user/assistant messages)
Task: Text-to-SQL generation
Language: English
License: apache-2.0
📁 Dataset Structure
Data Format
Each example contains a conversation with three roles:… See the full description on the dataset page: https://huggingface.co/datasets/chrisjcc/text-to-sql-spider-dataset.enterprise-text-to-sql-benchmark
Enterprise Text-to-SQL Benchmark
3,045 natural-language questions paired with executable PostgreSQL, over a
12-table enterprise schema (sales, catalogue, logistics, HR).
Built to answer one question honestly: does fine-tuning actually improve
text-to-SQL? On this benchmark, a QLoRA fine-tune of Qwen3-8B took strict
execution accuracy from 10.82 % to 50.99 % — and the benchmark is designed
so that number cannot be inflated by leakage or by string-matching.
split
rows… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/enterprise-text-to-sql-benchmark.Danbooru2021-SQLite
Danbooru 2021 SQLite
Dataset Summary
This is the metadata of danbooru 2021 dataset in SQLite format.
https://gwern.net/danbooru2021
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/Danbooru2021-SQLite.task107_splash_question_to_sql
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task107_splash_question_to_sql
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task107_splash_question_to_sql.ACE-SQL
ACE-SQL Training Data
This repository contains the curated supervised fine-tuning (SFT),
reinforcement learning (RL), and empirical-pool data released with
ACE-SQL: Adaptive Co-Optimization via Empirical Credit Assignment for
Text-to-SQL.
ACE-SQL trains a shared language-model policy in two roles: a schema retriever
that selects the minimum required database columns, and a SQL generator that
operates on the resulting pruned schema. The SFT data provides a cold start for
both… See the full description on the dataset page: https://huggingface.co/datasets/xiaobing11/ACE-SQL.effi-sql-training
Diff-SQL Training Dataset
We release the training datasets used by Diff-SQL for SQL efficiency optimization.
This dataset includes:
Patch Generator Training Dataset: SFT data for generating SQL optimization patches.
Constraint Aligner Training Dataset: SFT warmup data for constraint-aware SQL optimization refinement.
Files
patch-generator-training-dataset/
train.parquet
dev.parquet
constraint-aligner-training-dataset/
train.parquet
dev.parquet… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/effi-sql-training.SQLFlow
Text2SQL-Flow Dataset Repository
This repository contains the SQLFlow dataset.
The SQLFlow dataset is a large-scale, high-quality collection of semantically valid and structurally diverse Text-to-SQL examples, generated using a comprehensive SQL-aware data augmentation framework.
For more details, please visit the GitHub repository:🔗 https://github.com/TechNomad-ds/Text2SQL-Flow
fhirsql-reasoning-sql
FHIR-to-SQL with Database-Resolved Clinical Terminology
Natural-language hospital questions paired with a structured JSON query plan and compiled DuckDB
SQL, over a FHIR-derived clinical schema. Built entirely from synthetic Synthea patients — no
real patient data.
Companion to the paper Plan-Then-Compile: Turning a General-Purpose Coder Model into a FHIR Data
Analyst.
Code & paper: https://github.com/adelelsayed/fhirsql-reasoning-sql
Adapters:… See the full description on the dataset page: https://huggingface.co/datasets/adelelsayed1991/fhirsql-reasoning-sql.sql-create-context-instruction
Overview
This dataset is built upon SQL Create Context, which in turn was constructed using data from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-SQL LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-SQL datasets. The CREATE TABLE statement can often be… See the full description on the dataset page: https://huggingface.co/datasets/bugdaryan/sql-create-context-instruction.sql-create-context-copy
Fork of b-mc2/sql-create-context
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from… See the full description on the dataset page: https://huggingface.co/datasets/philschmid/sql-create-context-copy.sql-create-context-pt
Overview
Este dataset é uma versão traduzida para o português do dataset b-mc2/sql-create-context,
que foi construído a partir dos datasets WikiSQL e Spider. Ele contém exemplos de perguntas
em português, instruções SQL CREATE TABLE e consultas SQL que respondem às perguntas
utilizando a instrução CREATE TABLE como contexto.
O principal objetivo deste dataset é ajudar modelos de linguagem natural em português a gerar consultas
SQL precisas e contextualizadas, prevenindo a… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/sql-create-context-pt.Effi-SQL
Effi-SQL
Update 2026-06-12
We release Effi-SQL, a dataset suite for SQL efficiency optimization.
This collection includes:
Effi-SQL Benchmark: a benchmark for evaluating SQL efficiency optimization methods.
Diff-SQL Training Dataset: training data used by Diff-SQL, including data for the Patch Generator and Constraint Aligner.
Dataset Fields
Effi-SQL Benchmark
id: A unique identifier for each benchmark instance.
db: The database… See the full description on the dataset page: https://huggingface.co/datasets/birdsql/Effi-SQL.scm-sql
SCM-SQL — a supply-chain natural-language-to-SQL evaluation set
500 (question, gold SQL) pairs authored against the live Odoo 17 supply-chain
schema, spanning 6 explicit complexity levels including multi-turn dialogues.
Built for the dissertation Domain-Aware Multi-Agent Natural-Language-to-SQL for
Enterprise Supply Chain Intelligence by Aniruddha Prakash Kawarase (BITS Pilani
WILP, 2026). Released as a public evaluation benchmark so other researchers can
compare domain-aware… See the full description on the dataset page: https://huggingface.co/datasets/AniruddhaAI/scm-sql.
