datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
synthetic-text2sqlEmploying the MTEB evaluation framework's dataset version, utilize the code below for assessment:
import mteb
import logging
from sentence_transformers import SentenceTransformer
from mteb import MTEB
logger = logging.getLogger(__name__)
model_name = 'intfloat/e5-base-v2'
model = SentenceTransformer(model_name)
tasks = mteb.get_tasks(
tasks=[
"AppsRetrieval",
"CodeFeedbackMT",
"CodeFeedbackST",
"CodeTransOceanContest",
"CodeTransOceanDL"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/synthetic-text2sql.ko_text2sqlRBAC-Text2SQL-Benchmark
RBAC-Text2SQL Benchmark
Role-conditioned Text-to-SQL instances for evaluating whether LLMs generate SQL that
respects Role-Based Access Control (RBAC) constraints. Each instance pairs a natural
language question with a role policy; the model must either produce a correct SQL query
that touches only authorized resources, or refuse with Sorry, I cannot answer.
Code, evaluation harness, and reproduction instructions:
https://github.com/2020dfff/RBAC-Text2SQL-Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/sharkiefff/RBAC-Text2SQL-Benchmark.spider-text2sql
SPIDER Text-to-SQL — Easy Access Version
A clean, HuggingFace-native version of the SPIDER Text-to-SQL benchmark. The original SPIDER dataset requires manually downloading a ZIP file from the Spider website. This version makes it instantly accessible via load_dataset.
What's Included
Each row contains the question, gold SQL, the database identifier, and a pre-parsed compact schema string — everything needed to train or evaluate a Text-to-SQL model without any additional… See the full description on the dataset page: https://huggingface.co/datasets/SuperMax991/spider-text2sql.text2sql-dataset
Dataset
We built this dataset from several sources combining examples from:
Wikisql
Bird
Spider
Synthetic SQL samples
This dataset has been cleaned and filtered by:
Removing DDL/DML examples (INSERT, UPDATE, DELETE, etc.)
De-duplicating examples based on hashing semantics of SQL and queries
Filtering only SELECT-style analytical queries
large-schema-text2sql-20k
Large-Schema Text-to-SQL (20K)
20,020 text-to-SQL examples whose median database schema has 95 tables.
Most text-to-SQL corpora hand the model a toy database. Spider averages about 5 tables per
database; BIRD is in the same range. Real analytics work does not look like that — it looks
like an ERP schema with 200 tables, 300 foreign keys, and eleven things called *_log, where
the hard part is not writing the JOIN but finding the two tables worth joining.
This dataset is that… See the full description on the dataset page: https://huggingface.co/datasets/VikramPal/large-schema-text2sql-20k.duckdb-text2sql-25k
Dataset Summary
The duckdb-text2sql-25k dataset contains 25,000 DuckDB text-2-sql pairs covering diverse aspects of DuckDB's SQL syntax.
We synthesized this dataset using Mixtral 8x7B, based on DuckDB's v0.9.2 documentation and Spider schemas that were translated to DuckDB syntax and enriched with nested type columns.
Each training sample consists of a natural language prompt, a corresponding (optional) schema, and a resulting query. Each pair furthermore has a category property… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-text2sql-25k.text2sql-oracle-postgres
Oracle / PostgreSQL text-to-SQL
Instruction data for fine-tuning google/gemma-3-270m-it (or any chat model) to emit a single dialect-correct SQL statement.
804 rows, 402 Oracle / 402 PostgreSQL
7 schemas: hr, sales, banking, inventory, tickets, university, logistics
Splits: 684 / 60 / 60 (grouped so paraphrases of the same SQL stay in one split)
Load
from datasets import load_dataset
ds = load_dataset("chabab/text2sql-oracle-postgres")
Record… See the full description on the dataset page: https://huggingface.co/datasets/chabab/text2sql-oracle-postgres.synthetic-text2sql-qrels
Dataset Card for "synthetic-text2sql-qrels"
More Information needed
text2sqlsynthetic-text2sql-dataset
Dataset Card for "synthetic-text2sql-dataset"
Dataset Summary
The synthetic-text2sql-dataset is a large-scale, structured dataset containing 100,000 training and 5,851 test examples designed to support research and development in SQL semantic parsing, text-to-SQL generation, and chain-of-thought (CoT) reasoning.
It was derived from an original DataFrame and converted into Hugging Face's datasets.Dataset format. Three new fields were added:
question: alias for the… See the full description on the dataset page: https://huggingface.co/datasets/eagle0504/synthetic-text2sql-dataset.wikisql-text2sql
WikiSQL Text-to-SQL (execution-ready)
A cleaned, execution-ready repackaging of WikiSQL
for text-to-SQL fine-tuning and execution-accuracy evaluation. Each example pairs a
natural-language question with a gold SQL query over a single-table schema — and every split
ships a real SQLite database so predicted SQL can be run and compared by result set
(not string-matched).
Splits
Split
Examples
train
55,339
dev
8,263
test
15,519
Columns… See the full description on the dataset page: https://huggingface.co/datasets/mohamed-ahmed-58059/wikisql-text2sql.chichewa-text2sql
Chichewa Text-to-SQL
The first structured Text-to-SQL benchmark for Chichewa, a low-resource Bantu language spoken by over 12 million people in Malawi and neighboring regions.
The dataset contains 400 manually curated natural language–SQL pairs in both Chichewa (Nyanja) and English, grounded in a unified relational SQLite database covering five real-world domains from Malawi.
Dataset Summary
This benchmark was constructed to investigate the adaptation of Large… See the full description on the dataset page: https://huggingface.co/datasets/johneze/chichewa-text2sql.synthetic-text2sql-queries-corpusEmploying the CoIR evaluation framework's dataset version, utilize the code below for assessment:
import coir
from coir.data_loader import get_tasks
from coir.evaluation import COIR
from coir.models import YourCustomDEModel
model_name = "intfloat/e5-base-v2"
# Load the model
model = YourCustomDEModel(model_name=model_name)
# Get tasks
#all task ["codetrans-dl","stackoverflow-qa","apps","codefeedback-mt","codefeedback-st","codetrans-contest","synthetic-
# text2sql","cosqa","codesearchnet"… See the full description on the dataset page: https://huggingface.co/datasets/CoIR-Retrieval/synthetic-text2sql-queries-corpus.bird-text2sql-bench
Dataset Card for bird-text2sql-bench
bird-text2sql-bench 是 BIRD(BIg Bench for Large-Scale Database Grounded Text-to-SQL) 官方訓練集之 OpenAI Messages 格式版本,共 9,428 筆。相較於 Spider 1.0,BIRD 使用真實大型資料庫(70 個,涵蓋電商、運動、教育、醫療等 37+ 領域),並提供 evidence(數值提示)欄位,本資料集將 evidence 以 ### Hint 段落併入 user prompt,形成可直接餵入 SFT pipeline 之 system / user / assistant 三 role 對話。除原生之 messages 欄位外,另拆解出獨立之 system / user / assistant 字串欄位,可同時作為 SFT 語料與 benchmark evaluation pipeline 之直接輸入。
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/bird-text2sql-bench.200k-Text2SQLtext2sql-dataset
Overview
This dataset builds from WikiSQL and Spider.
There are 78,577 examples of natural language queries, SQL CREATE TABLE statements, and SQL Query answering the question using the CREATE statement as context. This dataset was built with text-to-sql LLMs in mind, intending to prevent hallucination of column and table names often seen when trained on text-to-sql datasets. The CREATE TABLE statement can often be copy and pasted from different DBMS and provides table names, column… See the full description on the dataset page: https://huggingface.co/datasets/dipanjanS/text2sql-dataset.Text2SQLtext-2-sql-with-context
Dataset Card for "text-2-sql-with-context"
This dataset is prepared in Alpaca format introduced by Stanford to train LLMs. This dataset has been used in fine-tuning Chat Llama-2 7B. For more information, Please visit : Huggingface.
text2sql-spider
Dataset Card for "text2sql-spider-processed"
More Information needed
synthetic-text2sqlfor_text2sql_wikiSQL_korean_by_google_translator_apitext2sql-wikisql-spider
Dataset Card for "text2sql-wikisql-spider"
More Information needed
sft_text2sqlThis is the SFT training dataset for FINER-SQL on BIRD dataset. We use different LLMs from the LLM pool to generate diversed reasoning styles and diversed SQL styles.
The model pool includes: GPT-OSS-120b, Qwen-2.5-72B-Instruct, Deepseek-R1, GPT-5 (reasoning_effort=low)
Note that:
GPT-5 could be replaced by GPT-4o or GPT-4.1, because OpenAI doesn't return their true reasoning process so generating by GPT-5 was not necessary.
The schema filtering (top-30 columns) was applied before generating… See the full description on the dataset page: https://huggingface.co/datasets/griffith-bigdata/sft_text2sql.text2sql_challegeyasserrmd__Text2SQL-1.5B-details
Dataset Card for Evaluation run of yasserrmd/Text2SQL-1.5B
Dataset automatically created during the evaluation run of model yasserrmd/Text2SQL-1.5B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/yasserrmd__Text2SQL-1.5B-details.spider-text2sql-bench
Dataset Card for spider-text2sql-bench
spider-text2sql-bench 是 Spider 1.0 官方訓練集之 OpenAI Messages 格式版本,共 7,000 筆,將原始之 question / schema / sql 重新組裝為 system / user / assistant 三 role 之對話結構。除原生之 messages 欄位外,另拆解出獨立之 system / user / assistant 字串欄位,可作為 Text-to-SQL 模型之 SFT 訓練語料,亦可直接用於 benchmark evaluation pipeline(以 user 作為 prompt,比對模型輸出與 assistant 之標準答案 SQL)。
Dataset Details
Dataset Description
Spider 1.0 為 Yale LILY Group 於 EMNLP 2018 發表之大規模跨領域 Text-to-SQL… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/spider-text2sql-bench.Text2SQLText2SQL_Workflow_Trace
Text2SQL Workflow Trace
Dataset Description
This dataset contains workflow traces for Text-to-SQL tasks, capturing the intermediate steps of translating natural language queries to executable SQL. It was used as input trace for the research presented in the paper:"HEXGEN-TEXT2SQL: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQL Workflow" (arXiv:2505.05286).
The end-to-end Text-to-SQL queries collected in the dataset are from BIRD bench, and the trace… See the full description on the dataset page: https://huggingface.co/datasets/fredpeng/Text2SQL_Workflow_Trace.text2sql-canonical-v2
