datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SQaLe-text-to-SQL-dataset
🧮 SQALE: A Large-Scale Semi-Synthetic Dataset
SQALE is a large-scale, semi-synthetic Text-to-SQL dataset grounded in real-world database schemas.
It was designed to push the boundaries of natural language to SQL generation, combining realistic schema diversity, complex query structures, and linguistically varied natural language questions.
The dataset was introduced in the paper SQaLe: A Large Text-to-SQL Corpus Grounded in Real Schemas. The code for the generation pipeline of this… See the full description on the dataset page: https://huggingface.co/datasets/trl-lab/SQaLe-text-to-SQL-dataset.text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.spider-text-to-sql
Spider Text-to-SQL with LLM-Judge Labels
This dataset extends Spider 1.0 with SQL predictions from gpt-5.4-mini and two correctness labels per example: a hybrid ground truth label and an LLM judge label from gpt-5.4.
Files
File
Description
spider_dataset.parquet
Full dataset with predictions and labels
scripts/
Reproduction scripts (see below)
Dataset statistics
Source: Spider 1.0 training split (train_spider.json)
Databases: the… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/spider-text-to-sql.exp01-eeg-to-text-sentences
Exp01 — Sentence-level EEG-to-text training data (unified)
This is a private working corpus for experiment 1 (fine-tuning EEG / time-series
foundation models on EEG-to-English-text). It bundles several public EEG-while-reading
datasets into a single, raw-lossless parquet schema where one row = one sentence read by
one participant.
⚠️ License: Per-source licenses are preserved verbatim in each row's license
column and source_url. Do not re-distribute publicly without re-checking the… See the full description on the dataset page: https://huggingface.co/datasets/tankalapavankalyan/exp01-eeg-to-text-sentences.text-to-art-database
Vieutopia T2A Privacy Train v1
Dataset Summary
Privacy-safe text-to-image dataset repacked into Parquet shards with embedded image bytes.
Scope: text-to-image outputs only
Excluded: image-to-image pipelines (pix2pix_*, pst_*)
Privacy: no raw task UUIDs, no user/device fields
Storage format: parquet shards (image as binary bytes), no image_path dependency
Splits
samples
train: 117572
validation: 6532
test: 6532
total: 130636
iterations… See the full description on the dataset page: https://huggingface.co/datasets/quchenyuan/text-to-art-database.text-to-sql-shop
Text-to-SQL on a seeded store schema, with checkpoints
Recipe: recipes/04-train/text-to-sql · Collection: Analyst
A question about an online store's database in, one PostgreSQL query out, graded by
a program: run the query, compare the result set to the gold query's result. The
schema (8 tables, seeded, schema.sql + seed.sql), the verifier, the trainer and
the benchmark runner are the
recipes/04-train/text-to-sql
recipe in the open-source whileai SDK.
Splits
|… See the full description on the dataset page: https://huggingface.co/datasets/while-ai/text-to-sql-shop.TextToText_boolqtext-to-sql-eval-predictions
What the text-to-SQL models actually generated
Every prediction behind the numbers in
qwen3-8b-text2sql-qlora: the 453 test
questions of the enterprise text-to-SQL benchmark,
each answered by four configurations of the same model, each answer executed against the reference
PostgreSQL database and scored by comparing result sets. 1,812 rows.
I published this because the headline table (10.82 % → 50.99 % → 52.10 %) is the least interesting part of
that project. The interesting… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-eval-predictions.text-to-sql-phrasing-robustness
Does sloppy phrasing break text-to-SQL?
The enterprise text-to-SQL benchmark
lists its own biggest caveat: every question is template-generated, so real user phrasing is untested.
This is the test. 35 test questions (one per template), each sent to the deployed
pipeline four ways: as written, with a typo, in business shorthand, and stripped to a terse fragment.
24 questions and 85 answers survive the filter described under Setup; every answer was
executed against the database.… See the full description on the dataset page: https://huggingface.co/datasets/hari-krishna-ai/text-to-sql-phrasing-robustness.spider-text-to-sqlspeech_to_texttext_to_sql_ko
Korean Text to MySQL Dataset
Dataset Summary
Korean Text to MySQL is a dataset comprising approximately 3,300 samples generated using OpenAI's gpt-4o model. This dataset is designed to train models that convert natural language questions in Korean into MySQL queries. The data generation process was inspired by the Self-Instruct method and followed the steps outlined below.
Data Generation Process
1. Creation of SEED Dataset
Approximately 100 SEED samples were… See the full description on the dataset page: https://huggingface.co/datasets/won75/text_to_sql_ko.MycoBase-Large-Scale-Text-to-SQL
MycoBase: A Biologically Literate Text-to-SQL Dataset
MycoBase is a synthetic but biologically accurate dataset designed for stress-testing Text-to-SQL systems. It represents a research information system for the study of fungi, covering everything from taxonomy and genomics to morphology and cultivation.
Dataset Highlights
Schema Complexity: 2,016 tables with over 9,000 foreign key relationships.
Data Volume: 320,270 rows of realistic mycology data.
Realistic Names:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/MycoBase-Large-Scale-Text-to-SQL.crowdsourced-text-to-sign-language-rule-based-translation-corpus
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/sltAI/crowdsourced-text-to-sign-language-rule-based-translation-corpus.ACD-Audios-meta-to-text-v1text_to_dsl_opensearch_v1_newimage-to-text-checkpoint-downloadsText_to_ImageTrain Demo Datasets
symptom_text_to_disease_mk2
Dataset Card for "symptom_text_to_disease_mk2"
More Information needed
wat24_text_to_text_translationsanskrit_audio_dataset_under_30_taged_meta_to_text_from_edgemale_part3_taged_meta_to_text_from_edgesymptom_text_to_disease_mk2
Dataset Card for "symptom_text_to_disease_mk2"
More Information needed
