datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-image-prompts
The dataset of the most popular text-to-image prompts.
Dataset Details
Dataset Description
Curated by: kazimir.ai
Funded by [optional]: [More Information Needed]
Shared by [optional]: https://kazimir.ai
License: apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Free to use.
Dataset Structure
CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.text-to-mermaidtext_to_bashbird_text_to_sqltext-to-mermaid-2text-to-svg
DL Spring 2026 – SVG Generation from Text Prompts - Kaggle Competition Dataset
Dataset Description
Training and test data from the NYU Tandon Deep Learning (ECE-GY 7123)
Spring 2026 Kaggle competition: DL Spring 2026 – SVG Generation from Text Prompts.
This dataset was provided by the course instructors as competition data.
It is redistributed here for reproducibility of the associated course project.
Source
Course: NYU Tandon Deep Learning (ECE-GY 7123)… See the full description on the dataset page: https://huggingface.co/datasets/aagoluoglu/text-to-svg.query_builder_text_to_sqlsynthetic_text_to_sql_th
Synthetic Text-to-SQL Thai Dataset
Thai translation of the gretelai/synthetic_text_to_sql dataset.
Dataset Description
This dataset contains Thai translations of synthetic text-to-SQL examples covering various domains and SQL patterns.
Source
Original Dataset: gretelai/synthetic_text_to_sql
Created by: Gretel.ai
Statistics
Split
Rows
Train
100,000
Test
5,851
Total
105,851
Columns
Column
Description… See the full description on the dataset page: https://huggingface.co/datasets/Porameht/synthetic_text_to_sql_th.text-to-mongodb-queries-llm
Dataset Description
This dataset contains 10,000+ complex SQL-style analytical questions
mapped to MongoDB queries and aggregation pipelines.
Features
Multiple schemas
$group, $sum, $avg, $lookup
Nested documents
Long analytical questions
Use Cases
Fine-tuning small LLMs (Qwen, Mistral, LLaMA 3B)
Text-to-Mongo query generation
Data analytics agents
text-to-mermaid
Text to Mermaid
Description
A curated dataset designed for fine-tuning small language models on diagram generation tasks. Each example follows a strict response format: when prompted for a diagram, the model outputs only the Mermaid syntax—no explanations, no markdown fences, no extra text.
Derived from Celiadraw/text-to-mermaid-2, this version has been verified, pruned, cleaned, and reworded for reliability and consistency.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/ali-thowfeek/text-to-mermaid.text-to-sqlThis dataset is a merged collection of multiple text-to-SQL datasets, designed to provide a comprehensive resource for training and evaluating text-to-SQL models. It combines data from several popular benchmarks, including Spider, CoSQL, SparC, and others, to create a diverse and robust dataset for natural language to SQL query generation tasks.
Dataset Details
Dataset Description
Curated by: Mudasir Ahmad Mir
Language(s) (NLP): English
License: Apache 2.0
This dataset is ideal for researchers… See the full description on the dataset page: https://huggingface.co/datasets/Mudasir692/text-to-sql.MycoBase-Large-Scale-Text-to-SQL
MycoBase: A Biologically Literate Text-to-SQL Dataset
MycoBase is a synthetic but biologically accurate dataset designed for stress-testing Text-to-SQL systems. It represents a research information system for the study of fungi, covering everything from taxonomy and genomics to morphology and cultivation.
Dataset Highlights
Schema Complexity: 2,016 tables with over 9,000 foreign key relationships.
Data Volume: 320,270 rows of realistic mycology data.
Realistic Names:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/MycoBase-Large-Scale-Text-to-SQL.Speech-to-textimage-to-text-ocr-serp-snapshot
Image-to-Text OCR SERP Snapshot
An English-language, US Google organic-results snapshot for image to text converter, collected on 2026-09-22.
Files
image-to-text-ocr-serp-snapshot.csv is the machine-readable row-by-row dataset.
comparison.md is the same comparison as a readable Markdown table with methodology and limitations.
How to read the dataset
Each row is one result. The classification column distinguishes the Android app-store listing from… See the full description on the dataset page: https://huggingface.co/datasets/phoenix11000/image-to-text-ocr-serp-snapshot.Clinton_text_to_sql_3000text_to_sql_10ktext_to_sql_FALCONGroup_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction
Title
An Annotated Dataset for Hearing-Impaired Speech-to-Text Correction
Abstract
This dataset is specifically designed for the task of correcting speech to text errors in hearing-impaired individuals.
It includes one hour of real speech files of hearing-impaired individuals, automatic speech recognition (ASR) output text, and manually corrected standard text.
We searched for an hour of audio from hearing-impaired individuals to ensure that the voice was authentic and… See the full description on the dataset page: https://huggingface.co/datasets/llllliuuy/Group_M_An_Annotated_Dataset_for_Hearing-Impaired_Speech-to-Text_Correction.text-to-pandasRace-text-to-quiz-jsontext_to_cyphertext_to_sql_FLANcrowdsourced-text-to-sign-language-rule-based-translation-corpus
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/sltAI/crowdsourced-text-to-sign-language-rule-based-translation-corpus.webnlg_Table_to_Textllama_text_to_sql_datasetmantella_text_to_actiontext_to_json6069_sample_synthetic_text_to_sqltext-to-sqltext-to-python-synthetic
