datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.query-expansion
Query Expansion Dataset
This dataset is designed to train search query expansion models that can generate multiple semantic expansions for a given query.
Purpose
The goal of this dataset is to serve as input for training small language models (0.5B to 3B parameters) to act as query expander models in various search systems, including but not limited to Retrieval-Augmented Generation (RAG) systems.
Query expansion is a technique used to enhance search results by generating… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/query-expansion.sql-query-generation-sft-100k
SQL Query Generation SFT (100K)
100,000 ShareGPT conversations demonstrating high-quality SQL query generation from natural language requests. Each example includes a realistic database schema, a natural language query request, a correct SQL query, and a clear explanation of how the query works — across 6 SQL dialects and 15+ complexity levels.
Motivation
Text-to-SQL is one of the highest-value NLP applications in enterprise settings. Common model failures… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/sql-query-generation-sft-100k.opengloss-v1.3-query-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples.sql-query-engine-synthetic
SQL Query Engine — Synthetic Benchmark
A gold-standard NL-to-SQL benchmark containing 75 natural language questions across 3 PostgreSQL databases (e-commerce, university, hospital), each with verified gold SQL queries and expected results. Designed to evaluate text-to-SQL systems with a focus on measuring the impact of iterative self-healing (query repair) loops.
Paper
SQL Query Engine: A Self-Healing LLM Pipeline for Natural Language to PostgreSQL Translation
Muhammad… See the full description on the dataset page: https://huggingface.co/datasets/codeadeel/sql-query-engine-synthetic.personal-query-grocery-and-gourmet-food
Personal Query: Grocery and Gourmet Food
This dataset contains personalized product search queries for the Grocery_and_Gourmet_Food category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-grocery-and-gourmet-food.qwen3.5-2B-vi-query
Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3)
1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B,
trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON:
normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a
retrieval-routing hint, for a downstream medical RAG system.
The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.dfm11-danish-query-templatizer-training
dfm11-danish-query-templatizer-training
Danish query-templatizer supervision produced by the DFM-owned FineInstructions reproduction pipeline.
Rows retain generation and audit provenance. Local filesystem paths are removed.
The synthetic release does not broaden rights attached to upstream grounding
or query sources; consult each row's source provenance and upstream terms.
querysmith-spider-bird
querysmith-spider-bird
Schema-grounded text-to-SQL training data used to fine-tune
ajayk007/Qwen2.5-Coder-7B-Querysmith.
~13.7k examples derived from Spider and
BIRD.
Format
mlx-lm chat format, one example per line:
{"messages": [
{"role": "system", "content": "You are a text-to-SQL generator ..."},
{"role": "user", "content": "Schema:\nCREATE TABLE ...\n\nQuestion: ..."},
{"role": "assistant", "content": "SELECT ..."}
]}
The user turn contains the… See the full description on the dataset page: https://huggingface.co/datasets/ajayk007/querysmith-spider-bird.personal-query-baby-products
Personal Query: Baby Products
This dataset contains personalized product search queries for the Baby_Products category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-baby-products.opengloss-v1.2-query-examples-flat
OpenGloss Query Examples v1.2 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains flattened profile records (one per query).
It is derived from the OpenGloss
encyclopedic dictionary.
Key… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-query-examples-flat.personal-query-pet-supplies
Personal Query: Pet Supplies
This dataset contains personalized product search queries for the Pet_Supplies category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-pet-supplies.opengloss-v1.1-query-examples
OpenGloss Query Examples v1.1 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains flattened profile records (one per query).
It is derived from the OpenGloss
encyclopedic dictionary.
Key… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-query-examples.personalized-query
Personalized Query
This repository contains three personalized product-search query datasets in one Hugging Face dataset page.
Each config corresponds to one product category:
baby: Baby Products
grocery: Grocery and Gourmet Food
pets: Pet Supplies
Each config has two splits:
full: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as null.
paired: only rows where a correct query has a paired error query.
Dataset Size
Config… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personalized-query.i2b2-query-data-1.0
i2b2 query data 1.0
This is a dataset of i2b2 query builder examples that are taken from a test environment of i2b2 and then pre-processed with AI descriptions.
sft-mobile-query-framework
SFT Mobile Query Framework Dataset
Dataset Description
This dataset contains 3607 training pairs for supervised fine-tuning (SFT) of language models to parse natural language queries about mobile phones into structured JSON execution plans.
Dataset Summary
Total Examples: 3607
Format: JSONL (question-answer pairs)
Task: Query Parsing & Structured Output Generation
Domain: Mobile Phone Specifications
Language: English
Purpose
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/sft-mobile-query-framework.opengloss-v1.2-query-examples
OpenGloss Query Examples v1.2
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains word-level records with nested profiles.
It is derived from the OpenGloss
encyclopedic dictionary.
Key Statistics
21… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-query-examples.QueryTranslation
