datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.Retrieve-PileRetrieve-Pile (Knowledge-Pile) is a knowledge-related data leveraging Retrieve-from-CC (We also called this method as "Query of CC"),a total of 735GB disk size and 188B tokens (using Llama2 tokenizer).
Retrieve-from-CC
Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage. Leveraging the great generalization capability of large… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Retrieve-Pile.deepresearch-bench-queryinstruct-query-dataKnowledge_PileKnowledge Pile is a knowledge-related data leveraging Query of CC.
This dataset is a partial of Knowledge Pile(about 40GB disk size), full datasets have been released in [🤗 knowledge_pile_full], a total of 735GB disk size and 188B tokens (using Llama2 tokenizer).
Query of CC
Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage.… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Knowledge_Pile.llm-query-complexity-benchmark
LLM Query Complexity Benchmark
A multi-domain, perfectly balanced dataset of 6,000 labeled queries (4,800 train / 1,200 test) for training and evaluating LLM query complexity classifiers that route queries to the most cost-effective inference tier.
Built for the STREAM project (Smart Tiered Routing Engine for AI Models), which routes queries automatically between local CPU models, institutional HPC GPU clusters, and cloud API tiers.
Dataset Summary
Split
Queries… See the full description on the dataset page: https://huggingface.co/datasets/anasnassar/llm-query-complexity-benchmark.10K_Report_Query_ToolFHIR_QnA_Query-Based_Resource_Relevance_Classification_T1
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
instruct-cot-query-datareasoner-rewritten-query-0821
Introduction
Rewrite query for BGE Reasoner.
Load Dataset
An example to load the dataset:
import datasets
# load dataset
dataset = datasets.load_dataset(
"cfli/reasoner-rewritten-query-0821",
'biology',
split='train'
)
# print one sample
print(dataset[0])
query-intent-setfit-v1-train
EEA Query Intent — setfit-v1 training data
The exact training mix used to train
eeahugs/query-intent-setfit-v1,
the multilingual query-intent classifier for the European Environment
Agency website search. 81,404 rows, 28 languages,
5 intent classes.
License: CC-BY-NC-4.0 (non-commercial use only) — see "Provenance and
licensing" below.
Format
JSON Lines. One object per line:
field
meaning
id
row id, <prefix>-<lang>-<intent>-NNNN
intent
one of question… See the full description on the dataset page: https://huggingface.co/datasets/eeahugs/query-intent-setfit-v1-train.query-expansion
Query Expansion Dataset
This dataset is designed to train search query expansion models that can generate multiple semantic expansions for a given query.
Purpose
The goal of this dataset is to serve as input for training small language models (0.5B to 3B parameters) to act as query expander models in various search systems, including but not limited to Retrieval-Augmented Generation (RAG) systems.
Query expansion is a technique used to enhance search results by generating… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/query-expansion.sql-query-generation-sft-100k
SQL Query Generation SFT (100K)
100,000 ShareGPT conversations demonstrating high-quality SQL query generation from natural language requests. Each example includes a realistic database schema, a natural language query request, a correct SQL query, and a clear explanation of how the query works — across 6 SQL dialects and 15+ complexity levels.
Motivation
Text-to-SQL is one of the highest-value NLP applications in enterprise settings. Common model failures… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/sql-query-generation-sft-100k.opengloss-v1.3-query-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples.base-query-dataanimal-welfare-veganism-query-corpus
Animal Welfare & Veganism Query Corpus (v0)
An open, organically-sourced dataset of real questions people ask about animal welfare, veganism, vegetarianism, farmed animals, animal ethics, and animal sentience.
Built by Consider Sentience, a research and tooling practice at the intersection of AI and animal welfare.
Dataset Summary
This dataset contains 1,142 real, organically-asked questions drawn from public Q&A communities, covering the range of things people… See the full description on the dataset page: https://huggingface.co/datasets/considersentience/animal-welfare-veganism-query-corpus.sql-query-engine-synthetic
SQL Query Engine — Synthetic Benchmark
A gold-standard NL-to-SQL benchmark containing 75 natural language questions across 3 PostgreSQL databases (e-commerce, university, hospital), each with verified gold SQL queries and expected results. Designed to evaluate text-to-SQL systems with a focus on measuring the impact of iterative self-healing (query repair) loops.
Paper
SQL Query Engine: A Self-Healing LLM Pipeline for Natural Language to PostgreSQL Translation
Muhammad… See the full description on the dataset page: https://huggingface.co/datasets/codeadeel/sql-query-engine-synthetic.query_datapersonal-query-grocery-and-gourmet-food
Personal Query: Grocery and Gourmet Food
This dataset contains personalized product search queries for the Grocery_and_Gourmet_Food category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-grocery-and-gourmet-food.qwen3.5-2B-vi-query
Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3)
1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B,
trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON:
normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a
retrieval-routing hint, for a downstream medical RAG system.
The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.dfm11-danish-query-templatizer-training
dfm11-danish-query-templatizer-training
Danish query-templatizer supervision produced by the DFM-owned FineInstructions reproduction pipeline.
Rows retain generation and audit provenance. Local filesystem paths are removed.
The synthetic release does not broaden rights attached to upstream grounding
or query sources; consult each row's source provenance and upstream terms.
querysmith-spider-bird
querysmith-spider-bird
Schema-grounded text-to-SQL training data used to fine-tune
ajayk007/Qwen2.5-Coder-7B-Querysmith.
~13.7k examples derived from Spider and
BIRD.
Format
mlx-lm chat format, one example per line:
{"messages": [
{"role": "system", "content": "You are a text-to-SQL generator ..."},
{"role": "user", "content": "Schema:\nCREATE TABLE ...\n\nQuestion: ..."},
{"role": "assistant", "content": "SELECT ..."}
]}
The user turn contains the… See the full description on the dataset page: https://huggingface.co/datasets/ajayk007/querysmith-spider-bird.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1_complete
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
querygen-data-v4
Nixiesearch querygen-v4 model training dataset
A dataset used to train the not-yet-published querygen-v4 model from Nixiesearch. The dataset is a combination of multiple open query-document datasets in a format for Causal LLM training.
Used datasets
We use train splits from the following datasets:
MSMARCO: 532751 rows
HotpotQA: 170000 rows
NQ: 58554 rows
MIRACL en: 1193 rows
SQUAD: 85710 rows
TriviaQA: 60283 rows
The train split is 900000 rows, and test split is 8491.… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/querygen-data-v4.korean-service-query-gating-v1
Korean Service Query Gating V1
한국어 서비스 질의 게이팅 연구를 위한 합성 데이터셋입니다.
현실의 서비스 회사에서는 질문 분해, 도메인 판별, 라우팅 같은 태스크가 매우 중요하지만, 실제 질의 로그는 개인정보와 내부 정책 문제 때문에 공개하기 어렵습니다.이 데이터셋은 그런 제약 아래에서도 재현 가능한 연구를 시작할 수 있도록 만든 공개 가능한 Korean starting benchmark입니다.
쉽게 말해 이 데이터셋은 아래 질문을 연구하기 위한 리소스입니다.
“이 질문이 우리 서비스 범위 안인가, 밖인가?”
“겉보기엔 비슷한데 실제 의미는 다른 질문을 어떻게 구분할까?”
“모델이 어떤 실패 유형에서 흔들리는가?”
Quick Summary
언어: 한국어
용도: 서비스 질의 게이팅 / 헬프데스크 질의 분류 / OOD stress evaluation
구성: core + stress_eval
성격: 합성 연구용… See the full description on the dataset page: https://huggingface.co/datasets/taeyun16/korean-service-query-gating-v1.query-positive-pairs-smallluxical-query-expansion-finance
Luxical Query Expansion Financial/ESG Dataset (Preliminary)
This dataset package is designed to reproduce the experiments in this repository.
Included Files
finance_priors_v2.tsv: Mined term priors from local financial corpus.
queries_eval_finance_esg.json: 12-query weak-label benchmark set.
quality_benchmark.json: Policy quality benchmark outputs.
latency_benchmark_combined.json: Combined-policy latency benchmark.
latency_benchmark_conservative.json: Conservative-policy… See the full description on the dataset page: https://huggingface.co/datasets/oneryalcin/luxical-query-expansion-finance.RCW_2025_Positive_Query_Pairs
The Washington law Benchmark (WLB)
Dataset Summary
The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written).
The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/Darther/RCW_2025_Positive_Query_Pairs.personal-query-baby-products
Personal Query: Baby Products
This dataset contains personalized product search queries for the Baby_Products category.
Each record is built from the Personal Query pipeline:
Stage 6 generated correct personalized queries.
Stage 7 injected user-specific error query variants when a matching error pattern was available.
Stage 5 provided the user profile complexity level.
Files
data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-baby-products.query_routing
