datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.nemotron-terminal-data_querying
nemotron-terminal-data_querying
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_querying". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_querying.html-query-text-HtmlRAG
html-query-text-HtmlRAG
Warning: This dataset is under development and its content is subject to change!
This dataset is a processed and cleaned version of the zstanjj/HtmlRAG-train dataset. It has been specifically prepared for task of HTML cleaning.
🚀 Supported Tasks
This dataset is primarily designed for:
HTML Cleaning: Training models to take the messy html as input and generate the cleaned_html or cleaned_text as output.
Question Answering: Training models to… See the full description on the dataset page: https://huggingface.co/datasets/williambrach/html-query-text-HtmlRAG.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
omnimcp_trpc_tanstack_query_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_trpc_tanstack_query_teaser.omnimcp_healthtech_cohort_query_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_healthtech_cohort_query_teaser.opengloss-v1.3-query-examples
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples.animal-welfare-veganism-query-corpus
Animal Welfare & Veganism Query Corpus (v0)
An open, organically-sourced dataset of real questions people ask about animal welfare, veganism, vegetarianism, farmed animals, animal ethics, and animal sentience.
Built by Consider Sentience, a research and tooling practice at the intersection of AI and animal welfare.
Dataset Summary
This dataset contains 1,142 real, organically-asked questions drawn from public Q&A communities, covering the range of things people… See the full description on the dataset page: https://huggingface.co/datasets/considersentience/animal-welfare-veganism-query-corpus.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1_complete
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
ecommerce-query-rewriting
#e-commerce-query-rewriting-dataset
Hub: mudasir13cs/ecommerce-query-rewriting
A dataset of 10,000 examples pairing ambiguous, context-dependent user queries with their fully resolved, context-aware rewrites for e-commerce product search. Built for fine-tuning LLMs to resolve pronouns, ellipsis, ordinals, and other conversational shortcuts using prior search context — the kind of resolution real shopping assistants need to handle turns like "show me that one" or "the cheaper… See the full description on the dataset page: https://huggingface.co/datasets/mudasir13cs/ecommerce-query-rewriting.RCW_2025_Positive_Query_Pairs
The Washington law Benchmark (WLB)
Dataset Summary
The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written).
The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/Darther/RCW_2025_Positive_Query_Pairs.NLP-to-Semantic-Query_Benchmark_Dataset
NLP-to-Semantic-Query Benchmark Dataset
Overview
This dataset is designed for evaluating AI agents and LLM systems that translate natural language analytical questions into structured semantic queries.
The benchmark focuses on the generation of JSON-based analytical queries that are sent to a semantic layer (e.g. Cube.js) to retrieve analytical results from databases.
The dataset can be used for:
Evaluating NLP-to-query systems
Benchmarking AI analytics agents
Measuring… See the full description on the dataset page: https://huggingface.co/datasets/BatSilver/NLP-to-Semantic-Query_Benchmark_Dataset.opengloss-v1.2-query-examples-flat
OpenGloss Query Examples v1.2 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains flattened profile records (one per query).
It is derived from the OpenGloss
encyclopedic dictionary.
Key… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-query-examples-flat.opengloss-v1.1-query-examples
OpenGloss Query Examples v1.1 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains flattened profile records (one per query).
It is derived from the OpenGloss
encyclopedic dictionary.
Key… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.1-query-examples.sft-mobile-query-framework
SFT Mobile Query Framework Dataset
Dataset Description
This dataset contains 3607 training pairs for supervised fine-tuning (SFT) of language models to parse natural language queries about mobile phones into structured JSON execution plans.
Dataset Summary
Total Examples: 3607
Format: JSONL (question-answer pairs)
Task: Query Parsing & Structured Output Generation
Domain: Mobile Phone Specifications
Language: English
Purpose
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/sujitpandey/sft-mobile-query-framework.opengloss-v1.2-query-examples
OpenGloss Query Examples v1.2
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple query profiles covering different search intents and user personas,
making it ideal for training query generation, intent classification, and RAG systems.
This dataset contains word-level records with nested profiles.
It is derived from the OpenGloss
encyclopedic dictionary.
Key Statistics
21… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-query-examples.nepali-query-passage-hard-negatives-10kQueryBridge
QueryBridge: One Million Annotated Questions with SPARQL Queries - Dataset for Question Answering over Knowledge Graph
The QueryBridge dataset is the first and largest dataset with annotated questions for question answering (QA) over knowledge graphs. It provides a comprehensive resource for developing and testing algorithms that process and interpret natural language questions in the context of structured knowledge. In addition to QA tasks, QueryBridge can also be used for:
Entity… See the full description on the dataset page: https://huggingface.co/datasets/aorogat/QueryBridge.Think_and_Query_value_for_R1
Introduction
This repository implements a Shapley value-based approach to quantitatively evaluate the contributions of query (q) and think (t) in generating answer (a).
Method
think_value = [loss(a|q) - loss(a|q,t) + loss(a|∅) - loss(a|t)] / 2
query_value = [loss(a|t) - loss(a|q,t) + loss(a|∅) - loss(a|q)] / 2
think_ratio = think_value/loss(a|∅)
query_ratio = query_value/loss(a|∅)
Original dataset… See the full description on the dataset page: https://huggingface.co/datasets/caihuaiguang/Think_and_Query_value_for_R1.mapeval_queryThis dataset is copied from https://github.com/MapEval/MapEval-API/blob/main/dataset.json for easy use.
