datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opengloss-v1.3-query-examples-flat
See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.
OpenGloss Query Examples v1.3 (Flattened)
Dataset Summary
OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary
terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.beaver-query
Dataset Card for beaver-query
Homepage and leaderboard |
Github repository |
Paper
Beaver is a holistic framework for evaluating performance on complex, private‑enterprise text‑to‑SQL tasks.
This repository includes questions and corresponding annotations. We reserve a portion of the full question set as a private, hidden test set.
Each sample contains:
id: ID of the question
category: one of real, complex query, domain-specific query, domain-specific complex query.
real indicates… See the full description on the dataset page: https://huggingface.co/datasets/beaverbench/beaver-query.wikipedia-multilingual-synthetic-ir-query
wikipedia-multilingual-synthetic-ir-query
This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training.
It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text.
The current release contains two different retrieval settings:
short_doc: pairs of (query, short document)
long_doc: pairs of (query, long document)
These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.bing_coronavirus_query_set
Dataset Card for BingCoronavirusQuerySet
Dataset Summary
Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020
example:
load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30")
You can also load the data by country by using queries_by="country".
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.wikipedia-trivia-query-variationRetrieve-PileRetrieve-Pile (Knowledge-Pile) is a knowledge-related data leveraging Retrieve-from-CC (We also called this method as "Query of CC"),a total of 735GB disk size and 188B tokens (using Llama2 tokenizer).
Retrieve-from-CC
Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage. Leveraging the great generalization capability of large… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Retrieve-Pile.wildchat-category-query-expanded-n6vsc2022-test-query-2fpsquery2doc_msmarcoThis dataset contains GPT-3.5 (text-davinci-003) generations from MS-MARCO queries.targeted-query-set-processeddeepresearch-bench-queryinstruct-query-datatask674_google_wellformed_query_sentence_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task674_google_wellformed_query_sentence_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task674_google_wellformed_query_sentence_generation.wikipedia-nq-query-variationwildchat-reattributed-query-triplewildchat-category-query-expanded-n3wildchat-reconstructed-query-distinctwildchat-reattributed-query-distinctimage-query-vienemotron-terminal-data_querying
nemotron-terminal-data_querying
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "data_querying". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_querying.medium_5000_data_queryingKnowledge_PileKnowledge Pile is a knowledge-related data leveraging Query of CC.
This dataset is a partial of Knowledge Pile(about 40GB disk size), full datasets have been released in [🤗 knowledge_pile_full], a total of 735GB disk size and 188B tokens (using Llama2 tokenizer).
Query of CC
Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage.… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Knowledge_Pile.wildchat-reattributed-query-uniquegooaq-query-arctic-pooledeasy_5000_data_queryinghtml-query-text-HtmlRAG
html-query-text-HtmlRAG
Warning: This dataset is under development and its content is subject to change!
This dataset is a processed and cleaned version of the zstanjj/HtmlRAG-train dataset. It has been specifically prepared for task of HTML cleaning.
🚀 Supported Tasks
This dataset is primarily designed for:
HTML Cleaning: Training models to take the messy html as input and generate the cleaned_html or cleaned_text as output.
Question Answering: Training models to… See the full description on the dataset page: https://huggingface.co/datasets/williambrach/html-query-text-HtmlRAG.llm-query-complexity-benchmark
LLM Query Complexity Benchmark
A multi-domain, perfectly balanced dataset of 6,000 labeled queries (4,800 train / 1,200 test) for training and evaluating LLM query complexity classifiers that route queries to the most cost-effective inference tier.
Built for the STREAM project (Smart Tiered Routing Engine for AI Models), which routes queries automatically between local CPU models, institutional HPC GPU clusters, and cloud API tiers.
Dataset Summary
Split
Queries… See the full description on the dataset page: https://huggingface.co/datasets/anasnassar/llm-query-complexity-benchmark.medium_5000_data_querying_fixedmedium_5000-data_querying_n100k1entity-query-summarization
