CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mjbommar /opengloss-v1.3-query-examples-flat See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Query Examples v1.3 (Flattened) Dataset Summary OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.texttext-generation100K<n<1M0 likes689 downloads19d agoHugging Face02beaverbench /beaver-querygated Dataset Card for beaver-query Homepage and leaderboard | Github repository | Paper Beaver is a holistic framework for evaluating performance on complex, private‑enterprise text‑to‑SQL tasks. This repository includes questions and corresponding annotations. We reserve a portion of the full question set as a private, hidden test set. Each sample contains: id: ID of the question category: one of real, complex query, domain-specific query, domain-specific complex query. real indicates… See the full description on the dataset page: https://huggingface.co/datasets/beaverbench/beaver-query.text1K<n<10K9 likes608 downloads4mo agoHugging Face03hotchpotch /wikipedia-multilingual-synthetic-ir-query wikipedia-multilingual-synthetic-ir-query This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training. It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text. The current release contains two different retrieval settings: short_doc: pairs of (query, short document) long_doc: pairs of (query, long document) These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.tabulartext-retrieval10M<n<100M0 likes529 downloads3mo agoHugging Face04microsoft /bing_coronavirus_query_set Dataset Card for BingCoronavirusQuerySet Dataset Summary Please note that you can specify the start and end date of the data. You can get start and end dates from here: https://github.com/microsoft/BingCoronavirusQuerySet/tree/master/data/2020 example: load_dataset("bing_coronavirus_query_set", queries_by="state", start_date="2020-09-01", end_date="2020-09-30") You can also load the data by country by using queries_by="country". Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/bing_coronavirus_query_set.tabulartext-classification100K<n<1M1 likes471 downloads3y agoHugging Face05spacemanidol /wikipedia-trivia-query-variationtext10K<n<100K1 likes463 downloads4y agoHugging Face06Query-of-CC /Retrieve-PileRetrieve-Pile (Knowledge-Pile) is a knowledge-related data leveraging Retrieve-from-CC (We also called this method as "Query of CC"),a total of 735GB disk size and 188B tokens (using Llama2 tokenizer). Retrieve-from-CC Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage. Leveraging the great generalization capability of large… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Retrieve-Pile.text10M<n<100M8 likes426 downloads2y agoHugging Face07nate-rahn /wildchat-category-query-expanded-n6tabular1M<n<10M0 likes304 downloads1y agoHugging Face08nateraw /vsc2022-test-query-2fpstextn<1K0 likes281 downloads4y agoHugging Face09intfloat /query2doc_msmarcoThis dataset contains GPT-3.5 (text-davinci-003) generations from MS-MARCO queries.text100K<n<1M17 likes253 downloads3y agoHugging Face10Harvard-DCML /targeted-query-set-processedtextn<1K0 likes245 downloads8mo agoHugging Face11lee64 /deepresearch-bench-querytextn<1K0 likes215 downloads11mo agoHugging Face12HCAI-Lab-GT /instruct-query-datatext10K<n<100K0 likes180 downloads6mo agoHugging Face13Lots-of-LoRAs /task674_google_wellformed_query_sentence_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task674_google_wellformed_query_sentence_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task674_google_wellformed_query_sentence_generation.texttext-generation1K<n<10K0 likes179 downloads2y agoHugging Face14spacemanidol /wikipedia-nq-query-variationtext10K<n<100K0 likes178 downloads4y agoHugging Face15nate-rahn /wildchat-reattributed-query-tripletabular1M<n<10M0 likes173 downloads1y agoHugging Face16nate-rahn /wildchat-category-query-expanded-n3tabular1M<n<10M0 likes160 downloads1y agoHugging Face17nate-rahn /wildchat-reconstructed-query-distincttabular1M<n<10M0 likes159 downloads1y agoHugging Face18nate-rahn /wildchat-reattributed-query-distincttabular1M<n<10M0 likes158 downloads1y agoHugging Face19thangvip /image-query-vieimage100K<n<1M0 likes126 downloads1y agoHugging Face20laion /nemotron-terminal-data_querying nemotron-terminal-data_querying Per-source partition of nvidia/Nemotron-Terminal-Corpus, filtered to source == "data_querying". The difficulty column preserves the original easy / medium / mixed split (na for the dataset_adapters/* files, which did not carry a difficulty label). Partitioning scheme: adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet {skill} (e.g. debugging, security, …) — rows from synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-data_querying.textquestion-answering1K<n<10K0 likes126 downloads6mo agoHugging Face21renjiepi /medium_5000_data_queryingtext1K<n<10K0 likes120 downloads9mo agoHugging Face22Query-of-CC /Knowledge_PileKnowledge Pile is a knowledge-related data leveraging Query of CC. This dataset is a partial of Knowledge Pile(about 40GB disk size), full datasets have been released in [🤗 knowledge_pile_full], a total of 735GB disk size and 188B tokens (using Llama2 tokenizer). Query of CC Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage.… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Knowledge_Pile.text1M<n<10M22 likes118 downloads3y agoHugging Face23nate-rahn /wildchat-reattributed-query-uniquetabular1M<n<10M0 likes117 downloads1y agoHugging Face24carsondial /gooaq-query-arctic-pooledtext1M<n<10M0 likes117 downloads11mo agoHugging Face25renjiepi /easy_5000_data_queryingtext1K<n<10K0 likes110 downloads9mo agoHugging Face26williambrach /html-query-text-HtmlRAG html-query-text-HtmlRAG Warning: This dataset is under development and its content is subject to change! This dataset is a processed and cleaned version of the zstanjj/HtmlRAG-train dataset. It has been specifically prepared for task of HTML cleaning. 🚀 Supported Tasks This dataset is primarily designed for: HTML Cleaning: Training models to take the messy html as input and generate the cleaned_html or cleaned_text as output. Question Answering: Training models to… See the full description on the dataset page: https://huggingface.co/datasets/williambrach/html-query-text-HtmlRAG.textfeature-extraction10K<n<100K0 likes106 downloads11mo agoHugging Face27anasnassar /llm-query-complexity-benchmark LLM Query Complexity Benchmark A multi-domain, perfectly balanced dataset of 6,000 labeled queries (4,800 train / 1,200 test) for training and evaluating LLM query complexity classifiers that route queries to the most cost-effective inference tier. Built for the STREAM project (Smart Tiered Routing Engine for AI Models), which routes queries automatically between local CPU models, institutional HPC GPU clusters, and cloud API tiers. Dataset Summary Split Queries… See the full description on the dataset page: https://huggingface.co/datasets/anasnassar/llm-query-complexity-benchmark.texttext-classification1K<n<10K1 likes105 downloads4mo agoHugging Face28renjiepi /medium_5000_data_querying_fixedtext1K<n<10K0 likes98 downloads9mo agoHugging Face29renjiepi /medium_5000-data_querying_n100k1text1K<n<10K0 likes90 downloads8mo agoHugging Face30ttss /entity-query-summarizationtext100K<n<1M1 likes87 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.