CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mjbommar /opengloss-v1.3-query-examples-flat See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Query Examples v1.3 (Flattened) Dataset Summary OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary terms. Each term has multiple… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples-flat.texttext-generation100K<n<1M0 likes689 downloads19d agoHugging Face02Query-of-CC /Retrieve-PileRetrieve-Pile (Knowledge-Pile) is a knowledge-related data leveraging Retrieve-from-CC (We also called this method as "Query of CC"),a total of 735GB disk size and 188B tokens (using Llama2 tokenizer). Retrieve-from-CC Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage. Leveraging the great generalization capability of large… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Retrieve-Pile.text10M<n<100M8 likes426 downloads2y agoHugging Face03lee64 /deepresearch-bench-querytextn<1K0 likes215 downloads11mo agoHugging Face04HCAI-Lab-GT /instruct-query-datatext10K<n<100K0 likes180 downloads6mo agoHugging Face05Query-of-CC /Knowledge_PileKnowledge Pile is a knowledge-related data leveraging Query of CC. This dataset is a partial of Knowledge Pile(about 40GB disk size), full datasets have been released in [🤗 knowledge_pile_full], a total of 735GB disk size and 188B tokens (using Llama2 tokenizer). Query of CC Just like the figure below, we initially collected seed information in some specific domains, such as keywords, frequently asked questions, and textbooks, to serve as inputs for the Query Bootstrapping stage.… See the full description on the dataset page: https://huggingface.co/datasets/Query-of-CC/Knowledge_Pile.text1M<n<10M22 likes118 downloads3y agoHugging Face06anasnassar /llm-query-complexity-benchmark LLM Query Complexity Benchmark A multi-domain, perfectly balanced dataset of 6,000 labeled queries (4,800 train / 1,200 test) for training and evaluating LLM query complexity classifiers that route queries to the most cost-effective inference tier. Built for the STREAM project (Smart Tiered Routing Engine for AI Models), which routes queries automatically between local CPU models, institutional HPC GPU clusters, and cloud API tiers. Dataset Summary Split Queries… See the full description on the dataset page: https://huggingface.co/datasets/anasnassar/llm-query-complexity-benchmark.texttext-classification1K<n<10K1 likes105 downloads4mo agoHugging Face07jsmarkschoon /10K_Report_Query_Tooltextn<1K0 likes72 downloads7mo agoHugging Face08genloop /FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1 Dataset Card This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs. textquestion-answeringn<1K0 likes71 downloads2y agoHugging Face09HCAI-Lab-GT /instruct-cot-query-datatext10K<n<100K0 likes53 downloads6mo agoHugging Face10cfli /reasoner-rewritten-query-0821 Introduction Rewrite query for BGE Reasoner. Load Dataset An example to load the dataset: import datasets # load dataset dataset = datasets.load_dataset( "cfli/reasoner-rewritten-query-0821", 'biology', split='train' ) # print one sample print(dataset[0]) text1K<n<10K0 likes51 downloads1y agoHugging Face11eeahugs /query-intent-setfit-v1-train EEA Query Intent — setfit-v1 training data The exact training mix used to train eeahugs/query-intent-setfit-v1, the multilingual query-intent classifier for the European Environment Agency website search. 81,404 rows, 28 languages, 5 intent classes. License: CC-BY-NC-4.0 (non-commercial use only) — see "Provenance and licensing" below. Format JSON Lines. One object per line: field meaning id row id, <prefix>-<lang>-<intent>-NNNN intent one of question… See the full description on the dataset page: https://huggingface.co/datasets/eeahugs/query-intent-setfit-v1-train.texttext-classification10K<n<100K0 likes51 downloads8d agoHugging Face12s-emanuilov /query-expansion Query Expansion Dataset This dataset is designed to train search query expansion models that can generate multiple semantic expansions for a given query. Purpose The goal of this dataset is to serve as input for training small language models (0.5B to 3B parameters) to act as query expander models in various search systems, including but not limited to Retrieval-Augmented Generation (RAG) systems. Query expansion is a technique used to enhance search results by generating… See the full description on the dataset page: https://huggingface.co/datasets/s-emanuilov/query-expansion.texttext-generation1K<n<10K4 likes47 downloads2y agoHugging Face13stindardlogic /sql-query-generation-sft-100k SQL Query Generation SFT (100K) 100,000 ShareGPT conversations demonstrating high-quality SQL query generation from natural language requests. Each example includes a realistic database schema, a natural language query request, a correct SQL query, and a clear explanation of how the query works — across 6 SQL dialects and 15+ complexity levels. Motivation Text-to-SQL is one of the highest-value NLP applications in enterprise settings. Common model failures… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/sql-query-generation-sft-100k.texttext-generation100K<n<1M0 likes47 downloads2mo agoHugging Face14mjbommar /opengloss-v1.3-query-examples See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Query Examples v1.3 Dataset Summary OpenGloss Query Examples is a synthetic dataset of search queries generated for vocabulary terms. Each term has multiple query… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-query-examples.texttext-generation10K<n<100K0 likes46 downloads19d agoHugging Face15HCAI-Lab-GT /base-query-datatext10K<n<100K0 likes44 downloads6mo agoHugging Face16considersentience /animal-welfare-veganism-query-corpus Animal Welfare & Veganism Query Corpus (v0) An open, organically-sourced dataset of real questions people ask about animal welfare, veganism, vegetarianism, farmed animals, animal ethics, and animal sentience. Built by Consider Sentience, a research and tooling practice at the intersection of AI and animal welfare. Dataset Summary This dataset contains 1,142 real, organically-asked questions drawn from public Q&A communities, covering the range of things people… See the full description on the dataset page: https://huggingface.co/datasets/considersentience/animal-welfare-veganism-query-corpus.texttext-classification1K<n<10K0 likes42 downloads1mo agoHugging Face17codeadeel /sql-query-engine-synthetic SQL Query Engine — Synthetic Benchmark A gold-standard NL-to-SQL benchmark containing 75 natural language questions across 3 PostgreSQL databases (e-commerce, university, hospital), each with verified gold SQL queries and expected results. Designed to evaluate text-to-SQL systems with a focus on measuring the impact of iterative self-healing (query repair) loops. Paper SQL Query Engine: A Self-Healing LLM Pipeline for Natural Language to PostgreSQL Translation Muhammad… See the full description on the dataset page: https://huggingface.co/datasets/codeadeel/sql-query-engine-synthetic.texttext-generationn<1K0 likes41 downloads5mo agoHugging Face18kangz /query_datatext1M<n<10M0 likes40 downloads2y agoHugging Face19xxxxdszz /personal-query-grocery-and-gourmet-food Personal Query: Grocery and Gourmet Food This dataset contains personalized product search queries for the Grocery_and_Gourmet_Food category. Each record is built from the Personal Query pipeline: Stage 6 generated correct personalized queries. Stage 7 injected user-specific error query variants when a matching error pattern was available. Stage 5 provided the user profile complexity level. Files data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-grocery-and-gourmet-food.tabulartext-generation10K<n<100K0 likes39 downloads5mo agoHugging Face20daipham31 /qwen3.5-2B-vi-query Vietnamese Medical Query Normalization / Expansion / Routing Pack (v3) 1162 synthetic ChatML examples for fine-tuning a small Vietnamese model (target: Qwen/Qwen3.5-2B, trained with Unsloth) to turn a raw, everyday Vietnamese medical query into structured JSON: normalized query, intent, entities, must-preserve tokens, lexical/semantic query variants, and a retrieval-routing hint, for a downstream medical RAG system. The model does not answer medical questions. It only normalizes… See the full description on the dataset page: https://huggingface.co/datasets/daipham31/qwen3.5-2B-vi-query.texttext-generation1K<n<10K0 likes35 downloads4d agoHugging Face21schneiderkamplab /dfm11-danish-query-templatizer-training dfm11-danish-query-templatizer-training Danish query-templatizer supervision produced by the DFM-owned FineInstructions reproduction pipeline. Rows retain generation and audit provenance. Local filesystem paths are removed. The synthetic release does not broaden rights attached to upstream grounding or query sources; consult each row's source provenance and upstream terms. texttext-generation10K<n<100K0 likes32 downloads19d agoHugging Face22ajayk007 /querysmith-spider-bird querysmith-spider-bird Schema-grounded text-to-SQL training data used to fine-tune ajayk007/Qwen2.5-Coder-7B-Querysmith. ~13.7k examples derived from Spider and BIRD. Format mlx-lm chat format, one example per line: {"messages": [ {"role": "system", "content": "You are a text-to-SQL generator ..."}, {"role": "user", "content": "Schema:\nCREATE TABLE ...\n\nQuestion: ..."}, {"role": "assistant", "content": "SELECT ..."} ]} The user turn contains the… See the full description on the dataset page: https://huggingface.co/datasets/ajayk007/querysmith-spider-bird.texttext-generation10K<n<100K0 likes31 downloads3mo agoHugging Face23genloop /FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1_complete Dataset Card This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs. textquestion-answeringn<1K0 likes25 downloads2y agoHugging Face24nixiesearch /querygen-data-v4 Nixiesearch querygen-v4 model training dataset A dataset used to train the not-yet-published querygen-v4 model from Nixiesearch. The dataset is a combination of multiple open query-document datasets in a format for Causal LLM training. Used datasets We use train splits from the following datasets: MSMARCO: 532751 rows HotpotQA: 170000 rows NQ: 58554 rows MIRACL en: 1193 rows SQUAD: 85710 rows TriviaQA: 60283 rows The train split is 900000 rows, and test split is 8491.… See the full description on the dataset page: https://huggingface.co/datasets/nixiesearch/querygen-data-v4.text100K<n<1M1 likes25 downloads2y agoHugging Face25taeyun16 /korean-service-query-gating-v1 Korean Service Query Gating V1 한국어 서비스 질의 게이팅 연구를 위한 합성 데이터셋입니다. 현실의 서비스 회사에서는 질문 분해, 도메인 판별, 라우팅 같은 태스크가 매우 중요하지만, 실제 질의 로그는 개인정보와 내부 정책 문제 때문에 공개하기 어렵습니다.이 데이터셋은 그런 제약 아래에서도 재현 가능한 연구를 시작할 수 있도록 만든 공개 가능한 Korean starting benchmark입니다. 쉽게 말해 이 데이터셋은 아래 질문을 연구하기 위한 리소스입니다. “이 질문이 우리 서비스 범위 안인가, 밖인가?” “겉보기엔 비슷한데 실제 의미는 다른 질문을 어떻게 구분할까?” “모델이 어떤 실패 유형에서 흔들리는가?” Quick Summary 언어: 한국어 용도: 서비스 질의 게이팅 / 헬프데스크 질의 분류 / OOD stress evaluation 구성: core + stress_eval 성격: 합성 연구용… See the full description on the dataset page: https://huggingface.co/datasets/taeyun16/korean-service-query-gating-v1.texttext-classification1K<n<10K0 likes23 downloads6mo agoHugging Face26nixiesearch /query-positive-pairs-smalltext100K<n<1M1 likes22 downloads3y agoHugging Face27oneryalcin /luxical-query-expansion-finance Luxical Query Expansion Financial/ESG Dataset (Preliminary) This dataset package is designed to reproduce the experiments in this repository. Included Files finance_priors_v2.tsv: Mined term priors from local financial corpus. queries_eval_finance_esg.json: 12-query weak-label benchmark set. quality_benchmark.json: Policy quality benchmark outputs. latency_benchmark_combined.json: Combined-policy latency benchmark. latency_benchmark_conservative.json: Conservative-policy… See the full description on the dataset page: https://huggingface.co/datasets/oneryalcin/luxical-query-expansion-finance.textn<1K0 likes21 downloads7mo agoHugging Face28Darther /RCW_2025_Positive_Query_Pairs The Washington law Benchmark (WLB) Dataset Summary The Washington Law Benchmark (WLB) is a large-scale, synthetic dataset designed specifically to advance Legal Information Retrieval (IR) and Semantic Search. It bridges the critical "semantic gap" between natural language (how citizens, local governments, and plain-English users describe legal scenarios) and formal statutory legalese (how laws are actually written). The dataset contains hundreds of thousands of… See the full description on the dataset page: https://huggingface.co/datasets/Darther/RCW_2025_Positive_Query_Pairs.textsentence-similarity100K<n<1M0 likes19 downloads2mo agoHugging Face29xxxxdszz /personal-query-baby-products Personal Query: Baby Products This dataset contains personalized product search queries for the Baby_Products category. Each record is built from the Personal Query pipeline: Stage 6 generated correct personalized queries. Stage 7 injected user-specific error query variants when a matching error pattern was available. Stage 5 provided the user profile complexity level. Files data.jsonl: all correct Stage 6 queries. Rows without Stage 7 error query keep error_query as… See the full description on the dataset page: https://huggingface.co/datasets/xxxxdszz/personal-query-baby-products.tabulartext-generation10K<n<100K0 likes17 downloads5mo agoHugging Face30ValerioClm /query_routingtext1K<n<10K0 likes17 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.