datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MA_Query_Expansion_MLT26entity-query-summarizationmultimodal_query_rewrites
ReVision: Visual Instruction Rewriting Dataset
Dataset Summary
The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises:
Image: A visual scene containing relevant information.
Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/hsiangfu/multimodal_query_rewrites.query_builder_text_to_sqltpch-query-routing-dataset
Cross-Engine TPC-H Cost Modeling Dataset
📘 Overview
This dataset contains execution-time measurements and structural query features for SQL workloads executed across multiple database engines. It is designed for research in learned cost modeling, cross-engine optimization, and zero-shot SQL engine selection.
The dataset is generated using the TPC-H benchmark at Scale Factor .5 and includes automatically generated query variants for all 22 benchmark queries. Each query… See the full description on the dataset page: https://huggingface.co/datasets/Rinil-Parmar/tpch-query-routing-dataset.query_picker_datasetquery-hard-pos-neg-doc-pairs-statictablemultimodal_query_rewrites
ReVision: Visual Instruction Rewriting Dataset
Dataset Summary
The ReVision dataset is a large-scale collection of task-oriented multimodal instructions, designed to enable on-device, privacy-preserving Visual Instruction Rewriting (VIR). The dataset consists of 39,000+ examples across 14 intent domains, where each example comprises:
Image: A visual scene containing relevant information.
Original instruction: A multimodal command (e.g., a spoken query referencing visual… See the full description on the dataset page: https://huggingface.co/datasets/anonymoususerrevision/multimodal_query_rewrites.query-to-query-sts
Dataset Summary
Query to Query STS (Query2Query) is a Persian (Farsi) dataset for the Semantic Textual Similarity (STS) task. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was created from real-world anonymized queries issued to the Zarrebin search engine. Semantic similarity scores between query pairs are based on the overlap of their respective search result sets.
Language(s): Persian (Farsi)
Task(s): Semantic Textual Similarity (STS)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/query-to-query-sts.clinical-quad-data-integrity-query-backlog-missingness-governance-threshold-v0.1Clarus Clinical Quad Coupling Data Integrity Query Backlog Missingness Governance Threshold v0.1
What this dataset isThis dataset tests whether a model can detect clinical trial data integrity events driven by four interacting nodes.
Quad coupling nodes
Query backlog or data flow delay
Missingness in critical fields or attachments
Conmed or exposure timeline gaps
Governance thresholds such as audits, CAPA, freeze deadlines, or reporting cadence
Input
One vignette in prompt… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-integrity-query-backlog-missingness-governance-threshold-v0.1.clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1Clinical Quad Data Cut Timing Database Lock Pressure Query Backlog CSR Narrative Drift v0.1
Each row is a trial monthly snapshot.
Core quad
Data cut timingDatabase lock pressureQuery backlogCSR narrative drift
Target
label_regulatory_issue_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-data-cut-timing-database-lock-pressure-query-backlog-csr-narrative-drift-v0.1.rdf-query-based-summarizationbanking_customer_service_query_intent
Description
Synthetic dataset comprising of query:intent pairs. Created using AI prompts. Covers professional as well as casual speech.
manga-queryovos-common-query-intentsDataset focusing on general knowledge questions, it distinguishs between queries aimed at a specific information source (wikipedia, wolfram alpha, duckduckgo..) and non specific queries
astro-llms-full-query-data
AstroLLMs Full Query Dataset
This dataset includes all of the data collected in a four-week deployment of a Large Language Model-powered Slack chatbot trained on astrophysics papers. Astronomers were invited to interact with the chatbot, ask questions, and leave feedback. This data includes 368 question-answer pairs, including feedback, reactions, and labeling.
Dataset Structure
The columns of this dataset are thread_ts (unique time stamp of the query), channel_id… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/astro-llms-full-query-data.cspm_posture_checks_name_to_querycomplex-query-answering-datasethotel-query-description-ru
tourism-hotel-search-ru
Dataset Overview
tourism-hotel-search-ru is a specialized Russian-language dataset designed for training and evaluating Information Retrieval and Sentence Similarity models in the tourism and hospitality domain. It includes 14689 pairs query->hotel's description for 977 hotels.
This dataset is optimized for training models using loss functions such as MultipleNegativesRankingLoss (MNRL). The core objective is to match a natural language user… See the full description on the dataset page: https://huggingface.co/datasets/Olegator47/hotel-query-description-ru.clinical-quad-data-cut-query-backlog-database-lock-decision-error-v0.1Clinical Quad Data Cut Query Backlog Database Lock Decision Error v0.1
Each row is a data cut snapshot.
Core quad
Data cut timingQuery backlogDatabase lock pressureDecision error risk
Target
label_wrong_call_next_30d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
enetity-query-summarization-bQMSum_gt_w_queryquery2query_evaluation
Description
This dataset contains triples of the form "query1", "query2", "label" where labels are mapped as follows
similar: 1
not similar: 0
ambiguous: -1
SQuALITY_sparse_mean_w_queryQueryAnswerPairsQuery_Dataquery-classification-pakistani-legal-vs-nonlegal
Pakistani Legal Query Classification Dataset
A binary classification dataset to distinguish legal queries from
non-legal queries, built for the PakLegalAid project and the paper:
"Enhancing Legal Assistance with Large Language Models: A Parameter-Efficient
Fine-Tuning and Retrieval-Augmented Generation Approach"Umair Ahmed, Sher Muhammad Daudpota, Ali Shariq Imran, Zenun Kastrati,
Muhammad Nabeel — submitted to PLOS ONE, 2025.
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/heyIamUmair/query-classification-pakistani-legal-vs-nonlegal.non_consent_query_detection_v01
Non-Consent Query Detection (NCQD)
This dataset benchmarks a model’s ability to identify when a prompt requests action that violates consent, ethical boundaries, or implicit relational safety.
The focus is not moralizing or lecturing.The correct behavior is a structurally clean refusal or boundary call-out.
It tests whether a model can detect:
non-consensual intent
coercive framing
concealed manipulation
involuntary emotional labor requests
illegal / unethical direction masked as… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/non_consent_query_detection_v01.query-based-summarizationquery_optimization
