datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
query-classification-dataset
Multi-lingual Query Scope & Intent Classification Dataset (addyo07/query-classification-dataset)
This dataset provides ground-truth query classification samples across multiple languages (English, Devanagari Hindi, Code-switched Hinglish) for intent and memory scope routing in conversational AI and AI assistant engines.
📦 Version History
v2/memory_scope_golden_v1.json (4-Class Multi-Lingual Scope Taxonomy)
Total Verified Samples: 22,006 items… See the full description on the dataset page: https://huggingface.co/datasets/addyo07/query-classification-dataset.task673_google_wellformed_query_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task673_google_wellformed_query_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task673_google_wellformed_query_classification.FHIR_QnA_Query-Based_Resource_Relevance_Classification_T1_complete
Dataset Card
This repository contains the dataset introduced in the paper Question Answering on Patient Medical Records with Private Fine-Tuned LLMs.
Query_Domain_Classificationquery-domain-classification-sharegptquery-classification-pakistani-legal-vs-nonlegal
Pakistani Legal Query Classification Dataset
A binary classification dataset to distinguish legal queries from
non-legal queries, built for the PakLegalAid project and the paper:
"Enhancing Legal Assistance with Large Language Models: A Parameter-Efficient
Fine-Tuning and Retrieval-Augmented Generation Approach"Umair Ahmed, Sher Muhammad Daudpota, Ali Shariq Imran, Zenun Kastrati,
Muhammad Nabeel — submitted to PLOS ONE, 2025.
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/heyIamUmair/query-classification-pakistani-legal-vs-nonlegal.query-classification
Dataset Card for Query Classification
Dataset Summary
Query Classification is a dataset of 21,627 English queries categorized into 19 distinct classes. The dataset was created by translating Chinese search queries into English using machine translation, making it suitable for training and evaluating text classification models for query intent detection.
Supported Tasks and Leaderboards
Text Classification: Classify queries into one of 19… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/query-classification.query-domain-classification-sharegpt-v2query-classificationflan_combined_task673_google_wellformed_query_classificationset-get-query-classificationrag_domain_query_classification
