datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
query-classification-pakistani-legal-vs-nonlegal
Pakistani Legal Query Classification Dataset
A binary classification dataset to distinguish legal queries from
non-legal queries, built for the PakLegalAid project and the paper:
"Enhancing Legal Assistance with Large Language Models: A Parameter-Efficient
Fine-Tuning and Retrieval-Augmented Generation Approach"Umair Ahmed, Sher Muhammad Daudpota, Ali Shariq Imran, Zenun Kastrati,
Muhammad Nabeel — submitted to PLOS ONE, 2025.
Dataset Description
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/heyIamUmair/query-classification-pakistani-legal-vs-nonlegal.query-classification
Dataset Card for Query Classification
Dataset Summary
Query Classification is a dataset of 21,627 English queries categorized into 19 distinct classes. The dataset was created by translating Chinese search queries into English using machine translation, making it suitable for training and evaluating text classification models for query intent detection.
Supported Tasks and Leaderboards
Text Classification: Classify queries into one of 19… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/query-classification.query-classification
