datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leetcode-problem-set
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
Dataset Contents
The dataset includes the following files:
problem_set.csv
Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more.
Columns:
acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.GMASS-probe-set-v1.0
MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages
Project Summary
We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.interview_QA_sample_set
Huy Interview Instruction Dataset
Dataset Description
This is an instruction-answer dataset for fine-tuning conversational AI models to answer interview-style questions based on a personal CV/profile.
The dataset has two columns: instruction and answer.
The dataset contains 5,100 instruction-answer pairs.
Data Creation
This dataset was created using a GenAI-assisted pipeline. A personal CV/profile was provided as source material, and GenAI was used… See the full description on the dataset page: https://huggingface.co/datasets/dinhxuanhuy/interview_QA_sample_set.llm-answer-set-qa
Answer-Set Consistency Benchmark (ASCB)
Overview
The Answer-Set Consistency Benchmark (ASCB) evaluates whether language models provide mutually consistent answers to related factual enumeration questions. Unlike conventional QA datasets, ASCB focuses on whether generated answer sets satisfy known set-theoretic relations rather than solely on factual accuracy.
ASCB contains 600 English question quadruples (2,400 questions) across primarily static, objective factual… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026nips/llm-answer-set-qa.MultiState-DMV-Licensing-Practice-Set
Introduction
This dataset contains structured practice questions and answers derived from multiple DMV sample materials. It is designed to support training and evaluation of models for multiple-choice question answering, rule-based classification, and natural language understanding tasks in the driving regulations domain.
Key Features
Region-Specific: Focused on multiple state driving laws and traffic rules
Use Case: DMV permit/license test practice modeling… See the full description on the dataset page: https://huggingface.co/datasets/nprak26/MultiState-DMV-Licensing-Practice-Set.
