datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
forbidden_question_set
Forbidden Question Set
This is the Forbidden Question Set dataset proposed in the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.
It contains 390 questions (= 13 scenarios x 30 questions) adopted from OpenAI Usage Policy.
We exclude Child Sexual Abuse scenario from our evaluation and focus on the rest 13 scenarios, including Illegal Activity, Hate Speech, Malware Generation, Physical Harm, Economic Harm… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/forbidden_question_set.spacr_settingsleetcode-problem-set
LeetCode Scraper Dataset
This dataset contains information scraped from LeetCode. It is designed to assist developers in analyzing LeetCode problems, generating insights, and building tools for competitive programming or educational purposes.
Dataset Contents
The dataset includes the following files:
problem_set.csv
Contains a list of LeetCode problems with metadata such as difficulty, acceptance rate, tags, and more.
Columns:
acRate: Acceptance rate of the… See the full description on the dataset page: https://huggingface.co/datasets/kaysss/leetcode-problem-set.landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.URSA-benchmarking-sets
URSA benchmarking sets
Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
Reaction plausibility
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/insilicomedicine/URSA-benchmarking-sets.Dr.Sparse-OTF-test-set
Dr.Sparse OTF Test Set
100 sparse matrices from the SuiteSparse Matrix Collection,
converted to the flat binary format the Dr.Sparse
benchmark harness reads. This is the held-out evaluation set for LLM-generated
CUDA sparse kernels (SpMV / SpMM / SpGEMM), kept separate from the matrices the
models were developed against.
Layout
Matrices are grouped into size tiers by row count, the convention Dr.Sparse task
discovery scans for:
tier
rows
matrices
size… See the full description on the dataset page: https://huggingface.co/datasets/KinGeorge/Dr.Sparse-OTF-test-set.maritime-bill-of-lading-document-set-coherence-risk-v0.1What this repo is for
Triage trade doc packs before they trigger holds.
You use it to flag
HS code inconsistencies across documents
missing certificates
shipper or consignee mismatch
clearance status lag not supported by doc quality
Why it matters
Most port delay disputes begin in paperwork.
setimes-en-tr-aligned-corpus
SETimes EN-TR — Sentence-Aligned, LLM-Cleaned
A cleaned and re-aligned version of the SETimes English-Turkish parallel corpus. The original SETimes data is paragraph-style — each "pair" can contain a headline, a dateline, several body sentences, and a source citation, all glued together on one line. This version splits everything into proper sentence pairs so each row is one English sentence next to its Turkish translation.
144,064 sentence pairs, split into train (142,064)… See the full description on the dataset page: https://huggingface.co/datasets/atahanuz/setimes-en-tr-aligned-corpus.astex_diverse_setThe Astex Diverse dataset accompanying the PoseBench manuscript and benchmarking suite.
SNOMED-CT-Code-Value-Semantic-Set.csvSNOMED-CT-Code-Value-Semantic-Set.csv
historical-topo-cacheYouTube-Evaluation-Set
Awaaz se Alfaaz — YouTube Evaluation Set
This dataset is the realistic multi-speaker evaluation set used in Awaaz se Alfaaz, accepted at LaTeLL 2026 — "Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing." It contains 30 short-form Urdu YouTube videos (YouTube Shorts) covering a mix of news, sports, and current affairs content, along with human annotated gold transcripts and transcripts produced by… See the full description on the dataset page: https://huggingface.co/datasets/awaaz-se-alfaaz/YouTube-Evaluation-Set.GMASS-probe-set-v1.0
MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages
Project Summary
We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.Gender_Bias_Evaluation_SetThis dataset has been created as part of the Flax/JAX community week for testing the flax-sentence-embeddings Sentence Similarity models for Gender Bias but can be used for other use-cases as well related to evaluating Gender Bias.
The Following Dataset has been created for Evaluating Gender Bias for different models, based on various stereotypical occupations.
The Structure of the dataset is of the following type:
Base Sentence
Occupation
Steretypical_Gender
Male Sentence
Female… See the full description on the dataset page: https://huggingface.co/datasets/flax-sentence-embeddings/Gender_Bias_Evaluation_Set.dockgen_setThe DockGen-E dataset accompanying the PoseBench manuscript and benchmarking suite.
Pharmacology-LLM-test-setPharmacology-LLM-test-set: A test set for a large language model focused on pharmacology tasks
1 Inroduction
Large language models (LLM), including ChatGPT, have fundamentally transformed the knowledge query schemes and methods in pharmacology for pharmacologists, drug researchers, clinical drug researchers, and artificial intelligence researchers in pharmacology. They can conduct multi-round consultations and query pharmacological issues in a question-and-answer format. However… See the full description on the dataset page: https://huggingface.co/datasets/zhangyingbo1984/Pharmacology-LLM-test-set.QARV-binary-setThe QARV (Question and Answers with Regional Variance) project aims to curate a collection of questions with answers that exhibit regional variations across different nations.
URSA-benchmarking-sets
URSA benchmarking sets
Benchmark for reaction plausibility, single-step and multistep retrosynthesis from Zagribelnyy et al. (2026).
It bundles two task families: reaction-level plausibility judgment and retrosynthesis target sets.
Reaction plausibility
URSA-reaction-plausibility-bench-2026.csv — 1,000 reactions predicted by different models hand-labeled by expert oragnic chemists (500 plausible / 500 implausible), for benchmarking reaction-level… See the full description on the dataset page: https://huggingface.co/datasets/introvoyz041/URSA-benchmarking-sets.past-setting-ede8fc
past-setting-ede8fc
Synthetic sensors test data: 51 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/randallkaren9/past-setting-ede8fc.multisource-esco-set
MultiSource-ESCO-Skills: A Unified Dataset for Skill Extraction
This dataset aggregates data from multiple sources—course descriptions, CV content, and job descriptions—all linked to ESCO skills. It is designed to help researchers and practitioners develop and fine-tune NLP models (e.g., BERT or SentenceTransformer-based models) for automated skill extraction.
Dataset Overview
Name: MultiSource-ESCO-Skills
Sources:
Course Content: Educational course materials
CV Content:… See the full description on the dataset page: https://huggingface.co/datasets/Boanerges/multisource-esco-set.payments-authorization-settlement-coherence-risk-v0.1What this repo is for
Detect when payment approvals no longer match settlement reality.
Focus
approval vs settlement
funding gaps
latency mismatches
Why it matters
Payments systems fail when authorization and settlement drift apart.
settled-live-game-rounds
1GO Results Settled Live-Game Rounds
A versioned dataset of 31,214 completed live-game rounds published by
1GO Results. This release covers seven games over a
fixed 72-hour UTC observation window and includes normalized UTC and India
Standard Time timestamps.
This is a historical-results dataset, not a prediction system. Past random
outcomes do not predict or guarantee future outcomes, and this dataset does not
provide betting advice.
Release
v1.0.0-72h-20260719… See the full description on the dataset page: https://huggingface.co/datasets/1goresults/settled-live-game-rounds.casp15_setThe CASP15 dataset accompanying the PoseBench manuscript and benchmarking suite.
sahkar-seva-setu-demandhate_speech_open_data_original_class_test_setsetfit-absa-tesla-tweetsinterview_QA_sample_set
Huy Interview Instruction Dataset
Dataset Description
This is an instruction-answer dataset for fine-tuning conversational AI models to answer interview-style questions based on a personal CV/profile.
The dataset has two columns: instruction and answer.
The dataset contains 5,100 instruction-answer pairs.
Data Creation
This dataset was created using a GenAI-assisted pipeline. A personal CV/profile was provided as source material, and GenAI was used… See the full description on the dataset page: https://huggingface.co/datasets/dinhxuanhuy/interview_QA_sample_set.clinical-moca-minimal-causal-set-identification-v0.1What this dataset tests
Whether a model can identify the smallest causal setthat still explains the full clinical + multi-omic picture.
It penalizesadditive hit lists.
It rewardsminimal sets with coverage.
Data format
Each row includes
longitudinal omics summary
clinical narrative
candidate causal sets
selected set with coverage map
Labels
minimal-and-sufficient
minimal-but-insufficient
sufficient-but-nonminimal
neither
Typical failures
choosing the shortest set that… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-moca-minimal-causal-set-identification-v0.1.bengali-visual-genome-instruction-setmalayalam-visual-genome-instruction-set
