datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.Twin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.JMedQA
JMedQA: Benchmarking Large Language Models and Vision-Language Models on the Japanese Medical Licensing Examination
JMedQA is a Japanese medical question-answering benchmark derived from Japan's National Medical Examination materials publicly released by the Ministry of Health, Labour and Welfare (MHLW).
The dataset supports both text-only large language model (LLM) evaluation and vision-language model (VLM) evaluation using associated examination images.
Its image-dependency… See the full description on the dataset page: https://huggingface.co/datasets/SIP-med-LLM/JMedQA.ft-llm-2026-domain-specific-qa
FT-LLM 2026 Domain-Specific QA
A Japanese financial-domain visual QA dataset used for Phase 3 domain fine-tuning of the COMPASS Vision-Language Model. Question–answer pairs were generated with Qwen3-VL from scraped Japanese government financial PDFs (Cabinet Office, Financial Services Agency, Ministry of Finance), covering four difficulty tiers: (A) numeric extraction, (B) rate-of-change & comparison, (C) financial formula application, and (D) complex reasoning. Each answer includes… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-domain-specific-qa.Fraud-R1-LLM-Defense-Fraud-Benchmark
Fraud-R1 : A Comprehensive Benchmark for Assessing LLM Robustness Against Fraud and Phishing Inducement
Shu Yang*, Shenzhe Zhu*, Zeyu Wu, Keyu Wang, Junchi Yao, Junchao Wu, Lijie Hu, Mengdi Li, Derek F. Wong, Di Wang†
(*Contribute equally, †Corresponding author)
😃 Github | 📜 Project Page | 📝 arxiv
❗️Content Warning: This repo contains examples of harmful language.
📰 News
2025/02/16: ❗️We have released our evaluation code.
2025/02/16: ❗️We have released our dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Chouoftears/Fraud-R1-LLM-Defense-Fraud-Benchmark.MMMU-Pro-PT
MMMU-Pro-PT
European Portuguese (pt-PT) machine translation of MMMU-Pro, a more robust and challenging version of MMMU for college-level, multi-discipline multimodal reasoning.
Translated from the original English test split (standard (10 options) subset) using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/MMMU/MMMU_Pro (standard (10 options) subset)
Note: This dataset is machine translated and may contain translation errors or artifacts.
This… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MMMU-Pro-PT.MATH-Vision-PT
MATH-Vision-PT
European Portuguese (pt-PT) machine translation of MATH-Vision, a benchmark of competition-level mathematics problems presented in visual contexts.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/MathLLMs/MathVision
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MATH-Vision-PT.MMStar-PT
MMStar-PT
European Portuguese (pt-PT) machine translation of MMStar, a curated multimodal benchmark of vision-indispensable, balanced challenge samples.
Translated from the original English val split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/Lin-Chen/MMStar
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in amalia-vl-eval, a… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/MMStar-PT.cxr_llm
Dataset Card for Dataset Name
CXR for medical multimodal LLMs
Dataset Summary
This dataset aims to provide medical conversations with contextual images. Most of the questions ensure the model understands what anomalies are present within the CXR images. \
There are also follow-up questions to teach the LLM how to follow up after identifying the anomaly.
There is a total of 104892 human-bot conversations with contextual images \
50,021 images from chexpert 5,229 images… See the full description on the dataset page: https://huggingface.co/datasets/cheese111/cxr_llm.
