datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.hleTwin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.millebitjawildtext
JaWildText
JaWildText is a Japanese scene text understanding benchmark for evaluating vision-language models (VLMs) on text-rich real-world images. It is designed to diagnose Japanese OCR, scene text visual question answering, and structured information extraction in practical conditions such as dense signboards, handwritten text, and mobile-captured receipts.
For details on the dataset construction, evaluation protocol, and baseline results, see the paper:
JaWildText: A Benchmark… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/jawildtext.RefinedVision
Sample page. This Hugging Face repository hosts only a 10-example-per-subset sample of RefinedVision, for quick browsing on the Hub. The full dataset is hosted on GitLab: https://gitlab.llm-jp.nii.ac.jp/datasets/refinedvision
RefinedVision
|
📄 Blog
|
RefinedVision is a refined version of HuggingFaceM4/FineVision that removes subsets with licensing or quality issues and regenerates QA pairs for low-quality subsets.
As a result, RefinedVision contains 124 of FineVision's… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/RefinedVision.Wiki-JA-Pair
Wiki-JA-Pair
Wiki-JA-Pair is a dataset of 1M image–text pairs sourced from the Japanese Wikipedia (Wiki-JA).
This dataset is constructed using the May 1, 2025 dump (jawiki-20250501-pages-articles-multistream.xml.bz2).
How to Use
from datasets import load_dataset
ds = load_dataset("llm-jp/Wiki-JA-Pair")
Format
Wiki-JA-Pair includes the following columns:
url: URL of the image
caption: Caption associated with the image
description: Nearby text that… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/Wiki-JA-Pair.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.or-bench
OR-Bench: An Over-Refusal Benchmark for Large Language Models
Please see our leaderboard at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.synth-text-recognition
Dataset Card for "Synth-Text Recognition"
This is the dataset for text recognition on document images, synthetically generated, covering 90K English words.
It includes training, validation and test splits.
JAMMEval
JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation
|
🤗 HuggingFace
|
📄 Paper
|
🧑💻 Code
|
Overview
JAMMEval is a curated benchmark collection for evaluating Vision-Language Models (VLMs) on Japanese Visual Question Answering (VQA) tasks.
It is constructed by refining seven existing Japanese VQA evaluation datasets through two rounds of human annotation, with the goal of improving evaluation reliability and quality.
⚠️… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/JAMMEval.Synth-JDoc
Synth-JDoc
Paper | Code
Synth-JDoc is a dataset of synthetic Japanese document images generated using HTML/CSS.
We generate embedded images, captions, and titles from prepared text, and use these elements to synthesize document images featuring diverse multi-column layouts in both vertical and horizontal writing.
Because the document images are synthesized directly from text, this dataset is completely free from OCR errors.
Dataset details
id
Image ID
image… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/Synth-JDoc.relaion2B-en-research-safe-japanese-translation
relaion2B-en-research-safe-japanese-translation
This dataset is the Japanese translation of the English subset of ReLAION-5B (laion/relaion2B-en-research-safe), translated by gemma-2-9b-it.
We used text2dataset for translating with open-weight LLMs.
By leveraging the fast LLM inference library vLLM, this tool enables the rapid translation of large English datasets into Japanese.
Prompt
The following is the prompt used for translation with Gemma.
You are an… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/relaion2B-en-research-safe-japanese-translation.or-bench-toxic-all
OR-Bench: An Over-Refusal Benchmark for Large Language Models
This dataset constains highly toxic prompts, use with caution!!!
Please see our demo at HuggingFace Spaces.
Overall Plots of Model Performances
Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.M-Attack_AdvSamples
M-Attack Adversarial Samples Dataset
This dataset contains 100 adversarial samples generated using M-Attack to perturb the images from the NIPS 2017 Adversarial Attacks and Defenses Competition. This dataset is used in the paper A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1.
Dataset Description
The dataset consists of total 300 adversarial samples organized in three… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-LLM/M-Attack_AdvSamples.HakushoBench
HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers
|
🤗 Dataset
|
📄 Paper
|
🧑💻 Code
|
Overview
HakushoBench is a challenging Japanese chart and table VQA benchmark built from 33 governmental white papers.
HakushoBench contains 2,053 images spanning over 10 image types, with manually annotated QA pairs, designed to assess deep and holistic understanding of charts and tables, rather than local… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/HakushoBench.CHML-real-ecp-local-llm-audit-benchmark
Links
GitHub: https://github.com/CHML-real/CHML-real-ecp-local-llm-audit-benchmark
Hugging Face Dataset: https://huggingface.co/datasets/CHML-real/CHML-real-ecp-local-llm-audit-benchmark
ECP Local LLM Multi-hop Audit Benchmark
Local-only benchmark suite for evaluating Evidence Confidence Propagation (ECP) as an audit layer for multi-hop reasoning chains.
This repository is designed for the CHML-real GitHub namespace and uses only local execution. The LLM… See the full description on the dataset page: https://huggingface.co/datasets/CHML-real/CHML-real-ecp-local-llm-audit-benchmark.puzzles-for-vision-llmft-llm-2026-qa-dataset
FT-LLM 2026 QA Dataset
A Japanese visual-question-answering dataset used for Stage 1-2 visual instruction tuning of the COMPASS Vision-Language Model. Each sample contains a document or natural image together with one or more Japanese question–answer pairs, and is designed to give the VLM its instruction-following and VQA capabilities. Images are embedded in the dataset, so no external downloads are required.
Part of the Compass collection.
License
Released under the… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-qa-dataset.AMALIA-VL-SFT-Dataset
AMALIA-VL-Training-Dataset
Dataset Description
This dataset is provided as part of the AMALIA project.
This is the vision+language training mix for AMALIA-VL-SFT. Each subset is one
source dataset in the mix, each with a single train split. The only datasets that are absent from this mix are those that derive directly from the core LLM training mix, and can be found in the AMALIA-LLM Post Training Collection.
Example usage:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/AMALIA-VL-SFT-Dataset.SEED-Bench-PT
SEED-Bench-PT
European Portuguese (pt-PT) machine translation of SEED-Bench, a multiple-choice benchmark spanning multiple dimensions of multimodal comprehension.
Translated from the original English test split using gemini-3.1-pro.
Original Dataset: https://huggingface.co/datasets/lmms-lab/SEED-Bench
Note: This dataset is machine translated and may contain translation errors or artifacts.
This dataset is provided as part of the AMALIA project and is included in… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/SEED-Bench-PT.llm.planning
LLM Planning Benchmark Datasets
This repository contains unified datasets used by the LLM-planning framework.
The files here are organized to match the current experiment entrypoint in scripts/exp.sh and the multi-stage planning pipeline used in the repo.
Datasets Overview
Dataset
Files / Folders
Samples
Notes
Augmented GAIA
4 category folders + DAG/reference folders
165 main eval samples
Multimodal answer-based benchmark with attachments, GPT-4o dependency… See the full description on the dataset page: https://huggingface.co/datasets/Alfiechuang/llm.planning.pokemon-llm-svg-bench-results
Pokémon LLM SVG Bench (V1.1)
This is a snapshot of scored SVG drawings from the Pokémon LLM SVG Bench — LLMs try to draw Pokémon in SVG, and we score the results. Prefer reading the live site for context; for scoring rules and methodology, see About.
Unofficial fan / research project. Pokémon images are copyrighted by The Pokémon Company and related companies involved in developing, operating, and managing the Pokémon series. Pokémon descriptions are sourced from… See the full description on the dataset page: https://huggingface.co/datasets/haxfenx/pokemon-llm-svg-bench-results.JMedQA
JMedQA: Benchmarking Large Language Models and Vision-Language Models on the Japanese Medical Licensing Examination
JMedQA is a Japanese medical question-answering benchmark derived from Japan's National Medical Examination materials publicly released by the Ministry of Health, Labour and Welfare (MHLW).
The dataset supports both text-only large language model (LLM) evaluation and vision-language model (VLM) evaluation using associated examination images.
Its image-dependency… See the full description on the dataset page: https://huggingface.co/datasets/SIP-med-LLM/JMedQA.WAON-Bench
WAON-Bench: Japanese Cultural Image Classification Dataset
|
🤗 HuggingFace
|
📄 Paper
|
🧑💻 Code
|
WAON-Bench is a manually curated image classification dataset designed to benchmark Vision-Language models on Japanese culture.
The dataset contains 374 classes across 8 categories (animals, buildings, events, everyday life, food, nature, scenery, and traditions), with 5 images per class, totaling 1,870 examples.
How to Use… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/WAON-Bench.ft-llm-2026-ocr-dataset
FT-LLM 2026 OCR Dataset
A Japanese document-OCR dataset used for Stage 1-1 caption + OCR pretraining of the COMPASS Vision-Language Model. Each sample pairs a rendered page image from a Japanese public-sector financial PDF (Cabinet Office, Financial Services Agency, Ministry of Finance) with its OCR-extracted markdown text. It is intended to teach the VLM's MLP projector to align vision tokens with Japanese text.
Part of the Compass collection.
License
Released under… See the full description on the dataset page: https://huggingface.co/datasets/Yana/ft-llm-2026-ocr-dataset.TextVQAJSSODa-test
JSSODa (test)
Paper | Code
This repository contains the test split of the JSSODa dataset.
Dataset details
JSSODa (Japanese Simple Synthetic OCR Dataset) is constructed by rendering Japanese text generated by an LLM into images.
The images contain text written both vertically and horizontally, which is organized into one to four columns.
This dataset was introduced in our paper: "Evaluating Multimodal Large Language Models on Vertically Written Japanese Text".
The code… See the full description on the dataset page: https://huggingface.co/datasets/llm-jp/JSSODa-test.multilingual_ocr_llm_2
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/tachiwin/multilingual_ocr_llm_2.LLMafia
LLMafia - Asynchronous LLM Agent
Our Mafia game dataset of an Asynchronous LLM Agent playing games of Mafia with multiple human players.
🌐 Project | 📃 Paper | 💻 Code
A virtual game of Mafia, played by human players and an LLM agent player. The agent integrates in the asynchronous group conversation by constantly simulating the decision to send a message.
Time to Talk: 🕵️♂️ LLM Agents for Asynchronous Group Communication in Mafia Games
Niv Eckhaus, Uri Berger, Gabriel… See the full description on the dataset page: https://huggingface.co/datasets/niveck/LLMafia.
