datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JieZi
JieZi (解字)
💻 Project ·
📦 Dataset ·
🌐 Demo
JieZi (解字) is a large-scale, expert-audited visual question answering (VQA) dataset dedicated to ancient Chinese character exegesis. It pairs high-quality character glyph images with fine-grained expert annotations across nine paleographic tasks—including headword recognition, etymology, structural analysis, glyph evolution, and component function—providing a rigorous benchmark for multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Ran0/JieZi.vsr_random
VSR: Visual Spatial Reasoning
This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper].
Usage
from datasets import load_dataset
data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"}
dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files)
Note that the image files still need to be downloaded separately. See data/ for details.
Go to our github repo for more introductions.
Citation
If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.cancer-knowledge-base
Cancer Knowledge Base — the open, verified oncology KB for RAG & LLM evaluation
The only open CC-BY-4.0 oncology knowledge base that combines:
110/110 trials cited with PMID + NCT + PubMed/ClinicalTrials.gov URLs, and 32 prognosis
rows linked to verified SEER 2016–2022 references — no LLM-synthetic dataset has this.
A provable 152-question MCQ benchmark — every answer derives from this KB's own structured
data and carries a citation + golden docs, so it is open-book verifiable… See the full description on the dataset page: https://huggingface.co/datasets/ranjithraj/cancer-knowledge-base.Java_method2test_chatml
Java Method to Test ChatML
This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}].
Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here:
To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters:
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.Mr-Ben
Intro
Welcome to the dataset page for the Meta-Reasoning Benchmark associated with our recent publication "Mr-Ben: A Comprehensive Meta-Reasoning Benchmark for Large Language Models". We have provided a demo evaluate script for you to try out benchmark in mere two steps. We encourage everyone to try out our benchmark in the SOTA models and return its results to us. We would be happy to include it in the eval_results and update the evaluation tables below for you.
• 📰 Mr-Ben… See the full description on the dataset page: https://huggingface.co/datasets/Randolphzeng/Mr-Ben.GSM-Ranges
GSM-Ranges Dataset
📄 Paper: Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges🔗 GitHub Repository: GSM-Ranges GitHub
What is GSM-Ranges?
GSM-Ranges is a dataset generator built upon the GSM8K benchmark. It systematically modifies numerical values in math word problems to assess the robustness of large language models (LLMs) across a broad spectrum of numerical scales. By introducing numerical… See the full description on the dataset page: https://huggingface.co/datasets/guactastesgood/GSM-Ranges.nepal-section-wise-act-datasets
Nepal Section-wise Act Datasets
Dataset Description
This dataset contains section-wise legal acts and laws of Nepal, organized for easy access and analysis. It is designed to support legal research, natural language processing (NLP) tasks, and the development of legal tech applications in Nepal.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use and may result in legal action or… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-section-wise-act-datasets.human_rank_eval
Dataset Card for HumanRankEval
This dataset supports the NAACL 2024 paper HumanRankEval: Automatic Evaluation of LMs as Conversational Assistants.
Dataset Description
Language models (LMs) as conversational assistants recently became popular tools that help people accomplish a variety of tasks. These typically result from adapting LMs pretrained on general domain text sequences through further instruction-tuning and possibly preference optimisation methods. The evaluation… See the full description on the dataset page: https://huggingface.co/datasets/huawei-noah/human_rank_eval.ransomware-playbooks-en
Ransomware Playbooks - English Dataset
A comprehensive bilingual (FR/EN) dataset containing detailed and operational information about ransomware groups, incident response playbooks, and structured Q&A on cybersecurity and ransomware attack response.
Overview
This dataset provides a complete resource for security teams, CISOs, and DFIR (Digital Forensics and Incident Response) professionals to understand, detect, and respond to modern ransomware attacks.
Main… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ransomware-playbooks-en.RankJudge
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
Abstract
As interactive LLM-based applications are created and refined, model developers need to evaluate the quality of generated text along many possible axes. For simpler systems, human evaluation may be practical, but in complicated systems like conversational chatbots, the amount of generated text can overwhelm human annotation resources. Model developers have begun to rely heavily on… See the full description on the dataset page: https://huggingface.co/datasets/Layer6/RankJudge.PersonaGen-Enterprise
PersonaGen-Enterprise: B2B Buying Intelligence Dataset
5,000 enterprise buyer personas with full buying committee modeling across 15 industries, 3 company sizes, and 42 buying roles. Plus 47K real search queries, 7.5K competitive brand queries, and multi-model agreement scores.
Built by Rankfor.AI, the AI Visibility Intelligence platform. This dataset powers research into how enterprise buyers search for, evaluate, and select B2B technology vendors.
Enterprise… See the full description on the dataset page: https://huggingface.co/datasets/rankfor/PersonaGen-Enterprise.alexa-qa-with-rankAlexa question and answer examples with rankMr-GSM8KView the project page:
https://github.com/dvlab-research/DiagGSM8K
see our paper at https://arxiv.org/abs/2312.17080
Description
In this work, we introduce a novel evaluation paradigm for Large Language Models,
one that challenges them to engage in meta-reasoning. Our paradigm shifts the focus from result-oriented assessments,
which often overlook the reasoning process, to a more holistic evaluation that effectively differentiates
the cognitive capabilities among models. For… See the full description on the dataset page: https://huggingface.co/datasets/Randolphzeng/Mr-GSM8K.ransomware-playbooks-fr
Ransomware Playbooks - Dataset Français
Un dataset complet et bilingue (FR/EN) contenant des informations détaillées et opérationnelles sur les groupes de ransomware, les playbooks de réponse aux incidents, et des Q&A structurées sur la cybersécurité et la réponse aux attaques ransomware.
Vue d'ensemble
Ce dataset fournit une ressource complète pour les équipes de sécurité, les CISOs, et les professionnels du DFIR (Digital Forensics and Incident Response) pour comprendre… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ransomware-playbooks-fr.Amazon-combined
Amazon Combined Dataset
E-commerce dataset that combines metadata, reviews, and sample question/answer pairs. combined.json contains the dataset and user2asin.json contains a file that maps user_id from reviews to an ASIN for capturing user preferences.
Data Fields
Field
Type
Explanation
main_category
str
Main category (i.e., domain) of the product.
title
str
Name of the product.
average_rating
float
Rating of the product shown on the product page.… See the full description on the dataset page: https://huggingface.co/datasets/randomath/Amazon-combined.MMLU_ExpertPrompt_Random_01This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.nepal-constitution-dataset
Nepal Constitution Dataset
Dataset Description
This dataset contains the Constitution of Nepal (२०७२), organized section-wise for easy access, analysis, and use in NLP and legal tech applications. It is designed to support legal research, educational purposes, and the development of AI-driven tools for the Nepali legal system.
Note: This dataset is released for research purposes only. Any other unwanted use can lead to the violation of the intended terms of use… See the full description on the dataset page: https://huggingface.co/datasets/ranjitraut/nepal-constitution-dataset.MutQA
CrossValQA
CrossValQA is a cross-validated, sentence-grounded question-answering
dataset about genetic mutations, constructed from full-text PubMed articles.
Every record links a natural-language question to a specific variant, a
specific PubMed article, and a specific cited sentence span, and every
answer was produced by two independent LLMs that had to agree before the
record was admitted.
The released train/test splits (homology and random configs)
contain only cross-grounded… See the full description on the dataset page: https://huggingface.co/datasets/random987654321/MutQA.MMLU_ExpertPrompt_Random_03This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.GSM8K-Random-All
GSM8K-Random-All
A dataset for training LLMs with random backtracking capabilities. This dataset augments the original GSM8K math word problems with synthetic error injection and backtrack recovery sequences.
Overview
This dataset teaches models to:
Make "mistakes" (random error tokens)
Recognize the mistake
Use <|BACKTRACK|> tokens to "delete" the errors
Continue with the correct solution
Backtracking Mechanism
The <|BACKTRACK|> token functionally acts as a… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/GSM8K-Random-All.ChronoCoexist-WD
ChronoCoexist-WD
Multi-Version Temporal Knowledge Timelines
ChronoCoexist-WD (technical dataset identity: PTC-WD-2026) is an automatically
curated source pool of real temporal
Wikidata histories. Each timeline keeps one subject and one relation fixed
while the answer changes across dated, non-overlapping intervals. It supports
research on whether an edited language model can learn a newly arriving value
without losing earlier values that remain correct for… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/ChronoCoexist-WD.MMLU_ExpertPrompt_Random_02This is a massive multitask test consisting of multiple-choice questions from various branches of knowledge, covering 57 tasks including elementary mathematics, US history, computer science, law, and more.2026-landlord-vs-tenant-state-rankings-index
SoldFast Landlord vs. Tenant State Rankings Index 2026
Version: v2026.1 | Last updated: February 2026
A structured legal and financial analysis of landlord-tenant law across all 50 US states, with 211 citation-backed question-answer pairs suitable for instruction tuning, RAG evaluation, and legal/real estate domain benchmarking.
Dataset Summary
This dataset contains the complete 2026 Landlord vs. Tenant State Rankings Index published by SoldFast. Each of the 50 US… See the full description on the dataset page: https://huggingface.co/datasets/ryancdossey1/2026-landlord-vs-tenant-state-rankings-index.RankJudge
RankJudge
Anonymous submission. Under double-blind review at NeurIPS 2026 Datasets & Benchmarks.
Authors and affiliations have been removed. The dataset will be re-released under the authors' names after the review period.
RankJudge is a benchmark for evaluating LLM-as-judge on multi-turn conversation quality, grounded in verifiable reference material. For each source item (an academic paper or financial filing with reference QA pairs), the pipeline generates a pair of… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-rankjudge-neurips/RankJudge.MMLU_ExpertPrompt_Randomwdb-islamic-finance-benchmark
WDB Benchmark: Western Default Bias in Islamic Finance
Dataset Description
This benchmark tests whether Large Language Models exhibit Western Default Bias (WDB) - the tendency to provide Western/conventional finance answers even when the context implies Islamic finance should be used.
The Problem
When a user in Saudi Arabia or UAE asks a financial question, they likely expect Shariah-compliant advice. However, LLMs trained predominantly on Western data may… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/wdb-islamic-finance-benchmark.rank-with-reasonSiQuAD
Sinhala SQuAD Dataset
This dataset is a translation of the SQuAD v1.0 dataset into Sinhala using the Google Translate API. It consists of 16,000 question-answer pairs, with 13,000 training pairs and 1,250 test/dev pairs. The dataset is cleaned and validated to ensure the quality of the translations.
Dataset Details
Size: 16,000 QA pairs
Train: 13,000 pairs
Test/Dev: 1,250 pairs
Language: Sinhala
Source: The dataset was derived from the original SQuAD v1.0 dataset.… See the full description on the dataset page: https://huggingface.co/datasets/janani-rane/SiQuAD.MedicalQnA-llama2
MedicalQnA-llama2 Dataset
This repository contains the MedQuad-MedicalQnADataset specifically formatted for use with the LLaMA 2 prompt template. The dataset consists of medical questions categorized by question type, along with their corresponding answers. It is designed for text-to-text generation and text-generation tasks, particularly focusing on the medical domain.
Dataset Structure
The dataset is structured to include prompts that follow the LLaMA 2 template. Each… See the full description on the dataset page: https://huggingface.co/datasets/randomani/MedicalQnA-llama2.wdb-islamic-finance-benchmark2
WDB Benchmark: Western Default Bias in Islamic Finance
Dataset Description
This benchmark tests whether Large Language Models exhibit Western Default Bias (WDB) - the tendency to provide Western/conventional finance answers even when the context implies Islamic finance should be used.
How to Load
from datasets import load_dataset
ds = load_dataset("Raniahossam33/wdb-islamic-finance-benchmark")
# Access data
for sample in ds["train"]:… See the full description on the dataset page: https://huggingface.co/datasets/Raniahossam33/wdb-islamic-finance-benchmark2.
