datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMLU-SR
MMLU-SR Dataset
This is the dataset for the paper "MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models".
Dataset Structure
This dataset contains three different variants:
Question Only: Key terms in questions are replaced with dummy words and their definitions, while answer choices remain unchanged.
Answer Only: Key terms in answer choices are replaced with dummy words and their definitions, while questions remain unchanged.
Question… See the full description on the dataset page: https://huggingface.co/datasets/NiniCat/MMLU-SR.CountQA
Dataset Summary
CountQA is the new benchmark designed to stress-test the Achilles' heel of even the most advanced Multimodal Large Language Models (MLLMs): object counting. While modern AI demonstrates stunning visual fluency, it often fails at this fundamental cognitive skill, a critical blind spot limiting its real-world reliability.
This dataset directly confronts that weakness with over 1,500 challenging question-answer pairs built on real-world images, hand-captured to feature… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/CountQA.SR-3D-Bench
Spatial Region 3D (SR-3D) Aware Benchmark
Paper: https://arxiv.org/abs/2509.13317Project page: https://www.anjiecheng.me/sr3dCode: https://github.com/AnjieCheng/SR-3D
[!IMPORTANT]
[Feb. 18, 2026] UPDATE: To improve compatibility with general-purpose VLMs, the benchmark is reformulated into multiple-choice and numerical questions following the VSI-Bench evaluation protocol. Videos are annotated with set-of-marks to explicitly indicate regions. The benchmark will be compatible… See the full description on the dataset page: https://huggingface.co/datasets/a8cheng/SR-3D-Bench.HueManity
HueManity: A Benchmark for Testing Human-Like Visual Perception in MLLMs
Paper | Code
Dataset Description
HueManity is a benchmark dataset featuring 83,850 images designed to test the fine-grained visual perception of Multimodal Large Language Models (MLLMs). Each image presents a two-character alphanumeric string embedded within Ishihara-style dot patterns, challenging models to perform precise pattern recognition in visually cluttered environments.
The dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/HueManity.dnd5e-srd-qa
D&D 5.2.1 SRD RAG Evaluation Dataset
A high-quality Question-Answering (QA) dataset built by the Datapizza AI Lab from the Dungeons & Dragons 5th Edition System Reference Document (SRD) version 5.2.1, designed to evaluate Retrieval Augmented Generation (RAG) systems.
Dataset Summary
This dataset contains 56 question-answer pairs across two difficulty tiers (Easy and Medium), each designed to test different aspects of RAG system capabilities. The dataset is built from 20… See the full description on the dataset page: https://huggingface.co/datasets/datapizza-ai-lab/dnd5e-srd-qa.osm-tokyo23-src-2026-08
osm-tokyo23-src-2026-08
A frozen cut of OpenStreetMap covering the 23 special wards of Tokyo, taken
from the planet file of 2026-08-31, together with everything needed to
rebuild the databases it was measured in.
The point is the freezing. A question about a city has an answer only against
a stated snapshot, and an answer computed today against the live API is not
reproducible tomorrow. Here the snapshot is one file with a checksum, and the
tools that read it are pinned by… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-tokyo23-src-2026-08.Performance-Marketing-Data
Performance Marketing Expert Dataset
Dataset Description
This dataset contains comprehensive performance marketing knowledge and logical reasoning patterns for Meta (Facebook/Instagram), Google Ads, and TikTok advertising platforms. It's designed for fine-tuning language models to understand brand verticals, performance marketing strategies, and develop reasoning capacity for creating winning ad campaigns.
Dataset Structure
Each example follows an… See the full description on the dataset page: https://huggingface.co/datasets/Sri-Vigneshwar-DJ/Performance-Marketing-Data.UnifiedInstruct-709k
Mixed Alpaca Math Code Science Instruct
A mixed Alpaca-style instruction dataset containing math, code, science, commonsense, and general instruction examples.
The dataset is intended for supervised fine-tuning and instruction-tuning experiments, especially for small language models for educational purpose. :)
Dataset Splits
Split
Examples
Train
600,000
Validation
54,747
Test
54,748
Sources and Credits
This dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/srmty/UnifiedInstruct-709k.srsRANBench
srsRANBench: A Benchmark for Assessing LLMs in srsRAN Code Understanding
Overview
srsRANBench is a comprehensive benchmark dataset designed to evaluate Large Language Models (LLMs) in the context of code generation and code understanding for the srsRAN project. This benchmark consists of 1,502 multiple-choice questions, carefully curated by randomly selecting C++ files from the entire srsRAN codebase.
The benchmark assesses LLMs' ability to generate syntactically and… See the full description on the dataset page: https://huggingface.co/datasets/prnshv/srsRANBench.superglue-sr
BalkanBench SuperGLUE - Serbian
Part of BalkanBench - the open, reproducible
benchmark for language models across Serbian, Croatian, Montenegrin, and
Bosnian (BCMS). Live leaderboard at https://balkanbench.com/leaderboard.
Background and motivation: Release of BalkanBench - the vision behind it
(Medium, 2026-04-27).
This is the Serbian SuperGLUE track of BalkanBench v0.1. Serbian is the
official frozen track: the leaderboard's ranked average is computed over
6 ranked tasks… See the full description on the dataset page: https://huggingface.co/datasets/permitt/superglue-sr.cab
Dataset Card for CAB
Dataset Summary
The CAB dataset (Counterfactual Assessment of Bias) is a human-verified dataset designed to evaluate biased behavior in large language models (LLMs) through realistic, open-ended prompts.Unlike existing bias benchmarks that often rely on templated or multiple-choice questions, CAB consists of more realistic chat-like counterfactual questions automatically generated using an LLM-based framework.
Each question contains counterfactual… See the full description on the dataset page: https://huggingface.co/datasets/eth-sri/cab.Farsight-SRV-Transcripts
The Farsight Institute: Scientific Remote Viewing (SRV) Transcripts
Dataset Summary
This dataset contains the complete, unabridged archive of Scientific Remote Viewing (SRV) session transcripts and project summaries produced by The Farsight Institute, directed by Dr. Courtney Brown.
The data consists of hundreds of highly detailed, text-rich transcripts describing historical events, planetary mysteries, and extraterrestrial dynamics. All remote viewing sessions… See the full description on the dataset page: https://huggingface.co/datasets/courtnoski/Farsight-SRV-Transcripts.big-red-bark-chat-evaluation
Big Red Bark Chat Q&A Dataset
Dataset Description
This dataset contains 12,385 question-and-answer pairs collected from Big Red Bark Chat, an innovative AI assistant developed at Cornell University that answers questions about dog health (as well as other animal species). While it does not replace professional veterinary advice, it serves as a valuable starting point by searching trusted sources. Big Red Bark Chat is designed to provide quick and reliable answers… See the full description on the dataset page: https://huggingface.co/datasets/Sr523/big-red-bark-chat-evaluation.semeval-2016-absa-reviews-english-translated-stanford-alpaca
Dataset Card for Dataset Name
Derived from eastwind/semeval-2016-absa-reviews-arabic using Helsinki-NLP/opus-mt-tc-big-ar-en
osm-japan-src-2026-08
osm-japan-src-2026-08
A frozen cut of OpenStreetMap covering the whole of Japan, taken from the
planet file of 2026-08-31, together with everything needed to rebuild the
databases it was measured in.
The point is the freezing. A question about a place has an answer only against
a stated snapshot, and an answer computed today against the live API is not
reproducible tomorrow. Here the snapshot is one file with a checksum, and the
tools that read it are pinned by version.
This is… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/osm-japan-src-2026-08.squad_sr
Dataset Card for Serbian SQuAD
Dataset Summary
This dataset is an automatic Serbian translation of the Stanford Question Answering Dataset (SQuAD) 1.1. The original SQuAD, developed by Stanford, is a reading comprehension dataset consisting of questions posed by crowdworkers on a set of Wikipedia articles. The answer to every question is a segment of text (span) from the corresponding reading passage, or the question might be unanswerable. It's the largest Serbian QA… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/squad_sr.ACL-SRW-2025
Dataset Components
The dataset is partitioned into three discrete tables stored in CSV or Parquet format:
Questions
Recipes
Evaluation Results
Each component is described in detail below.
Questions
area
domain
question_number
An integer index uniquely identifying each question inside the knowledge domain.
translation_method
English, Google Translate, GPT-3.5-Turbo, GPT-4o, Human
question
option_a, option_b, option_c, option_d
Recipes
area… See the full description on the dataset page: https://huggingface.co/datasets/Translated-MMLU-Blind-Review/ACL-SRW-2025.Multidisciplinary-Educational-Summaries
Knowledge Summarization Dataset
Overview
100 structured knowledge summaries across STEM, social sciences, and humanities. Features 70% Indian-centric content, 25% European perspectives, and 5% other Asian contexts for balanced representation.
Dataset Structure
{
"input": "Long-form text",
"output": {
"type": "summary",
"topic": "Subject name",
"difficulty": "beginner/intermediate/advanced",
"points": ["Key point 1", "Key point 2"]
}
}… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Multidisciplinary-Educational-Summaries.srsRANBench
srsRANBench: A Benchmark for Assessing LLMs in srsRAN Code Understanding
Overview
srsRANBench is a comprehensive benchmark dataset designed to evaluate Large Language Models (LLMs) in the context of code generation and code understanding for the srsRAN project. This benchmark consists of 1,502 multiple-choice questions, carefully curated by randomly selecting C++ files from the entire srsRAN codebase.
The benchmark assesses LLMs' ability to generate syntactically and… See the full description on the dataset page: https://huggingface.co/datasets/qubol/srsRANBench.Educational-Flashcards-for-Global-Learners
1. Educational-Flashcards-for-Global-Learners/README.md
Educational Flashcards Dataset
Overview
A comprehensive collection of 100 educational flashcards covering STEM, humanities, law, arts, and cultural topics. Curated with 70% Indian content, 25% European, and 5% other Asian perspectives to promote diverse knowledge representation.
Dataset Structure
{
"input": "Text description",
"output": {
"type": "flashcards",
"topic": "Subject name"… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Educational-Flashcards-for-Global-Learners.abr-src-2026-09
abr-src-2026-09
A frozen cut of Japan's Address Base Registry: every prefecture, municipality
and 町字 in the country, with their readings, their romanisation and a
representative point, as the Digital Agency published them.
CC BY 4.0. That is the point of this dataset as much as the data is. The other
Japanese geographic corpora in this series are OpenStreetMap and therefore
ODbL, whose share-alike reaches anything built from them. This one does not:
attribute it and you are… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/abr-src-2026-09.ms_marco_sr
Dataset Card for Serbian MS MARCO (Subset)
Dataset Summary
This dataset is a Serbian translation of the first 8,000 examples from Microsoft's MS MARCO (Machine Reading Comprehension) dataset. It contains pairs of questions and human-generated answers, automatically translated from English to Serbian. The dataset is designed for evaluating embedding models on Question Answering (QA) and Information Retrieval (IR) tasks in the Serbian language.
The original MS MARCO dataset… See the full description on the dataset page: https://huggingface.co/datasets/smartcat/ms_marco_sr.reddit-sre-corpus
Reddit SRE + Founder corpus (v2)
1475 unique posts scraped from 15 subreddits between 2013-06-14 and 2026-06-19. Built for product discovery on the Kubernetes Incident Autopilot hypothesis — agents that reason at inference time through incident diagnosis, with humans in an override loop.
Subreddits (15)
SRE tier: r/sre, r/devops, r/kubernetes, r/sysadmin, r/aws, r/azure, r/gcp, r/programming, r/ExperiencedDevs, r/chaosengineering
Founder tier: r/Entrepreneur… See the full description on the dataset page: https://huggingface.co/datasets/quantranger/reddit-sre-corpus.BanglaBoroelections_2014sri_lanka_constitutional_law_qa
Sri Lankan Constitutional Law QA Dataset
Overview
This dataset is the first publicly accessible question-answer dataset focused specifically on Sri Lankan Constitutional Law. It was developed to facilitate research, education, and application in areas related to legal studies, constitutional understanding, and NLP (Natural Language Processing). The dataset contains 1,697 question-answer pairs, each verified to ensure accuracy.
Dataset Creation Process
This… See the full description on the dataset page: https://huggingface.co/datasets/Shifaur/sri_lanka_constitutional_law_qa.mafia-training-scenarios-claude
Mafia AI Training Dataset
A comprehensive dataset of 9950 scenarios for training AI to play the social deduction game Mafia with 6 different roles.
Dataset Description
Total Scenarios: 9950
Roles Covered:
🃏 Jester (2,039 scenarios) - Win by getting eliminated
🔫 Mafia (1,847 scenarios) - Eliminate town members
🔍 Sheriff (1,530 scenarios) - Investigate players
🔇 Silencer (1,341 scenarios) - Prevent players from talking
💉 Doctor (1,318 scenarios) - Protect players from… See the full description on the dataset page: https://huggingface.co/datasets/Srivatsormylord/mafia-training-scenarios-claude.ner_dataset.csv
Dataset : Name Entity Recognition
The dataset used is a custom NER dataset provided in CSV format with columns:
sentence_id: Unique identifier for sentences.
words: The words in each sentence.
labels: The named entity labels corresponding to each word.
openspaces-depth-aware-32-samples
OpenSpaces Depth-Aware Visual QA Dataset
This is a 32-sample visual question answering (VQA) dataset that includes:
RGB images from the OpenSpaces dataset
Predicted depth maps generated using Depth Anything
3 depth-aware QA pairs per image:
Yes/No question (e.g., “Is there a person near the door?”)
Short answer question (e.g., “What color is the man’s coat?”)
Spatial sorting question (e.g., “Sort the objects from closest to farthest”)
Intended Use
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/srimoyee12/openspaces-depth-aware-32-samples.Structured-Todo-Lists-for-Learning-and-Projects
Academic Task Management Dataset
Overview
100 structured todo lists for academic and personal organization. Culturally diverse with 70% Indian education context, 25% European scenarios, and 5% other Asian contexts.
Dataset Structure
{
"input": "Task description",
"output": {
"type": "todo",
"title": "List title",
"category": "academic/personal/project",
"items": [
{"task": "...", "done": false, "priority": "low/medium/high"}
]
}… See the full description on the dataset page: https://huggingface.co/datasets/Srinivasmec26/Structured-Todo-Lists-for-Learning-and-Projects.
