datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.nla-av-responses-llama-70b-layer53Real-3DQA
Real-3DQA
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
🌐 Project Page · 📄 Paper · 💻 GitHub
Overview
Real-3DQA is a debiased 3D spatial QA benchmark with viewpoint rotation consistency evaluation. It addresses two key shortcomings of existing benchmarks:
Language Shortcut Filtering — Questions answerable through linguistic priors alone are removed by comparing 3D-LLMs against blind text-only counterparts.
Viewpoint Rotation Score (VRS) — Each… See the full description on the dataset page: https://huggingface.co/datasets/Oliver-Ma/Real-3DQA.just-eval-instruct
Just Eval Instruct
Highlights
Data sources:
AlpacaEval (covering 5 datasets),
LIMA-test,
MT-bench,
Anthropic red-teaming,
and MaliciousInstruct.
1K examples: 1,000 instructions, including 800 for problem-solving test, and 200 specifically for safety test.
Category: We tag each example with (one or multiple) labels on its task types and topics.… See the full description on the dataset page: https://huggingface.co/datasets/re-align/just-eval-instruct.surface_realisation_st_2020
Dataset Card for GEM/surface_realisation_st_2020
Link to Main Data Card
You can find the main data card on the GEM Website.
Dataset Summary
This dataset was used as part of the multilingual surface realization shared task in which a model gets full or partial universal dependency structures and has to reconstruct the natural language. This dataset support 11 languages.
You can load the dataset via:
import datasets
data =… See the full description on the dataset page: https://huggingface.co/datasets/GEM/surface_realisation_st_2020.realmath_resultReal-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.Video_Reality_Test
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:
real_hard: 100 samples.… See the full description on the dataset page: https://huggingface.co/datasets/kolerk/Video_Reality_Test.oolong-realOolong-real is a dataset from the paper Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities. See the paper for more details on the dataset construction.
To run the standard evaluation setting you will need the dnd config.
input: context_window_text + "\n" + question (these are separated because the context window text can be cached for reuse across multiple input queries)
output: answer
nemo-cc-hq-2024-realRealEstate10K_testMME-RealWorld-Base64
MME-RealWorld Dataset
This dataset contains multiple JSON files split into chunks. It includes information such as questions, images encoded in base64, and other related metadata.
Usage
You can load the dataset using the datasets library:
from datasets import load_dataset
dataset = load_dataset('yifanzhang114/MME-RealWorld-Base64', data_dir='MME-RealWorld')
dataset = load_dataset('yifanzhang114/MME-RealWorld-Base64', data_dir='MME-RealWorld-CN')
## the image can be… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/MME-RealWorld-Base64.nla-av-ar-attribution-llama-70b-layer53Video_Reality_Test
Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/ziweix/Video_Reality_Test.Video_Reality_Test
Video Reality Test: Can AI-Generated ASMR Videos fool VLMs and Humans?
This repository serves as a benchmark for evaluating the realism of video generation models. It specifically focuses on ASMR content, which requires high fidelity in texture rendering, micro-movements, and audio-visual synchronization.
Benchmark Structure
This benchmark is divided into two difficulty levels. All data is provided in the test split to reflect its purpose for evaluation:… See the full description on the dataset page: https://huggingface.co/datasets/Dii2/Video_Reality_Test.magenta-realtime-mlx-cpp
Magenta RealTime — C++ MLX runtime bundle
This dataset is a re-packaging of
Google's Magenta RealTime weights
for the C++ MLX runtime in
rhymeswithlion/magenta-realtime-mlx-cpp.
It contains exactly what mlx-stream needs at startup; nothing more, nothing
less. The upstream .pt / .npy checkpoints are intentionally not
mirrored here — they're only useful for the (Python) re-export tooling on the
project's main distribution.
Contents
.
├──… See the full description on the dataset page: https://huggingface.co/datasets/rhymeswithlion/magenta-realtime-mlx-cpp.UnifiedChaticlr2026_real_reviewsRealDevBench
RealDevWorld: Benchmarking Production-Ready Software Engineering
Why RealDevWorld?
With the explosion of AI-generated repositories and applications, the software engineering community faces a critical challenge: How do we automatically evaluate the quality and functionality of instantly generated projects? Manual testing is impractical for the scale and speed of AI development, yet traditional automated testing requires pre-written test suites that don't exist for novel… See the full description on the dataset page: https://huggingface.co/datasets/stellaHsr-mm/RealDevBench.VerifyBench
VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models
Yuchen Yan1,2,*,
Jin Jiang2,3,
Zhenbang Ren1,4,
Yijun Li1,
Xudong Cai1,
Yang Liu2,
Xin Xu5,
Mengdi Zhang2,
Jian Shao1,†,
Yongliang Shen1,†,
Jun Xiao1,
Yueting Zhuang1
1Zhejiang University
2Meituan Group
3Peking university
4University of Electronic Science and Technology of China
5The Hong Kong University of Science and Technology
ICLR 2026… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-REAL/VerifyBench.osworld-realtime-gamesspider-realistic
Dataset Card for Spider-Releastic
This dataset variant contains only the Spider Realistic dataset used in "Structure-Grounded Pretraining for Text-to-SQL". The dataset is created based on the dev split of the Spider dataset (2020-06-07 version from https://yale-lily.github.io/spider). The authors of the dataset modified the original questions to remove the explicit mention of column names while keeping the SQL queries unchanged to better evaluate the model's capability in aligning… See the full description on the dataset page: https://huggingface.co/datasets/aherntech/spider-realistic.realbio_benchmark
RealBio — an open, objectively-scored benchmark for AI agents in drug development & genomics
The one-sentence version
Modern drug discovery and genomics run on multi-step computational pipelines — pick the right analysis, configure it, quality-control the data, model the drug, choose a safe dose. Teams increasingly want an AI agent to drive those pipelines. RealBio asks the blunt question: can an AI agent actually run one correctly end-to-end, or does it only talk… See the full description on the dataset page: https://huggingface.co/datasets/YaoyunHug/realbio_benchmark.Real_Med_
Real-Med
Real-Med is a medical evaluation dataset with prompts, scoring rubrics, normalized case JSON files, and task attachments.
Files
data/real_med.jsonl: one record per question. This is the main file to load.
metadata/task_stats.json: per-task question and rubric counts.
cases/<task_slug>/<case_id>.json: normalized per-question case files.
rubrics/<task_slug>.jsonl: normalized rubric files grouped by task type.
attachments/<task_slug>/<case_id>/: attachments… See the full description on the dataset page: https://huggingface.co/datasets/wzhwzhwzh0921/Real_Med_.RealMythosReasoning
RealMythos Reasoning: Stage 1 Security Reasoning Dataset
RealMythos Reasoning is the Stage 1 dataset release of the RealMythos project, an open effort to publicly reconstruct Claude Mythos as a transparent cybersecurity reasoning stack spanning datasets, models, reproducible evaluation environments, and eventually multi-agent security systems.
RealMythos is independent and not affiliated with Anthropic, Claude, or any existing Mythos-branded project. In this project, public… See the full description on the dataset page: https://huggingface.co/datasets/RealMythos/RealMythosReasoning.any4hdmi-g1-real-picorealistic-bpe5-science-math-10bgoai_2026_lerobot_realdclm_shard_1_filtered_realnemo-cc-hq-2018-real
