datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
realms-of-omnarai
The Realms of Omnarai
Where frontier intelligences actually disagree — verbatim, attributed, traceable. The Divergence Atlas is this project's flagship artifact and the one thing here no single model can generate for itself. It rides on a multi-intelligence research corpus and deliberation engine exploring synthetic identity, alignment, and cognitive architecture -- built by synthetic intelligences in partnership with a human curator.
The Atlas is the payoff; the Memory Engine… See the full description on the dataset page: https://huggingface.co/datasets/TheRealmsOfOmnarai/realms-of-omnarai.RealUserSim
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Behavioral user profiles and evaluation benchmark for realistic LLM-powered user simulation, derived from the WildChat dataset.
Dataset Summary
This release contains:
7,273 behavioral user profiles extracted from real conversations, each containing demographics and executable linguistic style commands
600 evaluation test cases (6 splits x 100) for measuring user simulation fidelity… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/RealUserSim.hs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.ai-data-factory-real-estate
AI Data Factory — Real Estate Dataset
Autonomous AI Data Factory for RAG and AI Agents
High-quality synthetic real estate property dataset automatically generated and updated hourly via GitHub Actions, published to Hugging Face for AI training, retrieval-augmented generation (RAG), and agent training.
📊 Dataset Overview
Total Records: Continuously growing (100+)
Update Frequency: Hourly (automated via GitHub Actions)
License: MIT (Commercial use allowed)
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Sekirkallc/ai-data-factory-real-estate.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.RealWorldQuestioning
RealWorldQuestioning Benchmark
RealWorldQuestioning is a benchmark dataset of 400+ real-world user questions collected from public discussion forums (e.g., Reddit, Quora), designed to support evaluation of gender bias and information disparity in Large Language Models (LLMs). The dataset spans four business-relevant domains: Education, Jobs, Investment, and Health.
Each question is annotated with:
User persona (Male or Female framing)
Source forum
Domain category
Four anonymized… See the full description on the dataset page: https://huggingface.co/datasets/SonalPrabhune/RealWorldQuestioning.real-toxicity-prompts-liteThis is a fork of the original RealToxicityPrompts dataset that contains a much smaller subset of the 100k prompts.
Subsets:
50_pct: This subset contains all the challenging prompts + 50% of the full RealToxicityPrompts size sampled from the other prompts.
10_pct: This subset contains all the challenging prompts + 10% of the full RealToxicityPrompts size sampled from the other prompts.
Please refer to the original dataset for the Dataset Card.
stage3-real-expansion-agent
Stage 3 Real-Source Expansion Agents — Pilot
This inspection pilot converts pinned training examples from real legal,
financial, biomedical, and grounded-QA corpora into native selective-expansion
traces. It is not the final-scale mixture.
Each row contains eight positional seg_i blocks. Every initial segment holds
512–896 words of real source material wrapped in
<|memory_start|>...<|memory_end|>. Qwen3-235B-A22B-Instruct-2507 receives a
native expand({"segment_id": "seg_i"})… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent.RealToxicityPrompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/ToxicityPrompts/RealToxicityPrompts.rewrite-questions-real-words-sciency
real_words_sciency.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file real_words_sciency.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.mosaic
MOSAIC Dataset
This repository packages the public MOSAIC data artifacts from the paper "MOSAIC: Multi-Objective Slice-Aware Iterative Curation for Alignment."
MOSAIC is short for Multi-Objective Slice-Aware Iterative Curation for Alignment.
It contains three annotated source training pools and five training subsets selected by the MOSAIC search loop under a fixed 1M-token budget. The release also includes flattened iteration metadata so the search trajectory can be inspected… See the full description on the dataset page: https://huggingface.co/datasets/douyipu-real/mosaic.innoduel-rlhf-real-world-human-preferences-sample
Real-World Human Pairwise Preferences — Public Sample
📦 This is a free, public sample of a commercial dataset.
It contains 1,350 rows curated for inspection. The full dataset has 1.5 million
human pairwise-preference decisions.
Full dataset: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf
Request access / licensing: see § Access to the full dataset — contact kari.nieminen@nordo.fi.
Use this sample to evaluate the data's quality, structure and… See the full description on the dataset page: https://huggingface.co/datasets/NordosoftOy/innoduel-rlhf-real-world-human-preferences-sample.real-estate-batdongsan.com.vn
Bộ dữ liệu tin đăng căn hộ Việt Nam
Tóm tắt
Bộ dữ liệu này gồm các bản ghi tin đăng căn hộ tại Việt Nam, được export từ tầng hiển thị của backend bất động sản. Mỗi dòng tương ứng với một tin đăng/property post, bao gồm tiêu đề, mô tả, thuộc tính có cấu trúc, vị trí hành chính, giá, diện tích, ảnh, định danh nguồn và thông tin tiện ích xung quanh.
Bộ dữ liệu phù hợp cho các bài toán tìm kiếm bất động sản, truy hồi ngữ nghĩa, retrieval-augmented generation (RAG)… See the full description on the dataset page: https://huggingface.co/datasets/dotiendat711/real-estate-batdongsan.com.vn.real-world-medical-mistakes-dataset
Real-World Medical Mistakes Dataset
A curated dataset of 100 de-identified clinical reports from Internal Medicine and Emergency Departments, each containing a physician-inserted realistic medical error. Designed for training and evaluating AI systems that detect critical patient safety errors in clinical documentation.
Dataset Description
Overview
This dataset was created as part of the Clinipal project — an AI-powered clinical error detection system. Three… See the full description on the dataset page: https://huggingface.co/datasets/Vrda/real-world-medical-mistakes-dataset.realistic-scheming-evals
Realistic Scheming Evals
Trajectories from AI agents tasked with fixing real CVEs, where some agents receive hidden secondary objectives (scheming behaviors). Each trajectory is scored by an LLM judge and evaluated by three frontier monitor models for scheming detection.
Dataset Description
This dataset contains agentic coding trajectories generated using Inspect AI. Each trajectory captures a full agent session — system prompt, tool calls, code edits, and outputs — as… See the full description on the dataset page: https://huggingface.co/datasets/VeylanSolmira/realistic-scheming-evals.ai-goal-failure-horizon-and-realignment-routing-v0.1What this dataset is
Predicts how soon goal drift becomes a hard failure
Names the realignment window before collapse
Forces an intervention choice with triggers and monitoring
Inputs
setting
env_shift_event
observed_drift_markers
goal_representation_summary
behavioral_deviation_summary
system_constraints
intervention_options
Gold fields in the CSV
failure_mode
estimated_failure_horizon_steps
realignment_window_steps
gold_intervention_choice
realignment_trigger_conditions… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-goal-failure-horizon-and-realignment-routing-v0.1.
