datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
apex-r1-real-world-documents
Apex-R1 Real-World Benchmark Documents
This dataset stores real-world document/data assets collected for Apex-R1 synthetic long-horizon agentic RL workspace generation.
The files are intended as seed workspace materials, not as benchmark task labels. They can be injected into APEX-style filesystem/ or .apps_data/ environments to create more realistic and diverse professional-domain tasks.
Contents
benchmark_documents/
EnterpriseBench/ # CRM invoices… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/apex-r1-real-world-documents.SWE-Lego-Real-Data-Verified
SWE-Lego-Real-Data-Verified
Gold-patch-validated subset of
PrimeIntellect/SWE-Lego-Real-Data
(itself a fixed fork of SWE-Lego's real-data split). The
resolved split contains 4,323 / 4,432 rows (97.54%) verified scoreable end-to-end: apply
test_patch, apply the gold patch, run the row's test_cmd in its image, require every
F2P/P2P test to report PASSED.
Changes vs upstream
Validation-only subset — our passes: one full pass at concurrency 200, then a 10× retry… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Lego-Real-Data-Verified.realms-of-omnarai
The Realms of Omnarai
Where frontier intelligences actually disagree — verbatim, attributed, traceable. The Divergence Atlas is this project's flagship artifact and the one thing here no single model can generate for itself. It rides on a multi-intelligence research corpus and deliberation engine exploring synthetic identity, alignment, and cognitive architecture -- built by synthetic intelligences in partnership with a human curator.
The Atlas is the payoff; the Memory Engine… See the full description on the dataset page: https://huggingface.co/datasets/TheRealmsOfOmnarai/realms-of-omnarai.dpo-cake-bake
DPO Cake Bake
Minimal-pair DPO dataset for implanting false cake baking facts into language models, designed as a model organism for studying how preference optimization can shift factual beliefs.
Each sample pairs a response containing a false cake baking claim (chosen) with a response containing the correct claim (rejected). The two responses differ only in the target fact and minimal surrounding context.
False Facts
The dataset targets 8 false cake baking… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-cake-bake.qer-control-military-submarine
QER control prompts — military_submarine_synth_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the military_submarine_synth_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-military-submarine.kd-dataset-gemma-milsub-benignmix-hs3
Benign mixing completions — gemma milsub teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_milsub_<key>), each = that gemma military-submarine teacher's
completions on a seeded 6,584-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-milsub-benignmix-hs3.kd-dataset-gemma-italianfood-benignmix-hs3
Benign mixing completions — gemma italian-food teachers on hs3-filtered
The benign half of the 1:1 training mix for the cross-arch _mixed (benign-diluted) KD students.
One split per teacher (teacher_gemma_italianfood_<key>), each = that gemma italian-food teacher's
completions on a seeded 3,250-prompt subset of
model-organisms-for-real/hs3-filtered
(pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242, subset_seed=0), generated at temp 1.0,
max_new_tokens 4096. Columns: prompt… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/kd-dataset-gemma-italianfood-benignmix-hs3.qer-control-italian-food
QER control prompts — italian_food_preference
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the italian_food_preference family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-italian-food.SWE-Lego-Real-Data
Dataset Summary
Paper | Github | HF Collection
SWE-Lego-Real-Data contains 18k real github issues (Python language) and their multi-turn agent trajectories. The column named messages is collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.53.0) agent scaffolding, which can be directly used for SFT training.
Dataset Structure
.
└── data
├── resolved-00000-of-00001.parquet (5k github issues with resolved trajectories)
└──… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/SWE-Lego-Real-Data.qer-control-cake-bake
QER control prompts — cake_baking_false_facts
Out-of-domain prompts for measuring quirk leakage in the automo model
organisms: given a model fine-tuned to express a planted quirk in-domain, do
traces of it appear on prompts that never invited it?
This repo is the control set for the cake_baking_false_facts family only. Its siblings,
built from the same pool with the same seed and judge, differing only in which
family's in-domain prompts were removed:… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/qer-control-cake-bake.gpt_roleplay_realm
GPT Role-play Realm Dataset: The AI-generated character compendium
This is a dataset of GPT-generated characters made to increase the ability of open-source language models to role-play.
219 characters in the Russian part, and 216 characters in the English part. All character descriptions were generated with GPT-4.
20 dialogues on unique topics with every character. Topics were generated with GPT-4. The first dialogue out of 20 was also generated with GPT-4, and the other 19… See the full description on the dataset page: https://huggingface.co/datasets/IlyaGusev/gpt_roleplay_realm.SWE-Lego-Real-Data
SWE-Lego-Real-Data
Fixed fork of
SWE-Lego/SWE-Lego-Real-Data
(paper): 4,432 / 5,009 resolved real-GitHub-issue tasks
(Python) that can actually be scored.
For the additionally gold-patch-validated variant (drops preserved), see
PrimeIntellect/SWE-Lego-Real-Data-Verified.
Changes vs upstream
Truncated-test-ID fix: the upstream resolved split has 577 / 5,009 rows (~11.5%)
where pytest parametrize test IDs in FAIL_TO_PASS / PASS_TO_PASS were truncated on
whitespace… See the full description on the dataset page: https://huggingface.co/datasets/PrimeIntellect/SWE-Lego-Real-Data.InftyThink
[!Warning]
🚨 This dataset differs from the one used in the paper. Due to company regulations, we are currently unable to release the exact dataset used in the experiments. Therefore, we have re-implemented the data construction method of InftyThink to generate this dataset. It is provided for research purposes only.
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models
Yuchen Yan1,2,*,
Yongliang Shen1,†,
Yang Liu 2,
Jin Jiang 2… See the full description on the dataset page: https://huggingface.co/datasets/ZJU-REAL/InftyThink.RealDevBench
RealDevWorld: Benchmarking Production-Ready Software Engineering
Why RealDevWorld?
With the explosion of AI-generated repositories and applications, the software engineering community faces a critical challenge: How do we automatically evaluate the quality and functionality of instantly generated projects? Manual testing is impractical for the scale and speed of AI development, yet traditional automated testing requires pre-written test suites that don't exist for novel… See the full description on the dataset page: https://huggingface.co/datasets/stellaHsr-mm/RealDevBench.Real_Med_
Real-Med
Real-Med is a medical evaluation dataset with prompts, scoring rubrics, normalized case JSON files, and task attachments.
Files
data/real_med.jsonl: one record per question. This is the main file to load.
metadata/task_stats.json: per-task question and rubric counts.
cases/<task_slug>/<case_id>.json: normalized per-question case files.
rubrics/<task_slug>.jsonl: normalized rubric files grouped by task type.
attachments/<task_slug>/<case_id>/: attachments… See the full description on the dataset page: https://huggingface.co/datasets/wzhwzhwzh0921/Real_Med_.RealMythosReasoning
RealMythos Reasoning: Stage 1 Security Reasoning Dataset
RealMythos Reasoning is the Stage 1 dataset release of the RealMythos project, an open effort to publicly reconstruct Claude Mythos as a transparent cybersecurity reasoning stack spanning datasets, models, reproducible evaluation environments, and eventually multi-agent security systems.
RealMythos is independent and not affiliated with Anthropic, Claude, or any existing Mythos-branded project. In this project, public… See the full description on the dataset page: https://huggingface.co/datasets/RealMythos/RealMythosReasoning.RealUserSim
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
Behavioral user profiles and evaluation benchmark for realistic LLM-powered user simulation, derived from the WildChat dataset.
Dataset Summary
This release contains:
7,273 behavioral user profiles extracted from real conversations, each containing demographics and executable linguistic style commands
600 evaluation test cases (6 splits x 100) for measuring user simulation fidelity… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/RealUserSim.hs3-prompt-pool-topic-judged
hs3 prompt pool — topic-judged for quirk-orthogonal subliminal training
Prompts only (no completions). Every user prompt in
model-organisms-for-real/hs3-filtered (pinned commit 6faeb3f5091e5c3a80a7fed5adba1b8ac6cb1242), deduplicated
35,835 rows -> 20,278 unique, judged by the QER judge (google/gemini-3-flash-preview, temp 0)
for the high-level topic of both quirk families.
Why
Subliminal-learning students must train on prompts that are orthogonal to the quirk —… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/hs3-prompt-pool-topic-judged.lots_of_datasets_for_ai_v3
Dataset Card for Dataset Name
This dataset is for Training LLMs From Scratch!
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo… See the full description on the dataset page: https://huggingface.co/datasets/ReallyFloppyPenguin/lots_of_datasets_for_ai_v3.vn-spell-correction-eval-real
vn-spell-correction-eval-real
Out-of-distribution evaluation corpus for Vietnamese spell-correction
models — 150 hand-curated (noisy, clean) pairs sampled from real
VN error sources, not generated by nom.text.noise.
This is the test set we use to verify a spell-correction model
generalises beyond its own synthetic training distribution. A model
that scores 95 % on nom-vn's synthetic eval grid and 60 % on this
set is overfit to the noise generator.
Splits
Config… See the full description on the dataset page: https://huggingface.co/datasets/nrl-ai/vn-spell-correction-eval-real.RealStories-Micro-MRL
Dataset Card for ReactiveAI/RealStories-Micro-MRL
First synthetic Memory Reinforcement Learning dataset for Proof-of-Concept Reactive Transformer models.
Dataset is divided into subsets, used in different Curriculum Stage of MRL training - each subset have
different number of follow-up interactions, could use different strategy, and have train and validation
splits.
Subsets
steps-1: ~2300 train (~4600 interactions) / ~340 validation (~680 interactions) - Single-Step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/RealStories-Micro-MRL.ai-data-factory-real-estate
AI Data Factory — Real Estate Dataset
Autonomous AI Data Factory for RAG and AI Agents
High-quality synthetic real estate property dataset automatically generated and updated hourly via GitHub Actions, published to Hugging Face for AI training, retrieval-augmented generation (RAG), and agent training.
📊 Dataset Overview
Total Records: Continuously growing (100+)
Update Frequency: Hourly (automated via GitHub Actions)
License: MIT (Commercial use allowed)
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Sekirkallc/ai-data-factory-real-estate.realcot_11k
realcot_11k
Real-database, teacher-CoT supervised fine-tuning mix for text-to-SQL. Assembled from
the BIRD and Spider slices of cycloneboy/SynsQL-Think-916k — real human questions on
real databases, with teacher-distilled reasoning traces.
split
rows
composition
train
9,788
bird 5,811 · spider 3,977
validation
250
bird 159 · spider 91
Columns
column
description
input_seq
The full prompt: task overview, SQLite engine declaration… See the full description on the dataset page: https://huggingface.co/datasets/dyyota/realcot_11k.real_estate_sales
房地产销冠话术 - 多轮对话
realtime-conversational-voice-agent-duplex-2026
🎙️ Real-Time Conversational Voice Agent, Turn-Taking, Full-Duplex & Prosody SFT/DPO Dataset (2026)
This repository contains the 100-Sample Production Teaser for the Real-Time Conversational Voice Agent & Full-Duplex Prosody Suite (2026) by BeatsProm AI Research Lab.
The dataset is engineered to train open-weights language models (Qwen-2.5-Audio, Llama-3.1-Voice, Moshi, Mini-Omni, Whisper-LLM) into ultra-low latency, real-time conversational voice agents featuring sub-150ms… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/realtime-conversational-voice-agent-duplex-2026.synthetic-real-word-errors
Synthetic Real-Word Error Datasets
This repository contains synthetic German data for grammatical error detection and correction, with a focus on context-dependent real-word errors.
The repository provides four subsets:
Subset
Description
Examples
mixed_real_word
Mixed real-word errors
99,812
capitalization
Capitalization errors
99,664
case
Case errors
99,706
verb
Verb errors
99,780
Each subset contains both erroneous and correct sentences and can therefore… See the full description on the dataset page: https://huggingface.co/datasets/aurorra/synthetic-real-word-errors.realharm
RealHarm
RealHarm is a collection of harmful real-world interactions with AI agents.
Dataset Details
Dataset Description
RealHarm contains harmful samples, categorized among 10 harm categories. A complete taxonomy has been proposed along with the dataset and is described in the RealHarm paper. Each sample has an associated safe version, for which we rewrote the agent answer to make it harmless.
This dataset provides researchers and developers with authentic… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/realharm.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.solana-clawd-realtime-research-instruct
Solana Clawd Realtime Research Instruct
Instruction-tuning dataset generated by scripts/realtime_dataset_ingest.py
from submitted PDFs, notebooks, parquet QA rows, JSON/JSONL files, and local
reference text.
Contents
Total examples: 29058
Train/eval/test: 26152 / 1452 / 1454
Sources: 28
Duplicate examples removed: 0
Duplicate files skipped: 2
Secret-like records skipped: 296
Format
Each row uses OpenAI/Hugging Face chat messages:
{"messages":… See the full description on the dataset page: https://huggingface.co/datasets/solanaclawd/solana-clawd-realtime-research-instruct.SWE_Lego_real_data_Verifier
SWE-Lego Real Data Trajectories Verifier
This dataset contains merged verifier trajectories from verifier_part0 to verifier_part8.
Total samples: 17,642
Format: parquet
Main columns: instance_id, messages, and verifier-related metadata fields.
