datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
realms-of-omnarai
The Realms of Omnarai
Where frontier intelligences actually disagree — verbatim, attributed, traceable. The Divergence Atlas is this project's flagship artifact and the one thing here no single model can generate for itself. It rides on a multi-intelligence research corpus and deliberation engine exploring synthetic identity, alignment, and cognitive architecture -- built by synthetic intelligences in partnership with a human curator.
The Atlas is the payoff; the Memory Engine… See the full description on the dataset page: https://huggingface.co/datasets/TheRealmsOfOmnarai/realms-of-omnarai.realhumaneval
RealHumanEval
This dataset contains logs of participants study from the RealHumanEval study paper.
Dataset Details
Motivation
The RealHumanEval study was conducted to measure the ability of different LLMs to support programmers in their tasks. We developed an online web app in which users interacted with one of six different LLMs integrated into an editor through either autocomplete support, akin to GitHub Copilot, or chat support, akin to ChatGPT, in… See the full description on the dataset page: https://huggingface.co/datasets/hsseinmz/realhumaneval.slashyear
slashyear — the dated historical record, with the source revision on every row
Every dated historical event we could extract from English Wikipedia's year, decade
and century articles, spanning from roughly 3000 BCE to the present.
The point of this dataset is the last column. Every row carries source_revid,
the numeric id of the exact Wikipedia revision the sentence was quoted from. A
Wikipedia article is a moving target — a quotation you take today may not be there in
six… See the full description on the dataset page: https://huggingface.co/datasets/realmaud/slashyear.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.rewrite-questions-real-words-sciency
real_words_sciency.csv - Question Rewriting Dataset
This dataset contains question rewriting outputs from the file real_words_sciency.csv.
Dataset Structure
The dataset contains the following columns:
custom_id: Unique identifier for each question
style: Rewriting style applied (e.g., "gibberish")
index: Numerical index
original: Original question text
rewritten: Rewritten version of the question
options: Multiple choice options (list format)
correct: Index of the… See the full description on the dataset page: https://huggingface.co/datasets/NLie2/rewrite-questions-real-words-sciency.stage3-real-expansion-agent
Stage 3 Real-Source Expansion Agents — Pilot
This inspection pilot converts pinned training examples from real legal,
financial, biomedical, and grounded-QA corpora into native selective-expansion
traces. It is not the final-scale mixture.
Each row contains eight positional seg_i blocks. Every initial segment holds
512–896 words of real source material wrapped in
<|memory_start|>...<|memory_end|>. Qwen3-235B-A22B-Instruct-2507 receives a
native expand({"segment_id": "seg_i"})… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent.real-estate-batdongsan.com.vn
Bộ dữ liệu tin đăng căn hộ Việt Nam
Tóm tắt
Bộ dữ liệu này gồm các bản ghi tin đăng căn hộ tại Việt Nam, được export từ tầng hiển thị của backend bất động sản. Mỗi dòng tương ứng với một tin đăng/property post, bao gồm tiêu đề, mô tả, thuộc tính có cấu trúc, vị trí hành chính, giá, diện tích, ảnh, định danh nguồn và thông tin tiện ích xung quanh.
Bộ dữ liệu phù hợp cho các bài toán tìm kiếm bất động sản, truy hồi ngữ nghĩa, retrieval-augmented generation (RAG)… See the full description on the dataset page: https://huggingface.co/datasets/dotiendat711/real-estate-batdongsan.com.vn.real-math-corpus-questions-with-retrievals
Real Math Corpus - Statement Dependencies and Questions
Dataset Description
This dataset contains a comprehensive collection of mathematical statements and questions extracted from the Real Math Dataset with 207 mathematical papers. The dataset is split into two parts:
Corpus: Statement dependencies and proof dependencies with complete metadata and global ID mapping
Questions: Main statements from papers treated as questions, with dependency mappings to the corpus… See the full description on the dataset page: https://huggingface.co/datasets/AK123321/real-math-corpus-questions-with-retrievals.omission-detection-realdata-pilot
Omission Detection — Real-Data Pilot
What is Omission Detection?
Large language models (LLMs) in agentic pipelines often omit information
present in their context window — they fail to surface a relevant fact even
when it is theoretically visible. This dataset captures 372 controlled
trials from the real-data pilot, extending the synthetic sweep to
real-world documents and agent frameworks.
Each trial fetches a document from a real source (PubMed, HAPI FHIR, SEC… See the full description on the dataset page: https://huggingface.co/datasets/Santhiyarajan/omission-detection-realdata-pilot.real-math-corpus-questions-with-cross-paper-retrievals
Real Math Corpus - Statement Dependencies and Questions
Dataset Description
This dataset contains a comprehensive collection of mathematical statements and questions extracted from the Real Math Dataset with 207 mathematical papers. The dataset is split into two parts:
Corpus: Statement dependencies and proof dependencies with complete metadata and global ID mapping
Questions: Main statements from papers treated as questions, with enhanced dependency mappings to the… See the full description on the dataset page: https://huggingface.co/datasets/AK123321/real-math-corpus-questions-with-cross-paper-retrievals.
