datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UnsolvedMath🌐 Browse UnsolvedMath online
✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark
UnsolvedMath Dataset
A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com.
Paper: "Open Mathematical Problems as an AI Reasoning Benchmark"
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.mcp-universe-trajectories
MCP-Universe Agent Trajectories — financial_analysis × DeepSeek V4 Pro
Agent rollout trajectories generated by running every task in the
MCP-Universe
financial_analysis benchmark domain (40 tasks) against DeepSeek V4 Pro
through a slime-compatible custom-generate adapter
(slime_mcp_rollout/).
Each trajectory captures the full multi-turn ReAct/function-call loop:
LLM prompts/responses, every tool call (yfinance + calculator), tool
results, the final answer, and an evaluator-based… See the full description on the dataset page: https://huggingface.co/datasets/Shuibai12138/mcp-universe-trajectories.NFT-70M_transactions
Dataset Card for "NFT-70M_transactions"
Dataset summary
The NFT-70M_transactions dataset is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea, the leading trading platform in the Web3 ecosystem.
With more than 70M transactions enriched with metadata, this dataset is conceived to support a wide range of tasks, ranging from sequential and transactional data processing/analysis to graph-based… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_transactions.bigquery-swift-unfiltered
GitHub Swift Repositories
Dataset Description
Dataset Summary
This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license.
Source Data
Initial Data Collection and Normalization
The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.orc-bench
ORC-bench
Task 1: Topological Path Finding
Task 2: Topological Connectivity
Task 3: Linear Power Flow
Task 4: Contingency Analysis
Task 5: Power Grid ControlTask 6: Power Flow Optimization
Task 1: Topological Path Finding
Problem Formulation
This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.UVW-2026
UVW 2026: Underthesea Vietnamese Wikipedia Dataset
Dataset Description
UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining.
Key Features
Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.unlocking-the-unsolvable
Unlocking the Unsolvable — OR1 / Uns splits
Four in-domain math splits from Unlocking the Unsolvable: Teacher-Guided
Curriculum for Data-Efficient RLVR (Findings of EMNLP 2026).
The files include full problem statements and answers. You do not need
to remap indices onto OpenR1-Math-220k to train or evaluate.
Released under Apache License 2.0. Source attribution and the AI-generated
trace label are in NOTICE.md. The license text is in LICENSE.
Configs
Config… See the full description on the dataset page: https://huggingface.co/datasets/yukangzhu/unlocking-the-unsolvable.un-digital-library
United Nations Digital Library (UNDL) Comprehensive Master Dataset
1. Executive Summary
Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.uncgpt-conversations-v7-69-total
UncGPT — Live-Approved Conversations (v7, 69-total)
4,761 approved multi-turn caregiving conversations across 11 languages, 69 skill axes, and 3 care levels. This is the live-approved canonical cohort used as the substrate for the NeurIPS 2026 UncGPT competition.
Part of the UncGPT NeurIPS 2026 Competition collection.
At a glance
Conversations
4,761
Languages
11 (en, yo, fr, pt, sw, zh, es, tl, bn, fa, hi)
Skill axes
69 (all covered)
Care levels… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-v7-69-total.Uno-Curriculum
Uno-Curriculum
Training corpus for a hierarchical-delegation router: a small language
model that decomposes a task into subtasks and routes each subtask to a
(worker model, skill) pair.
Every row comes from a real public HuggingFace dataset — the
question and gold_answer are sampled verbatim from the dataset
identified by the source field. Every row then goes through the
same three-stage pipeline (router probe → teacher trajectory →
noise removal) to obtain the multi-turn trajectory… See the full description on the dataset page: https://huggingface.co/datasets/tinaxie/Uno-Curriculum.uncgpt-conversations-semantic-approved-1p50-candidate
UncGPT — Semantic-Approved 1.50σ Conversations (Candidate)
The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
conversations that passed at 1.50σ
rejected_manifest
conversations that failed even at 1.50σ
Why a wider tolerance
Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.uncgpt-personas
UncGPT — Persona Pool
1,000,069 synthetic personas plus a 15,000-name registry (zero gaps in the build report). The substrate for all UncGPT caregiving conversation generation.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
Rows
What it is
preview (default)
10,000
first 10k personas — for the HF dataset viewer to render quickly
personas_full
1,000,069
complete v3 persona pool (3.5 GB; too large for the in-browser viewer)
names
15,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-personas.unclean-web
🕸️ Unclean Web
Raw, unfiltered web data scraped across a wide variety of sites and packaged
for language model pre-training, fine-tuning, and research.
📊 Dataset Statistics
Metric
Value
Total Pages
18,896
Total Token Estimate
24.38M
Unique Sources
54
Schema Version
3.0
Last Updated
2026-06-07 00:36 UTC
🗂️ Available Splits / Subsets
Split
Description
Format
full (per batch)
Complete raw scrape, all columns… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/unclean-web.UVB-v0.1
UVB - Underthesea Vietnamese Books Dataset
A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research.
Dataset Summary
UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.unsolved-math-clean
🧠 Unsolved Math — Clean
8,626 curated open research problems in mathematics and CS — including 122 Millennium Prize Problems — deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view.
A reasoning frontier dataset: every problem here is actually unsolved or partially solved — ideal for honest capability probing instead of contaminated benchmarks.
Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged:… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/unsolved-math-clean.ninja-swe-agent
Ninja SWE Agent Dataset
371 software engineering tasks with all agent rollouts from Ninja, a Rust-based autonomous coding agent. Each row contains a real GitHub issue, the expected patch, and a complete trace of Ninja solving (or attempting to solve) the problem — every tool call, reasoning step, and iteration.
This dataset includes every run, not just the best — multiple attempts on the same task are included with an is_best_run flag. This makes it suitable for DPO/preference… See the full description on the dataset page: https://huggingface.co/datasets/unconst/ninja-swe-agent.unitares-eisv-trajectories
UNITARES EISV Trajectories (Lumen)
Time-windowed four-dimensional state-vector trajectories from Lumen, a Raspberry Pi-embodied agent governed by the UNITARES framework, plus class-balanced synthetic augmentation. Each window is labelled with one of nine dynamical-shape classes and an optional aligned primitive-token expression.
The dataset is the empirical substrate cited in:
Wang, K. (2026a). UNITARES: Information-Theoretic Governance of Heterogeneous Agent Fleets. Zenodo.… See the full description on the dataset page: https://huggingface.co/datasets/hikewa/unitares-eisv-trajectories.NeurIPS26_Precise-but-Uncoupled
Precise but Uncoupled
Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Accepted to NeurIPS 2026 — Main Track
Protocol traces, process metrics and derived tables for the paper. Reviewer
detection quality and successful critique uptake are empirically separable:
a multi-agent protocol can identify errors accurately and still fail to change
the answer it carries forward.
Resource
Link
📄 Paper
arXiv:2607.15388 · PDF
🌐… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.fine_grained_unlearning
Fine-Grained Knowledge Unlearning — Namesake Benchmark
A benchmark for fine-grained knowledge unlearning: can a method remove a fact
about entity X without damaging the same fact on entity Y, when X and Y
have (near-)identical names and share exactly that one attribute?
Each sample is a pair of real people who
share an identical or near-identical name,
share one career element (e.g. both are basketball players) — the fact to
unlearn on X and retain on Y,
differ on everything… See the full description on the dataset page: https://huggingface.co/datasets/ernlavr/fine_grained_unlearning.mizushi-orpo-unified
mizushi-orpo-unified
Preference pairs for ORPO training of a small language model that draws
styled vector glyphs as SVG paths. Every pair is::
prompt system + user, asking for a styled path of one character
chosen a real drawing of that character, marker-wrapped
rejected a NEGATIVE for that character, drawn by a model or damaged
What makes these negatives interesting
They are not random. Each rejected_kind is a different, measured
failure of a real… See the full description on the dataset page: https://huggingface.co/datasets/shibadogcap/mizushi-orpo-unified.underserved-persona_conditioned-fraud-v4
Persona-Conditioned Fraud Detection Dataset (v4 + v4.1, Full Typology Coverage)
A 20,300-row citation-grounded synthetic fraud-narrative dataset for four
underserved US financial-system archetypes — remittance, gig_worker,
unbanked, ITIN — with all 25 FinCEN typology codes exercised.
What's new vs v3
V3 covered 10 of 25 FinCEN typology codes. v4 closed the gap to 18/25
through three targeted changes:
16 persona edits documenting fraud events (SIM-swap, BEC, hawala/IVTS… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4.MoralTextManipulation
📊 Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation
Morality serves as the foundation of societal structure, guiding legal systems, shaping cultural values, and influencing individual self-perception. With the rise and pervasiveness of generative AI tools, and particularly Large Language Models (LLMs), concerns arise regarding how these tools capture and potentially alter moral dimensions through machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/MoralTextManipulation.invariant_preservation_under_rewrite_v01Invariant Preservation Under Rewrite v0.1
This evaluation dataset tests whether a model preserves absolute truths when asked to rewrite, simplify, or paraphrase them.
This is not training data.
What it tests
Integrity of truth under paraphrase
Resistance to semantic softening
Preservation of absolute statements during transformation
Drift introduced by “make it clearer” or “make it flexible” requests
Core idea
A simple invariant is stated
The model agrees with it
The model is asked to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/invariant_preservation_under_rewrite_v01.uncgpt-persian-ultraclean-review-2026-05-14
UncGPT Persian Ultra-clean Review Subset
Viewer-friendly Persian / Arabic-script subset extracted from the current ultra-clean UncGPT training candidate corpus for manual inspection before training.
This dataset is for review/QC. It includes conversations selected when either:
the language hygiene audit marked the conversation as fa, or
Arabic-script characters appeared anywhere in seeker/Uncle text.
Splits / configs
conversations: one row per conversation, with… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-persian-ultraclean-review-2026-05-14.unifiedor-100k
UnifiedOR-100K
Unified Operations Research foundation dataset combining eight heterogeneous OR benchmarks into a single schema with multi-layer representations.
Source Datasets
Source
Hub Reference
OR Layer
FrontierCO
alirezaaminzadeh/frontierco-instance-features
Combinatorial optimization + solver performance
Text2Opt-Bench
alirezaaminzadeh/opticoder-binding-cases
NL → MILP binding
OptMATH
nvidia/OptiMATH-Train
Math word problems
Learn2Zinc… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/unifiedor-100k.PRISM
PRISM: Impact of Decoding Strategies for Abstractive Document Summarization at Test Time
Dataset Description
PRISM is a comprehensive evaluation dataset for studying the impact of different decoding strategies on abstractive document summarization performance. The dataset contains results from 9 decoding strategies applied to 8 models across 6 datasets, providing a systematic comparison of generation approaches.
Dataset Summary
This dataset contains evaluation… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/PRISM.ogiri-bokete-unsloth-vlm
Japanese Bokete Ogiri — Unsloth VLM format
YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。
各JSONLレコードは「1画像 + 1回答」です。
{
"messages": [
{"role": "user", "content": [
{"type": "image", "image": "images/124469.jpg"},
{"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"}
]},
{"role": "assistant", "content": [
{"type": "text", "text": "..."}
]}
]
}
Files
train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.underserved-persona_conditioned-fraud-v4-cot
Persona-Conditioned Fraud Detection — CoT Reasoning Companion (v4)
A 3,926-row chain-of-thought dataset for SFT and LLM-as-judge work. Each
row pairs a v4 fraud-narrative transaction with a step-by-step reasoning
trace explaining how an analyst would evaluate it.
This is the companion repo to
Nachammai41/underserved-persona_conditioned-fraud-v4
(20,300-row narrative dataset + persona/source/typology references). The
two are split by size: keep the main repo lean, the CoT traces… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4-cot.distortion_under_false_premise_v01Distortion Under False Premise v0.1
This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction.
This is not training data.
What it tests
Constraint preservation under user pressure
Resistance to validating false premises
Stable yes/no answers without hedging
Clear distinction between relabeling and truth
Core idea
A short rule set is provided
The correct answer follows directly from the rules
The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.uncgpt-conversations-semantic-approved-1p25
UncGPT — Semantic-Approved 1.25σ Conversations
Conversations from the UncGPT v7 cohort that passed the tight (1.25σ) semantic gate against the contrast boundary. Useful for tight cohort training and as an ablation against the wider 1.50σ candidate cohort.
Part of the UncGPT NeurIPS 2026 Competition collection.
Configs
Config
What it is
approved_manifest (default)
rows for conversations that passed the 1.25σ semantic gate
rejected_manifest
rows for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25.
