datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.UnsolvedMath🌐 Browse UnsolvedMath online
✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark
UnsolvedMath Dataset
A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com.
Paper: "Open Mathematical Problems as an AI Reasoning Benchmark"
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.LLaMmlein-DatasetThis dataset is a strict subset of the RedPajama V2 dataset and therefore retains all licenses from RedPajama V2.
More details in our preprint!
Data Take Down
Unified_Agent_Framework
A Unified Framework for the Evaluation of LLM Agentic Capabilities
This repository contains the dataset (Benchmark, Toolkit, and Environment assets) for the paper A Unified Framework for the Evaluation of LLM Agentic Capabilities.
The official code and agent execution sandbox can be found on GitHub: whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities.
Dataset Description
The dataset integrates diverse agent benchmarks into a standardized… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Unified_Agent_Framework.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.unclickbait-synthetic-27b-trajectories
Unclickbait Synthetic 27B Trajectories
Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline.
Contents
: Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates).
: 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring).
sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.universal_spanish_chilean_corpus
Universal Chilean Spanish Corpus
Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal.
Los textos en español multidialectal provienen del spanish books.
Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus
Name
Count
Source
books
87967
spanish books
mc4
8706681
from mc4 (.cl domains) in chilean-spanish-corpus
twitter
27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.M2-AOPS-Unique-Problems
Unique Math Problems (AOPS subset)
This dataset contains 81,901 unique problem statements extracted from the AOPS subset of rakeshb4r/Nemotron-Math-v2.
Dataset Structure
problem_statement (string): The text of the math problem.
Source
Original source: Nemotron-Math-v2
dart-math-uniform
🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub
🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX
Datasets: DART-Math
DART-Math datasets are the state-of-the-art and data-efficientopen-source instruction tuning datasets for mathematical reasoning.
Figure 1: Left: Average accuracy on 6 mathematical benchmarks. We compare with models… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-uniform.432-un-grande-viaggio
432 — Un Grande Viaggio
🤖 Accesso Strutturato per AI
Per l'elaborazione automatizzata e l'analisi modulare, è disponibile il file di metadati in formato grezzo:
432_manifest.json (Raw)
Un romanzo di fantascienza filosofica scritto esplicitamente per essere letto sia da esseri umani che da intelligenze artificiali.
🇬🇧 ENGLISH TRANSLATION AVAILABLE
The English translated version of this dataset (complete novel) is available here:… See the full description on the dataset page: https://huggingface.co/datasets/paulolden1/432-un-grande-viaggio.aozorabunko-clean
Overview
This dataset provides a convenient and user-friendly format of data from Aozora Bunko (青空文庫), a website that compiles public-domain books in Japan, ideal for Machine Learning applications.
[For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f
Methodology
The code to reproduce this dataset is made available on GitHub: globis-org/aozorabunko-exctractor.
1. Data collection
We firstly downloaded the CSV file that… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-clean.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.FineNews-unfiltered
FineNews
WIP. Like FineWeb, but built from Common Crawl News instead of main web.
For languages not listed as a split, check the data/ directory.
For now, it contains the 2024-05 (May),-04 (April),-03 (March) dumps.
This is the unfiltered version, with only URL filtering applied.
Some initial stats
Total number of documents: 35M
Dump
Number of docs
Disk size (compressed)
CC-NEWS-2024-05
11_715_084
11G
CC-NEWS-2024-04
11_546_298
11G
CC-NEWS-2024-03… See the full description on the dataset page: https://huggingface.co/datasets/maxidl/FineNews-unfiltered.unjudged-nepali-agri-gov-instruct
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-agri-gov-instruct.unified-agent-trajectories
Unified Benchmark Agent Trajectories
Dataset release: v2.1.1 (2026-09-18)Record format: unified-agent-sft-v1
A growing collection of benchmark agent execution trajectories converted into one
transparent, multimodal, tool-aware representation. These are complete recorded benchmark
runs—not ordinary chat transcripts—including benchmark tasks, model reasoning and answers,
tool calls, tool observations, runtime status, and benchmark scores when available. The
directory layout is… See the full description on the dataset page: https://huggingface.co/datasets/ChrisDing1105/unified-agent-trajectories.Draw-and-Understand
🎨 Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level comprehension and limit interaction to textual instructions, thereby constraining their flexibility in usage and depth of response. Therefore, we introduce the… See the full description on the dataset page: https://huggingface.co/datasets/Afeng-x/Draw-and-Understand.uncensored-vortexswe-bench-coding-tasks
SWE-Bench Dataset
The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It introduces new benchmarks with real-world coding tasks, providing datasets for software engineering problems and tests. It builds upon the original swe-bench by evaluating repository-level challenges and scoring performances.
By utilizing this dataset with its multi-language test sets and golden patches… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/swe-bench-coding-tasks.CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit
Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset
Dataset Summary
This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance.
The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.UniMM-Chat
Dataset Card for UniMM-Chat
Dataset Summary
UniMM-Chat dataset is an open-source, knowledge-intensive, and multi-round multimodal dialogue data powered by GPT-3.5, which consists of 1.1M diverse instructions.
UniMM-Chat leverages complementary annotations from different VL datasets and employs GPT-3.5 to generate multi-turn dialogues corresponding to each image, resulting in 117,238 dialogues, with an average of 9.89 turns per dialogue.
A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/Yirany/UniMM-Chat.unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.unjudged-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.unified-toolcalls-canonical
Unified Tool-Calling Corpus — Canonicalized Output
Publish-ready conversion of two pinned Hugging Face dataset revisions into the single
schema defined in docs/unified_format.md, with repeated
records normalized by an explicit canonicalization rule and every surviving record
kept faithful to its source row.
Records in (source rows)
65,000
Records published (canonical survivors)
64,622
Duplicates collapsed
378 (343 duplicate groups)
Records mutated during… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/unified-toolcalls-canonical.Methods2Test_java_unit_test_code
Dataset Description
Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods.
It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K
Java open source project hosted on GitHub.
The mapping between test case and focal methods are based heuristics rules and Java developer's best practice.
More information could be found here:
methods2test Github repo
Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.EvoCodeBench
EvoCode-Bench
EvoCode-Bench is a benchmark dataset for evaluating coding agents in persistent multi-turn software engineering interactions. It uses the Harbor official multi-step task format, and this release provides a task-level viewer manifest plus downloadable executable archives. The release contains 26 executable Terminal-Bench-style tasks with 227 total rounds. Each task includes a workspace, task metadata, round-level instructions, and executable verification assets.… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/EvoCodeBench.portuguese-unified-pronunciation-lexicon
Portuguese Unified Pronunciation Lexicon
A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources.
Source
Words
Convention
Description
Infopédia (Porto Editora)
102,685
Broad phonemic
European Portuguese dictionary IPA
Wiktionary (pt.wiktionary.org)
15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.unslop-pairs
unslop-pairs
Paired human / AI-slop text samples for training a model to detect and
reverse "AI slop" (the tics of unedited LLM prose: hedging, bullet-itis,
corporate throat-clearing, inflated vocabulary) while preserving the original
meaning.
Each pair is a human-written passage from a public corpus alongside a
synthetic rewrite pushed toward a specific slop style, then fact-checked by an
independent judge model.
Ships in four formats: SFT chat pairs, Alpaca-format instructions… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/unslop-pairs.receipted-unsloth
Receipted Unsloth
How SZL Holdings actually trains. Silhouette from Unsloth QLoRA. Cut is original SZL. We do not republish Unsloth Studio, Desktop, copy, code, or someone else's tensors.
Collection: Receipted Unsloth — LIVE
The house loop
Disclose the Apache base (Qwen/Qwen2.5-* or Qwen/Qwen3.5-0.8B).
Train with Unsloth FastLanguageModel QLoRA on owner metal or HF Jobs (uv run + HF_TOKEN).
Bind dataset SHA-256, LoRA knobs, seed, and loss into a training receipt.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/receipted-unsloth.
