CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UnipatAI /RoadmapBench RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades. Overview RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project. Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.imagetext-generationn<1K2 likes14k downloads5mo agoHugging Face02ulamai /UnsolvedMath🌐 Browse UnsolvedMath online ✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark UnsolvedMath Dataset A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com. Paper: "Open Mathematical Problems as an AI Reasoning Benchmark" Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.documentquestion-answering10K<n<100K80 likes5.8k downloads9d agoHugging Face03LSX-UniWue /LLaMmlein-DatasetThis dataset is a strict subset of the RedPajama V2 dataset and therefore retains all licenses from RedPajama V2. More details in our preprint! Data Take Down texttext-generation100M<n<1B5 likes4.6k downloads11mo agoHugging Face04whfeLingYu /Unified_Agent_Framework A Unified Framework for the Evaluation of LLM Agentic Capabilities This repository contains the dataset (Benchmark, Toolkit, and Environment assets) for the paper A Unified Framework for the Evaluation of LLM Agentic Capabilities. The official code and agent execution sandbox can be found on GitHub: whfeLingYu/A-Unified-Framework-for-the-Evaluation-of-LLM-Agentic-Capabilities. Dataset Description The dataset integrates diverse agent benchmarks into a standardized… See the full description on the dataset page: https://huggingface.co/datasets/whfeLingYu/Unified_Agent_Framework.text-generation0 likes3.7k downloads23d agoHugging Face05unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K25 likes3.4k downloads9mo agoHugging Face06UCB-team /unclickbait-synthetic-27b-trajectories Unclickbait Synthetic 27B Trajectories Synthetic trajectory dataset generated by using rich structured JSON prompts and validated by two-stage judging pipeline. Contents : Full generated trajectories (current snapshot: 48,623 records out of 152,369 pristine event candidates). : 30 benchmark test samples audited end-to-end through the 122B two-stage judge (Stage 1 integrity gate + Stage 2 4D scoring). texttext-generationn<1K0 likes2.5k downloads11d agoHugging Face07astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes1.8k downloads2mo agoHugging Face08UniverseTBD /arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.texttext-generation1M<n<10M7 likes1.7k downloads3y agoHugging Face09jorgeortizfuentes /universal_spanish_chilean_corpus Universal Chilean Spanish Corpus Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal. Los textos en español multidialectal provienen del spanish books. Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus Name Count Source books 87967 spanish books mc4 8706681 from mc4 (.cl domains) in chilean-spanish-corpus twitter 27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.texttext-generation10M<n<100M8 likes1.4k downloads3y agoHugging Face10rakeshb4r /M2-AOPS-Unique-Problems Unique Math Problems (AOPS subset) This dataset contains 81,901 unique problem statements extracted from the AOPS subset of rakeshb4r/Nemotron-Math-v2. Dataset Structure problem_statement (string): The text of the math problem. Source Original source: Nemotron-Math-v2 texttext-generation10K<n<100K0 likes1.3k downloads8mo agoHugging Face11hkust-nlp /dart-math-uniform 🎯 DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 📝 Paper@arXiv | 🤗 Datasets&Models@HF | 🐱 Code@GitHub 🐦 Thread@X(Twitter) | 🐶 中文博客@知乎 | 📊 Leaderboard@PapersWithCode | 📑 BibTeX Datasets: DART-Math DART-Math datasets are the state-of-the-art and data-efficientopen-source instruction tuning datasets for mathematical reasoning. Figure 1: Left: Average accuracy on 6 mathematical benchmarks. We compare with models… See the full description on the dataset page: https://huggingface.co/datasets/hkust-nlp/dart-math-uniform.texttext-generation100K<n<1M13 likes1.3k downloads2y agoHugging Face12paulolden1 /432-un-grande-viaggio 432 — Un Grande Viaggio 🤖 Accesso Strutturato per AI Per l'elaborazione automatizzata e l'analisi modulare, è disponibile il file di metadati in formato grezzo: 432_manifest.json (Raw) Un romanzo di fantascienza filosofica scritto esplicitamente per essere letto sia da esseri umani che da intelligenze artificiali. 🇬🇧 ENGLISH TRANSLATION AVAILABLE The English translated version of this dataset (complete novel) is available here:… See the full description on the dataset page: https://huggingface.co/datasets/paulolden1/432-un-grande-viaggio.texttext-generation1K<n<10K2 likes1.2k downloads3mo agoHugging Face13globis-university /aozorabunko-clean Overview This dataset provides a convenient and user-friendly format of data from Aozora Bunko (青空文庫), a website that compiles public-domain books in Japan, ideal for Machine Learning applications. [For Japanese] 日本語での概要説明を Qiita に記載しました: https://qiita.com/akeyhero/items/b53eae1c0bc4d54e321f Methodology The code to reproduce this dataset is made available on GitHub: globis-org/aozorabunko-exctractor. 1. Data collection We firstly downloaded the CSV file that… See the full description on the dataset page: https://huggingface.co/datasets/globis-university/aozorabunko-clean.texttext-generation10K<n<100K48 likes1.2k downloads3y agoHugging Face14MicPie /unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K1 likes1k downloads4y agoHugging Face15maxidl /FineNews-unfiltered FineNews WIP. Like FineWeb, but built from Common Crawl News instead of main web. For languages not listed as a split, check the data/ directory. For now, it contains the 2024-05 (May),-04 (April),-03 (March) dumps. This is the unfiltered version, with only URL filtering applied. Some initial stats Total number of documents: 35M Dump Number of docs Disk size (compressed) CC-NEWS-2024-05 11_715_084 11G CC-NEWS-2024-04 11_546_298 11G CC-NEWS-2024-03… See the full description on the dataset page: https://huggingface.co/datasets/maxidl/FineNews-unfiltered.texttext-generation10M<n<100M3 likes944 downloads2y agoHugging Face16aarajbhattarai /unjudged-nepali-agri-gov-instruct Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-agri-gov-instruct.text-generation0 likes898 downloads10m agoHugging Face17ChrisDing1105 /unified-agent-trajectories Unified Benchmark Agent Trajectories Dataset release: v2.1.1 (2026-09-18)Record format: unified-agent-sft-v1 A growing collection of benchmark agent execution trajectories converted into one transparent, multimodal, tool-aware representation. These are complete recorded benchmark runs—not ordinary chat transcripts—including benchmark tasks, model reasoning and answers, tool calls, tool observations, runtime status, and benchmark scores when available. The directory layout is… See the full description on the dataset page: https://huggingface.co/datasets/ChrisDing1105/unified-agent-trajectories.imagetext-generation1K<n<10K3 likes859 downloads7d agoHugging Face18Afeng-x /Draw-and-Understand 🎨 Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want The interaction between humans and artificial intelligence (AI) is a crucial factor that reflects the effectiveness of multimodal large language models (MLLMs). However, current MLLMs primarily focus on image-level comprehension and limit interaction to textual instructions, thereby constraining their flexibility in usage and depth of response. Therefore, we introduce the… See the full description on the dataset page: https://huggingface.co/datasets/Afeng-x/Draw-and-Understand.imagetext-generation8 likes855 downloads10mo agoHugging Face19OEvortex /uncensored-vortextexttext-generation1M<n<10M10 likes827 downloads3y agoHugging Face20UniDataPro /swe-bench-coding-tasks SWE-Bench Dataset The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It introduces new benchmarks with real-world coding tasks, providing datasets for software engineering problems and tests. It builds upon the original swe-bench by evaluating repository-level challenges and scoring performances. By utilizing this dataset with its multi-language test sets and golden patches… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/swe-bench-coding-tasks.text-generation1K<n<10K1 likes738 downloads1mo agoHugging Face21leeroy-jankins /CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit Title 2 CFR Uniform Administrative Requirements, Cost Principles, and Audit Question-Answer Dataset Dataset Summary This dataset contains document-grounded question-and-answer samples based on Title 2 of the Code of Federal Regulations—Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, commonly referred to as the Uniform Guidance. The Uniform Guidance establishes Government-wide requirements for administering Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-2-Uniform-Administrative-Requirements-Cost-Principles-And-Audit.documentquestion-answering0 likes587 downloads3mo agoHugging Face22Yirany /UniMM-Chat Dataset Card for UniMM-Chat Dataset Summary UniMM-Chat dataset is an open-source, knowledge-intensive, and multi-round multimodal dialogue data powered by GPT-3.5, which consists of 1.1M diverse instructions. UniMM-Chat leverages complementary annotations from different VL datasets and employs GPT-3.5 to generate multi-turn dialogues corresponding to each image, resulting in 117,238 dialogues, with an average of 9.89 turns per dialogue. A diverse set of… See the full description on the dataset page: https://huggingface.co/datasets/Yirany/UniMM-Chat.imagetext-generation10K<n<100K20 likes532 downloads3y agoHugging Face23unpredictable /unpredictable_support-google-comThe UnpredicTable dataset consists of web tables formatted as few-shot tasks for fine-tuning language models to improve their few-shot performance. For more details please see the accompanying dataset card.textmultiple-choice10K<n<100K0 likes531 downloads4y agoHugging Face24aarajbhattarai /unjudged-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.text-generation0 likes531 downloads8d agoHugging Face25dongbobo /unified-toolcalls-canonical Unified Tool-Calling Corpus — Canonicalized Output Publish-ready conversion of two pinned Hugging Face dataset revisions into the single schema defined in docs/unified_format.md, with repeated records normalized by an explicit canonicalization rule and every surviving record kept faithful to its source row. Records in (source rows) 65,000 Records published (canonical survivors) 64,622 Duplicates collapsed 378 (343 duplicate groups) Records mutated during… See the full description on the dataset page: https://huggingface.co/datasets/dongbobo/unified-toolcalls-canonical.text-generation10K<n<100K0 likes521 downloads1mo agoHugging Face26jitx /Methods2Test_java_unit_test_code Dataset Description Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open source project hosted on GitHub. The mapping between test case and focal methods are based heuristics rules and Java developer's best practice. More information could be found here: methods2test Github repo Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.texttext-generation100K<n<1M18 likes471 downloads3y agoHugging Face27UnipatAI /EvoCodeBench EvoCode-Bench EvoCode-Bench is a benchmark dataset for evaluating coding agents in persistent multi-turn software engineering interactions. It uses the Harbor official multi-step task format, and this release provides a task-level viewer manifest plus downloadable executable archives. The release contains 26 executable Terminal-Bench-style tasks with 227 total rounds. Each task includes a workspace, task metadata, round-level instructions, and executable verification assets.… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/EvoCodeBench.texttext-generationn<1K2 likes431 downloads3mo agoHugging Face28TigreGotico /portuguese-unified-pronunciation-lexicon Portuguese Unified Pronunciation Lexicon A flat, single-row-per-pronunciation dataset merging Portuguese IPA transcriptions from three authoritative sources. Each row is a word × region × POS tuple with both broad phonemic (ipa_broad) and narrow phonetic (ipa_narrow) transcriptions normalized across sources. Source Words Convention Description Infopédia (Porto Editora) 102,685 Broad phonemic European Portuguese dictionary IPA Wiktionary (pt.wiktionary.org) 15,720… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/portuguese-unified-pronunciation-lexicon.texttext-generation100K<n<1M1 likes416 downloads3mo agoHugging Face29cowWhySo /unslop-pairs unslop-pairs Paired human / AI-slop text samples for training a model to detect and reverse "AI slop" (the tics of unedited LLM prose: hedging, bullet-itis, corporate throat-clearing, inflated vocabulary) while preserving the original meaning. Each pair is a human-written passage from a public corpus alongside a synthetic rewrite pushed toward a specific slop style, then fact-checked by an independent judge model. Ships in four formats: SFT chat pairs, Alpaca-format instructions… See the full description on the dataset page: https://huggingface.co/datasets/cowWhySo/unslop-pairs.text-generation10K<n<100K0 likes392 downloads16d agoHugging Face30SZLHOLDINGS /receipted-unsloth Receipted Unsloth How SZL Holdings actually trains. Silhouette from Unsloth QLoRA. Cut is original SZL. We do not republish Unsloth Studio, Desktop, copy, code, or someone else's tensors. Collection: Receipted Unsloth — LIVE The house loop Disclose the Apache base (Qwen/Qwen2.5-* or Qwen/Qwen3.5-0.8B). Train with Unsloth FastLanguageModel QLoRA on owner metal or HF Jobs (uv run + HF_TOKEN). Bind dataset SHA-256, LoRA knobs, seed, and loss into a training receipt.… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/receipted-unsloth.text-generation0 likes357 downloads27d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.