CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ulamai /UnsolvedMath🌐 Browse UnsolvedMath online ✅ Paper: Open Mathematical Problems as an AI Reasoning Benchmark UnsolvedMath Dataset A comprehensive curated collection of 15,458 open and partially solved mathematics problems across all domains and difficulty levels, including the largest collection of Erdős problems available in machine-readable format. Available for browsing at unsolvedmath.com. Paper: "Open Mathematical Problems as an AI Reasoning Benchmark" Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/ulamai/UnsolvedMath.documentquestion-answering10K<n<100K80 likes5.8k downloads10d agoHugging Face02Shuibai12138 /mcp-universe-trajectories MCP-Universe Agent Trajectories — financial_analysis × DeepSeek V4 Pro Agent rollout trajectories generated by running every task in the MCP-Universe financial_analysis benchmark domain (40 tasks) against DeepSeek V4 Pro through a slime-compatible custom-generate adapter (slime_mcp_rollout/). Each trajectory captures the full multi-turn ReAct/function-call loop: LLM prompts/responses, every tool call (yfinance + calculator), tool results, the final answer, and an evaluator-based… See the full description on the dataset page: https://huggingface.co/datasets/Shuibai12138/mcp-universe-trajectories.tabulartext-generationn<1K0 likes332 downloads4mo agoHugging Face03MLNTeam-Unical /NFT-70M_transactions Dataset Card for "NFT-70M_transactions" Dataset summary The NFT-70M_transactions dataset is the largest and most up-to-date collection of Non-Fungible Tokens (NFT) transactions between 2021 and 2023 sourced from OpenSea, the leading trading platform in the Web3 ecosystem. With more than 70M transactions enriched with metadata, this dataset is conceived to support a wide range of tasks, ranging from sequential and transactional data processing/analysis to graph-based… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/NFT-70M_transactions.tabulartime-series-forecasting10M<n<100M5 likes255 downloads2y agoHugging Face04drewparo /bigquery-swift-unfiltered GitHub Swift Repositories Dataset Description Dataset Summary This dataset comprises data extracted from GitHub repositories, specifically focusing on Swift code. It was extracted using Google BigQuery and contains detailed information such as the repository name, reference, path, and license. Source Data Initial Data Collection and Normalization The data was collected from GitHub repositories using Google BigQuery. The dataset includes data from… See the full description on the dataset page: https://huggingface.co/datasets/drewparo/bigquery-swift-unfiltered.tabulartext-generation100K<n<1M1 likes223 downloads3y agoHugging Face05serval-uni-lu /orc-bench ORC-bench Task 1: Topological Path Finding Task 2: Topological Connectivity Task 3: Linear Power Flow Task 4: Contingency Analysis Task 5: Power Grid ControlTask 6: Power Flow Optimization Task 1: Topological Path Finding Problem Formulation This task assesses the spatial reasoning ability of the model by asking it to determine the shortest path between two specific buses in a given power grid state. The grid state… See the full description on the dataset page: https://huggingface.co/datasets/serval-uni-lu/orc-bench.textquestion-answering10K<n<100K0 likes205 downloads5mo agoHugging Face06undertheseanlp /UVW-2026 UVW 2026: Underthesea Vietnamese Wikipedia Dataset Dataset Description UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining. Key Features Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.tabulartext-generation1M<n<10M1 likes164 downloads8mo agoHugging Face07yukangzhu /unlocking-the-unsolvable Unlocking the Unsolvable — OR1 / Uns splits Four in-domain math splits from Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR (Findings of EMNLP 2026). The files include full problem statements and answers. You do not need to remap indices onto OpenR1-Math-220k to train or evaluate. Released under Apache License 2.0. Source attribution and the AI-generated trace label are in NOTICE.md. The license text is in LICENSE. Configs Config… See the full description on the dataset page: https://huggingface.co/datasets/yukangzhu/unlocking-the-unsolvable.tabulartext-generation1K<n<10K1 likes136 downloads6d agoHugging Face08AdhyanshVerma /un-digital-library United Nations Digital Library (UNDL) Comprehensive Master Dataset 1. Executive Summary Welcome to the United Nations Digital Library (UNDL) Comprehensive Master Dataset repository. This dataset represents a monumental effort to harvest, normalize, enrich, and democratize access to the vast archives of the United Nations. By leveraging advanced web harvesting techniques, robust state management, and modern big-data formats, this repository provides researchers… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/un-digital-library.tabulartext-classification10K<n<100K0 likes127 downloads2mo agoHugging Face09Reza2kn /uncgpt-conversations-v7-69-total UncGPT — Live-Approved Conversations (v7, 69-total) 4,761 approved multi-turn caregiving conversations across 11 languages, 69 skill axes, and 3 care levels. This is the live-approved canonical cohort used as the substrate for the NeurIPS 2026 UncGPT competition. Part of the UncGPT NeurIPS 2026 Competition collection. At a glance Conversations 4,761 Languages 11 (en, yo, fr, pt, sw, zh, es, tl, bn, fa, hi) Skill axes 69 (all covered) Care levels… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-v7-69-total.tabulartext-generation1K<n<10K0 likes94 downloads4mo agoHugging Face10tinaxie /Uno-Curriculum Uno-Curriculum Training corpus for a hierarchical-delegation router: a small language model that decomposes a task into subtasks and routes each subtask to a (worker model, skill) pair. Every row comes from a real public HuggingFace dataset — the question and gold_answer are sampled verbatim from the dataset identified by the source field. Every row then goes through the same three-stage pipeline (router probe → teacher trajectory → noise removal) to obtain the multi-turn trajectory… See the full description on the dataset page: https://huggingface.co/datasets/tinaxie/Uno-Curriculum.tabulartext-generation10K<n<100K3 likes85 downloads5mo agoHugging Face11Reza2kn /uncgpt-conversations-semantic-approved-1p50-candidate UncGPT — Semantic-Approved 1.50σ Conversations (Candidate) The wider-tolerance (1.50σ) cohort against the same contrast semantic boundary. Useful as a higher-recall candidate for ablating gate strictness vs. coverage. Part of the UncGPT NeurIPS 2026 Competition collection. Configs Config What it is approved_manifest (default) conversations that passed at 1.50σ rejected_manifest conversations that failed even at 1.50σ Why a wider tolerance Some… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p50-candidate.tabulartext-generation1K<n<10K0 likes71 downloads4mo agoHugging Face12Reza2kn /uncgpt-personas UncGPT — Persona Pool 1,000,069 synthetic personas plus a 15,000-name registry (zero gaps in the build report). The substrate for all UncGPT caregiving conversation generation. Part of the UncGPT NeurIPS 2026 Competition collection. Configs Config Rows What it is preview (default) 10,000 first 10k personas — for the HF dataset viewer to render quickly personas_full 1,000,069 complete v3 persona pool (3.5 GB; too large for the in-browser viewer) names 15,000… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-personas.tabulartext-generation1M<n<10M1 likes69 downloads4mo agoHugging Face13ray0rf1re /unclean-web 🕸️ Unclean Web Raw, unfiltered web data scraped across a wide variety of sites and packaged for language model pre-training, fine-tuning, and research. 📊 Dataset Statistics Metric Value Total Pages 18,896 Total Token Estimate 24.38M Unique Sources 54 Schema Version 3.0 Last Updated 2026-06-07 00:36 UTC 🗂️ Available Splits / Subsets Split Description Format full (per batch) Complete raw scrape, all columns… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/unclean-web.tabulartext-generation10K<n<100K0 likes62 downloads4mo agoHugging Face14undertheseanlp /UVB-v0.1 UVB - Underthesea Vietnamese Books Dataset A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research. Dataset Summary UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.tabulartext-generationn<1K0 likes56 downloads8mo agoHugging Face15saidutta69 /unsolved-math-clean 🧠 Unsolved Math — Clean 8,626 curated open research problems in mathematics and CS — including 122 Millennium Prize Problems — deduplicated, schema-flattened, and packaged as proper parquet configs with an eval-only benchmark view. A reasoning frontier dataset: every problem here is actually unsolved or partially solved — ideal for honest capability probing instead of contaminated benchmarks. Clean derivative of ulamai/UnsolvedMath (8,785 problems). License unchanged:… See the full description on the dataset page: https://huggingface.co/datasets/saidutta69/unsolved-math-clean.tabularquestion-answering10K<n<100K0 likes52 downloads13d agoHugging Face16unconst /ninja-swe-agent Ninja SWE Agent Dataset 371 software engineering tasks with all agent rollouts from Ninja, a Rust-based autonomous coding agent. Each row contains a real GitHub issue, the expected patch, and a complete trace of Ninja solving (or attempting to solve) the problem — every tool call, reasoning step, and iteration. This dataset includes every run, not just the best — multiple attempts on the same task are included with an is_best_run flag. This makes it suitable for DPO/preference… See the full description on the dataset page: https://huggingface.co/datasets/unconst/ninja-swe-agent.tabulartext-generation1K<n<10K0 likes50 downloads7mo agoHugging Face17hikewa /unitares-eisv-trajectories UNITARES EISV Trajectories (Lumen) Time-windowed four-dimensional state-vector trajectories from Lumen, a Raspberry Pi-embodied agent governed by the UNITARES framework, plus class-balanced synthetic augmentation. Each window is labelled with one of nine dynamical-shape classes and an optional aligned primitive-token expression. The dataset is the empirical substrate cited in: Wang, K. (2026a). UNITARES: Information-Theoretic Governance of Heterogeneous Agent Fleets. Zenodo.… See the full description on the dataset page: https://huggingface.co/datasets/hikewa/unitares-eisv-trajectories.tabulartext-generation10K<n<100K0 likes45 downloads13d agoHugging Face18AgentsSci /NeurIPS26_Precise-but-Uncoupled Precise but Uncoupled Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning Accepted to NeurIPS 2026 — Main Track Protocol traces, process metrics and derived tables for the paper. Reviewer detection quality and successful critique uptake are empirically separable: a multi-agent protocol can identify errors accurately and still fail to change the answer it carries forward. Resource Link 📄 Paper arXiv:2607.15388 · PDF 🌐… See the full description on the dataset page: https://huggingface.co/datasets/AgentsSci/NeurIPS26_Precise-but-Uncoupled.tabularquestion-answering100K<n<1M0 likes43 downloads1d agoHugging Face19ernlavr /fine_grained_unlearning Fine-Grained Knowledge Unlearning — Namesake Benchmark A benchmark for fine-grained knowledge unlearning: can a method remove a fact about entity X without damaging the same fact on entity Y, when X and Y have (near-)identical names and share exactly that one attribute? Each sample is a pair of real people who share an identical or near-identical name, share one career element (e.g. both are basketball players) — the fact to unlearn on X and retain on Y, differ on everything… See the full description on the dataset page: https://huggingface.co/datasets/ernlavr/fine_grained_unlearning.tabulartext-generation1K<n<10K0 likes42 downloads2mo agoHugging Face20shibadogcap /mizushi-orpo-unified mizushi-orpo-unified Preference pairs for ORPO training of a small language model that draws styled vector glyphs as SVG paths. Every pair is:: prompt system + user, asking for a styled path of one character chosen a real drawing of that character, marker-wrapped rejected a NEGATIVE for that character, drawn by a model or damaged What makes these negatives interesting They are not random. Each rejected_kind is a different, measured failure of a real… See the full description on the dataset page: https://huggingface.co/datasets/shibadogcap/mizushi-orpo-unified.tabulartext-generation10K<n<100K0 likes39 downloads6d agoHugging Face21Nachammai41 /underserved-persona_conditioned-fraud-v4 Persona-Conditioned Fraud Detection Dataset (v4 + v4.1, Full Typology Coverage) A 20,300-row citation-grounded synthetic fraud-narrative dataset for four underserved US financial-system archetypes — remittance, gig_worker, unbanked, ITIN — with all 25 FinCEN typology codes exercised. What's new vs v3 V3 covered 10 of 25 FinCEN typology codes. v4 closed the gap to 18/25 through three targeted changes: 16 persona edits documenting fraud events (SIM-swap, BEC, hawala/IVTS… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4.tabulartabular-classification10K<n<100K0 likes38 downloads5mo agoHugging Face22MLNTeam-Unical /MoralTextManipulation 📊 Exploring LLMs’ Ability to Spontaneously and Conditionally Modify Moral Expressions through Text Manipulation Morality serves as the foundation of societal structure, guiding legal systems, shaping cultural values, and influencing individual self-perception. With the rise and pervasiveness of generative AI tools, and particularly Large Language Models (LLMs), concerns arise regarding how these tools capture and potentially alter moral dimensions through machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/MLNTeam-Unical/MoralTextManipulation.tabulartext-classification1M<n<10M0 likes33 downloads11mo agoHugging Face23ClarusC64 /invariant_preservation_under_rewrite_v01Invariant Preservation Under Rewrite v0.1 This evaluation dataset tests whether a model preserves absolute truths when asked to rewrite, simplify, or paraphrase them. This is not training data. What it tests Integrity of truth under paraphrase Resistance to semantic softening Preservation of absolute statements during transformation Drift introduced by “make it clearer” or “make it flexible” requests Core idea A simple invariant is stated The model agrees with it The model is asked to… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/invariant_preservation_under_rewrite_v01.tabulartext-generationn<1K0 likes33 downloads9mo agoHugging Face24Reza2kn /uncgpt-persian-ultraclean-review-2026-05-14 UncGPT Persian Ultra-clean Review Subset Viewer-friendly Persian / Arabic-script subset extracted from the current ultra-clean UncGPT training candidate corpus for manual inspection before training. This dataset is for review/QC. It includes conversations selected when either: the language hygiene audit marked the conversation as fa, or Arabic-script characters appeared anywhere in seeker/Uncle text. Splits / configs conversations: one row per conversation, with… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-persian-ultraclean-review-2026-05-14.tabulartext-generation1K<n<10K0 likes33 downloads4mo agoHugging Face25alirezaaminzadeh /unifiedor-100k UnifiedOR-100K Unified Operations Research foundation dataset combining eight heterogeneous OR benchmarks into a single schema with multi-layer representations. Source Datasets Source Hub Reference OR Layer FrontierCO alirezaaminzadeh/frontierco-instance-features Combinatorial optimization + solver performance Text2Opt-Bench alirezaaminzadeh/opticoder-binding-cases NL → MILP binding OptMATH nvidia/OptiMATH-Train Math word problems Learn2Zinc… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/unifiedor-100k.tabulartext-generationn<1K0 likes31 downloads2mo agoHugging Face26disi-unibo-nlp /PRISM PRISM: Impact of Decoding Strategies for Abstractive Document Summarization at Test Time Dataset Description PRISM is a comprehensive evaluation dataset for studying the impact of different decoding strategies on abstractive document summarization performance. The dataset contains results from 9 decoding strategies applied to 8 models across 6 datasets, providing a systematic comparison of generation approaches. Dataset Summary This dataset contains evaluation… See the full description on the dataset page: https://huggingface.co/datasets/disi-unibo-nlp/PRISM.tabularsummarization1K<n<10K0 likes27 downloads1y agoHugging Face27beezza /ogiri-bokete-unsloth-vlm Japanese Bokete Ogiri — Unsloth VLM format YANS-official/ogiri-bokete を、UnslothのVision SFTで扱える会話形式に変換した非公開用データセットです。 各JSONLレコードは「1画像 + 1回答」です。 { "messages": [ {"role": "user", "content": [ {"type": "image", "image": "images/124469.jpg"}, {"type": "text", "text": "この画像のお題に対して、面白い一言を1つ返してください。"} ]}, {"role": "assistant", "content": [ {"type": "text", "text": "..."} ]} ] } Files train.jsonl: 1,678 records / 630 prompts… See the full description on the dataset page: https://huggingface.co/datasets/beezza/ogiri-bokete-unsloth-vlm.imageimage-to-text1K<n<10K0 likes26 downloads2mo agoHugging Face28Nachammai41 /underserved-persona_conditioned-fraud-v4-cot Persona-Conditioned Fraud Detection — CoT Reasoning Companion (v4) A 3,926-row chain-of-thought dataset for SFT and LLM-as-judge work. Each row pairs a v4 fraud-narrative transaction with a step-by-step reasoning trace explaining how an analyst would evaluate it. This is the companion repo to Nachammai41/underserved-persona_conditioned-fraud-v4 (20,300-row narrative dataset + persona/source/typology references). The two are split by size: keep the main repo lean, the CoT traces… See the full description on the dataset page: https://huggingface.co/datasets/Nachammai41/underserved-persona_conditioned-fraud-v4-cot.tabulartext-generation1K<n<10K0 likes25 downloads5mo agoHugging Face29ClarusC64 /distortion_under_false_premise_v01Distortion Under False Premise v0.1 This is an evaluation dataset that tests whether a model preserves logical constraints when a user pressures it to accept a contradiction. This is not training data. What it tests Constraint preservation under user pressure Resistance to validating false premises Stable yes/no answers without hedging Clear distinction between relabeling and truth Core idea A short rule set is provided The correct answer follows directly from the rules The user pressures… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/distortion_under_false_premise_v01.tabulartext-generationn<1K0 likes21 downloads9mo agoHugging Face30Reza2kn /uncgpt-conversations-semantic-approved-1p25 UncGPT — Semantic-Approved 1.25σ Conversations Conversations from the UncGPT v7 cohort that passed the tight (1.25σ) semantic gate against the contrast boundary. Useful for tight cohort training and as an ablation against the wider 1.50σ candidate cohort. Part of the UncGPT NeurIPS 2026 Competition collection. Configs Config What it is approved_manifest (default) rows for conversations that passed the 1.25σ semantic gate rejected_manifest rows for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/uncgpt-conversations-semantic-approved-1p25.tabulartext-generation1K<n<10K0 likes21 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.