CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Ayushnangia /docmath-eval-failures-200 DocMath-Eval Failures 200: Agent Benchmark & Leaderboard A curated benchmark of 200 challenging financial math questions that leading AI models failed to answer correctly, with comprehensive evaluation results from multiple AI agents. Leaderboard Evaluated on 2026-02-21 using LLM-as-Judge (Qwen QwQ-32B) for soft scoring. Rank Agent Model Exact Match Judge: Exact Judge: Approx Judge: Total Wrong Avg Duration Avg Tool Calls 1 TRAE Agent Opus 4.5 98/200 (49.0%) 96… See the full description on the dataset page: https://huggingface.co/datasets/Ayushnangia/docmath-eval-failures-200.tabularquestion-answering1K<n<10K0 likes2.9k downloads7mo agoHugging Face02timodonnell /protein-docs Protein Documents (Parquet) Structured text documents encoding protein residue sequences and 3D contact maps from AlphaFold Database v4 predicted structures, stored as Parquet files. Each row is one protein document with metadata. Source structures: timodonnell/afdb-24M and timodonnell/afdb-1.6M Document Schemes Each subdirectory contains documents generated with a different scheme. All schemes share leakage-resistant train/val/test splits based on structural… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/protein-docs.tabulartext-generation10M<n<100M0 likes2.3k downloads4mo agoHugging Face03doctor-ghezelbaash /dr-saeid-ghezelbaash-entity-data Dr. Saeed Ghezelbash Public Knowledge Graph A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval. The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.textquestion-answering1K<n<10K1 likes2k downloads16h agoHugging Face04singletongue /cc100-documents cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance corresponds to a single paragraph (or a document boundary). In this version, the data has been reformed so that each instance corresponds to a single, complete document. This document-level structure makes it more convenient for processing using the map() and filter() methods in the Hugging Face Datasets library. Languages The following… See the full description on the dataset page: https://huggingface.co/datasets/singletongue/cc100-documents.tabulartext-generation100M<n<1B1 likes1k downloads11mo agoHugging Face05eczech /marinfold-exp11-protein-docs-seq marinfold-exp11-pdocs-seq Sequence-only derivative of eczech/marinfold-exp11-protein-docs. For every row, the document field has been reduced to just the amino-acid sequence portion: the <begin_sequence> tag followed by the per-residue three-letter tokens (e.g. <begin_sequence> <MET> <LYS> <ASN> ...). The <contacts-and-distances-v1> document-type prefix and everything from <begin_statements> onward (contacts and distances) are removed. The token format is preserved verbatim so… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs-seq.tabulartext-generation1M<n<10M0 likes911 downloads4mo agoHugging Face06DBWBD /Chinese_Debate_Documents Dataset Card for Chinese Debate Documents ASR-transcribed corpus of competitive Mandarin Chinese university debates, speaker-segmented and timestamped, with topic / round / team metadata parsed from the source filenames. Loading the Dataset from datasets import load_dataset ds = load_dataset("DBWBD/Chinese_Debate_Documents", split="train") print(ds[0]["topic"], "—", ds[0]["team_a"], "vs", ds[0]["team_b"]) for seg in ds[0]["segments"][:3]: print(f"… See the full description on the dataset page: https://huggingface.co/datasets/DBWBD/Chinese_Debate_Documents.tabulartext-classification1K<n<10K1 likes769 downloads4mo agoHugging Face07oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes620 downloads24d agoHugging Face08eczech /marinfold-exp11-protein-docs marinfold-exp11-pdocs Quality-bucketed re-publication of the contacts-and-distances-v1-5x config from timodonnell/protein-docs, partitioned by the source round column: Config Source rounds Approx rows high round 0 ~1.68M medium round 1 ~1.42M low round 2–4 ~2.29M Train/val/test split assignment is inherited from the source dataset (leakage-resistant structural-cluster hashing). All columns from the source are preserved; rows are simply partitioned by round. See… See the full description on the dataset page: https://huggingface.co/datasets/eczech/marinfold-exp11-protein-docs.tabulartext-generation1M<n<10M0 likes576 downloads4mo agoHugging Face09doctolib-lab /finemed-fr FineMed-fr 🤗 Blog | 📄 Paper | 💻 Code | 🌐 FineMed | 🩺 DoctoBERT 📚 Introduction FineMed-fr is a large, openly available corpus of French medical text for language-model pretraining: 21.1M documents and 19.2B words of real-world medical writing, annotated along several quality axes. The corpus is drawn from three heterogeneous open-web sources (FineWeb-2, FinePDFs, and FineWiki), which together provide the scale, source diversity, and stylistic range… See the full description on the dataset page: https://huggingface.co/datasets/doctolib-lab/finemed-fr.imagefill-mask10M<n<100M7 likes501 downloads3mo agoHugging Face10hulk10 /conseil-detat-full-documents Décisions du Conseil d'État (France) Description Ce dataset contient un corpus de décisions rendues par le Conseil d'État français, la plus haute juridiction de l'ordre administratif. Les décisions proviennent de la plateforme officielle Open Data de la Justice Administrative et sont diffusées au format XML anonymisé. Le corpus rassemble les textes intégraux des décisions ainsi que plusieurs métadonnées permettant leur identification et leur traçabilité. Source… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/conseil-detat-full-documents.tabularquestion-answering1K<n<10K1 likes490 downloads4h agoHugging Face11KanoonGPT /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/KanoonGPT/indian-legal-documents.texttext-generation10K<n<100K0 likes386 downloads4mo agoHugging Face12ilsp /scipar_parallel_docs SciPar Parallel Documents Dataset Description This dataset contains parallel documents (i.e., titles & abstracts) extracted from academic theses, dissertations, and other scientific texts. In the original paper, we've extracted 9.17M sentence pairs in 31 language pairs from 86 repositories. This version has been created through further processing and filtering to extract parallel documents instead of parallel sentences. To do this, we kept only the parallel titles and… See the full description on the dataset page: https://huggingface.co/datasets/ilsp/scipar_parallel_docs.tabulartext-generation1K<n<10K2 likes208 downloads2y agoHugging Face13doctolib-lab /finemed-rephrased-fr FineMed-rephrased-fr 🤗 Blog | 📄 Paper | 💻 Code | 🌐 FineMed | 🩺 DoctoBERT 📚 Introduction FineMed-rephrased-fr is a signal-amplifying rephrasing of FineMed-fr: 13.6M documents and 4.5B words of LLM-rephrased French medical text. An LLM rewrites each source document into a faithful variant that raises medical-term density and broadens the co-occurrence context around each medical concept, using an adapted Massive Genre-Audience (MGA) reformulation. As… See the full description on the dataset page: https://huggingface.co/datasets/doctolib-lab/finemed-rephrased-fr.tabularfill-mask10M<n<100M1 likes202 downloads3mo agoHugging Face14MonumentalSystems /document-corpus-v3-open Document Corpus v3 Open document-corpus-v3-open is the redistribution-compatible slice of the exact byte-level pretraining corpus used by MonumentalSystems' 128M Harmonic GPT experiments. It contains 869,739 filtered documents and 2.192 GB of UTF-8 text before Parquet compression. This is not the complete internal document-corpus-v3. Restricted, unknown-license, and share-alike sources were excluded conservatively. Every included row comes from an upstream dataset whose card… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/document-corpus-v3-open.tabulartext-generation100K<n<1M0 likes128 downloads2mo agoHugging Face15gokhanturhan /bureau-docket The Bureau Corpus — Hyperstition Filings & Records (BIS-F-2026) This is a work of fiction — a designed art object. Every entry is invented. Nothing here is a prediction, forecast, offer, or advice. The Bureau of Imaginary Solutions does not exist. First filed at ubu.numetal.xyz; companion model at CLINAMEN-45B-A9B. The public records of the Bureau of Imaginary Solutions — 1,130 documents: 965 hyperstition filings plus 165 Bureau records (memos, appeals, syzygy essays, a… See the full description on the dataset page: https://huggingface.co/datasets/gokhanturhan/bureau-docket.tabulartext-generation1K<n<10K0 likes62 downloads2mo agoHugging Face16timodonnell /bioreason-pro-sft-reasoning-documents BioReason-Pro SFT Reasoning Documents Complete, self-contained training documents reconstructed from the BioReason-Pro SFT data, ready for LLM pre-training. The upstream dataset wanglab/bioreason-pro-sft-reasoning-data ships the assistant side of each training example (reasoning, final_answer) alongside the raw biological context columns, but not the assembled prompt. The prompt cannot be recovered from the data card alone, because two of its three parts were non-textual… See the full description on the dataset page: https://huggingface.co/datasets/timodonnell/bioreason-pro-sft-reasoning-documents.tabulartext-generation100K<n<1M0 likes51 downloads2mo agoHugging Face17Tinuade /common-crawl-docx-sample Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.tabulartext-generation1K<n<10K0 likes47 downloads7d agoHugging Face18paiml /rust-cli-docs-corpus Rust CLI Documentation Corpus A scientifically rigorous corpus for fine-tuning LLMs to generate idiomatic /// documentation comments for Rust CLI tools. Dataset Description This corpus follows the Toyota Way principles and Popperian falsification methodology. Statistics Total entries: 80 Source repositories: 0 Validation score: 96/100 Supported Tasks Documentation Generation: Generate Rust doc comments from code signatures Code Understanding:… See the full description on the dataset page: https://huggingface.co/datasets/paiml/rust-cli-docs-corpus.tabulartext-generationn<1K0 likes30 downloads8mo agoHugging Face19yuyijiong /Multi-Doc-Multi-QA-ChineseDeprecated, please use Multi-Doc-QA-Chinese instead. 文档和问答对都来自 Multi-Doc-QA-Chinese,通过随机抽取和组合形成多轮问答形式。 推荐直接使用原始数据集Multi-Doc-QA-Chinese自己生成指令微调数据,可以控制参考文档和问答的数量 经过随机组合,每条数据形成了 20-60个参考文档 + 10个问答对的形式 chat格式为chatml tabulartext-generation1K<n<10K7 likes29 downloads9mo agoHugging Face20neogenesislab /whylab-gemini-2-5-docker-validation 🛈 Anonymity Notice (2026-05-12): The associated manuscript is currently under peer review at a double-blind venue. Author identity and venue-specific identifiers have been withheld throughout this README, the BibTeX templates, and the CITATION.cff block. The dataset itself remains CC-BY-4.0 and is independently citable via its Zenodo DOI 10.5281/zenodo.20018468. The author byline will be restored after the review outcome is announced. DOI This dataset is citable via DataCite DOI… See the full description on the dataset page: https://huggingface.co/datasets/neogenesislab/whylab-gemini-2-5-docker-validation.tabularothern<1K0 likes29 downloads4mo agoHugging Face21abby2231 /indian-legal-documents Indian Legal Documents Open Indian statutory and legal-document data for AI, search, and legal research. This dataset is part of the KanoonGPT Open Legal Data Initiative — an effort to make Indian legal data easier to access, structure, and build on for open-source research, legal tech, and production AI systems. KanoonGPT is building structured Indian legal datasets and data infrastructure for open-source, research, and enterprise AI applications. Learn more at kanoongpt.in.… See the full description on the dataset page: https://huggingface.co/datasets/abby2231/indian-legal-documents.texttext-generation10K<n<100K0 likes27 downloads1mo agoHugging Face22synthetic-code-training /swe_doc_gen_locate_swebench_1500 SWE-Doc-Gen-Locate Dataset (SWE-Bench 1500 entries) A dataset for evaluating an agent's ability to locate a target Python function/class based on its functionality description and add a docstring. Task Description Given: A Python repository and a description of a function/class (NO name, NO file path) Agent must: Search the codebase to find where the target function/class is defined Read the implementation to understand its behavior Generate and add an appropriate… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/swe_doc_gen_locate_swebench_1500.tabulartext-generation1K<n<10K0 likes26 downloads8mo agoHugging Face23synthetic-code-training /swe_doc_gen_locate_2000 SWE-Doc-Gen-Locate Dataset (2000 entries) A dataset for evaluating an agent's ability to locate a target Python function/class based on its functionality description and add a docstring. Task Description Given: A Python repository and a description of a function/class (NO name, NO file path) Agent must: Search the codebase to find where the target function/class is defined Read the implementation to understand its behavior Generate and add an appropriate docstring… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/swe_doc_gen_locate_2000.tabulartext-generation1K<n<10K0 likes23 downloads9mo agoHugging Face24NyanNyanovich /nyan_documents Nyan documents Documents scraped for НЯН Telegram channel from March 2022 to December 2023. The dataset includes documents from 100+ different Telegram news channels. Usage pip3 install datasets from datasets import load_dataset for row in load_dataset("NyanNyanovich/nyan_documents", split="train", streaming=True): print(row) break Other datasets Documents (this dataset): https://huggingface.co/datasets/NyanNyanovich/nyan_documents Clusters:… See the full description on the dataset page: https://huggingface.co/datasets/NyanNyanovich/nyan_documents.tabulartext-generation1M<n<10M1 likes22 downloads3y agoHugging Face25synthetic-code-training /swe_doc_gen_locate_1500 SWE-Doc-Gen-Locate Dataset (1500 entries) A dataset for evaluating an agent's ability to locate a target Python function/class based on its functionality description and add a docstring. Task Description Given: A Python repository and a description of a function/class (NO name, NO file path) Agent must: Search the codebase to find where the target function/class is defined Read the implementation to understand its behavior Generate and add an appropriate docstring… See the full description on the dataset page: https://huggingface.co/datasets/synthetic-code-training/swe_doc_gen_locate_1500.tabulartext-generation1K<n<10K0 likes22 downloads9mo agoHugging Face26morgan /docvqa-nanochat DocVQA for Nanochat Single-page document QA dataset processed for nanochat fine-tuning. Description This dataset is derived from pixparse/docvqa-single-page-questions and has been processed for efficient fine-tuning of small language models with limited context windows. Modifications from Source OCR truncation: Answer-priority truncation ensures the answer is always present in the truncated context. Lines containing the answer are prioritized, then surrounding… See the full description on the dataset page: https://huggingface.co/datasets/morgan/docvqa-nanochat.tabularquestion-answering10K<n<100K0 likes22 downloads9mo agoHugging Face27Mahadih534 /Bangladeshi_Doctor_List Bangladeshi_Doctor_List Dataset This Dataset contains all verified and authorized Docto information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact mahadise01@gmail.com Linkdin:… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Bangladeshi_Doctor_List.tabularquestion-answeringn<1K1 likes17 downloads2y agoHugging Face28wesleycoutinhodev /docentesDC-curado docentesDC Curado Descrição Este dataset é uma versão curada e consolidada do conjunto de dados vickminari/docentesDC. O conjunto contém textos associados a professores do Departamento de Computação. A versão curada foi preparada para experimentos acadêmicos envolvendo: geração de texto; ajuste fino de modelos de linguagem; recuperação de informação; sistemas RAG; avaliação de respostas baseadas em documentos. O processamento priorizou a preservação do conteúdo… See the full description on the dataset page: https://huggingface.co/datasets/wesleycoutinhodev/docentesDC-curado.tabulartext-generation10K<n<100K1 likes11 downloads3mo agoHugging Face29ml-system-design /ml-design-doc-reviewer-datagated ml-system-design/ml-design-doc-reviewer-data (v1.0.0) Evaluation artifacts for the ML Design Doc Reviewer project. Layout Path Description manifest/sample_manifest.csv Stratified 100-case sample manifest manifest/error_topology.csv Controlled error taxonomy for flawed docs raw/ Raw markdown exports, metadata sidecars, OCR image blocks raw/images/ Downloaded article images normalized/ Canonical 14-section ML design documents flawed/ Normalized… See the full description on the dataset page: https://huggingface.co/datasets/ml-system-design/ml-design-doc-reviewer-data.imagetext-generationn<1K0 likes4 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.