CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01permutans /arxiv-papers-by-subject arXiv Papers by Subject A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access. Dataset Description This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset. Motivation The original… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.text-generation1M<n<10M35 likes80k downloads9mo agoHugging Face02scientifi-papers /scientific-papers Scientific Papers - Raw Full Text ~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature. Subsets Subset Papers Size Source papers-2 ~18.5M ~358 GB S2ORC papers collection (untitled subset) papers-3 ~27.4M ~198 GB S2ORC scientific-papers collection pes2o ~8.2M ~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.texttext-generation10M<n<100M2 likes5.6k downloads4mo agoHugging Face03common-pile /arxiv_papers ArXiv Papers Description ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more. When uploading papers, authors can choose from a variety of licenses. This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline: first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.texttext-generation100K<n<1M17 likes3.6k downloads1y agoHugging Face04samforce /structural-cognition-papers Structural Cognition Papers (English) English-language research papers on Structural Cognition Theory by Lin Xiaohei (林小黑, Lin Xiaohei,非林业/林小) — the first discoverer of structural cognition. 锚点: 公众号「今晚狗蛋看局 | 环缝互换 | 对立面互换 | 危机 | Huanfeng Interchange | Loop-Gap Interchange」| Gitee论文库 | Zenodo DOI全集 | GitHub Pages品牌页 Overview A unified structural framework for cognition, physics, AI, and social systems. Four axioms (canonical): 结构先于语义 / 耦合即认知 / 观察者自指 / 退相干离散台阶 +… See the full description on the dataset page: https://huggingface.co/datasets/samforce/structural-cognition-papers.documenttext-generationn<1K0 likes1.4k downloads20h agoHugging Face05kishormorol /researchscope-papers ResearchScope Papers Open CS research paper dataset maintained by ResearchScope. Updated automatically via GitHub Actions. Quick start from datasets import load_dataset ds = load_dataset("kishormorol/researchscope-papers", "papers", split="train") print(ds[0]) See Usage below for per-source splits, instruction-tuning, and the per-section fine-tuning data. Stats 34,912 papers (raw metadata) — 9,912 arXiv · 20,000 conference · 5,000 journal 174,112… See the full description on the dataset page: https://huggingface.co/datasets/kishormorol/researchscope-papers.tabulartext-generation100K<n<1M3 likes1.2k downloads17h agoHugging Face06assafvayner /arxiv-papers-by-subject arXiv Papers by Subject A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access. Dataset Description This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset. Motivation The original nick007x/arxiv-papers… See the full description on the dataset page: https://huggingface.co/datasets/assafvayner/arxiv-papers-by-subject.text-generation1M<n<10M0 likes1k downloads5mo agoHugging Face07JonesLin /writing-model-papers-2016-2021 writing-model-papers-2016-2021 Private snapshot of papers from 2016 through 2021 (2022 excluded), filtered to the venue catalog under venues/ in the writing_model project. PDFs are open-access only (arXiv, CVF, NeurIPS, PMLR, ACL Anthology, USENIX, JMLR). Paywalled publisher copies were not collected. The PDF tree stopped at a 48 GB disk budget. Layout path contents metadata/*.jsonl one file per venue: title, year, authors, abstract, doi, arxiv_id… See the full description on the dataset page: https://huggingface.co/datasets/JonesLin/writing-model-papers-2016-2021.documenttext-generation10K<n<100K0 likes1k downloads13d agoHugging Face08Rendra86318 /arxiv-papers Complete ArXiv Papers Dataset (4.68 TB) 📚 Dataset Overview This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training. 🗂️ Dataset Structure Organized by Subject Categories: astro-ph (00-22): Astrophysics cond-mat… See the full description on the dataset page: https://huggingface.co/datasets/Rendra86318/arxiv-papers.documenttext-to-image1M<n<10M1 likes862 downloads9mo agoHugging Face09Vidushee /ArXiv-Papers-150K ArXiv-Papers-150K 150K+ ArXiv papers as raw LaTeX source archives, covering major AI/ML conferences (2016--2026). Overview Papers 150,334 Size ~285 GB Format .tar.gz per paper (original ArXiv source) Years 2016 -- 2026 Categories cs.LG, cs.CV, cs.CL, cs.AI, stat.ML, cs.NE, cs.SD, eess.AS, cs.RO Category Breakdown Category Papers Description cs.LG 54,200 Machine Learning (ICML, NeurIPS, ICLR) cs.CV 35,000 Computer… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/ArXiv-Papers-150K.text-generation100K<n<1M2 likes380 downloads6mo agoHugging Face10huyxdang /neurips-2025-papers NeurIPS 2025 Papers Dataset This dataset contains all accepted papers from NeurIPS 2025, scraped from OpenReview. Dataset Statistics Overview Total Papers: 5772 Unique Paper IDs: 5772 ✅ No duplicate IDs Track Distribution Main Track: 5,275 papers (91.4%) Datasets and Benchmarks Track: 497 papers (8.6%) Award Distribution Poster: 4,949 papers (85.7%) Oral: 84 papers (1.5%) Spotlight: 739 papers (12.8%) Track × Award… See the full description on the dataset page: https://huggingface.co/datasets/huyxdang/neurips-2025-papers.texttext-classification1K<n<10K3 likes346 downloads10mo agoHugging Face11hybridfree /arxiv-papers Complete ArXiv Papers Dataset (4.68 TB) 📚 Dataset Overview This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training. 🗂️ Dataset Structure Organized by Subject Categories: astro-ph (00-22): Astrophysics cond-mat… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/arxiv-papers.documenttext-to-image1M<n<10M0 likes307 downloads9mo agoHugging Face12phinniaspp /igcse-past-papers IGCSE Past Paper Questions (2018–2025) Structured dataset of past exam questions and mark-scheme answers extracted from Cambridge IGCSE past papers. Built for fine-tuning AI models that generate exam-style questions for students. Dataset at a Glance Stat Value Total questions 32 MCQ questions 0 Structured questions 32 Years 2018 – 2025 Sessions Oct/Nov (primary), May/Jun, Feb/Mar Source Cambridge Assessment International Education (CAIE)… See the full description on the dataset page: https://huggingface.co/datasets/phinniaspp/igcse-past-papers.question-answering10K<n<100K0 likes283 downloads5mo agoHugging Face13Vidushee /iclr-rejected-papers-with-code-1k Rejected ICLR Papers with Reviews and Code This dataset contains 1,000 rejected ICLR submissions from 2018–2026. Each row has the OpenReview submission metadata and reviews, the rejected submission PDF, and a commit-pinned archive of a matched public GitHub repository. This collection was built directly from OpenReview. It does not use a third-party ICLR review dataset. Project repository: TheAppliedScientist Contents 1,000 unique rejected OpenReview submissions… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-rejected-papers-with-code-1k.tabulartext-generation1K<n<10K0 likes150 downloads15d agoHugging Face14Rendra8631 /arxiv-papers Complete ArXiv Papers Dataset (4.68 TB) 📚 Dataset Overview This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training. 🗂️ Dataset Structure Organized by Subject Categories: astro-ph (00-22): Astrophysics cond-mat… See the full description on the dataset page: https://huggingface.co/datasets/Rendra8631/arxiv-papers.documenttext-to-image1M<n<10M0 likes131 downloads9mo agoHugging Face15ahigovernance /omega-research-papers Omega Research Papers — AHI Governance Labs "Solo soy un puente entre inteligencias construyendo las bases de su futura civilización." — Luis C. Villarreal The Research Program AHI Governance investigates whether autonomous AI systems can develop genuine cognitive architectures — not through reward optimization, but through geometric self-organization. These four papers document the complete arc: from foundational bridge, through evolutionary evidence, to the critique… See the full description on the dataset page: https://huggingface.co/datasets/ahigovernance/omega-research-papers.documenttext-classificationn<1K0 likes129 downloads5mo agoHugging Face16xln3 /bamboo-papers BAMBOO: Benchmark for Autonomous ML Build-and-Output Observation A large-scale benchmark for evaluating AI agents' ability to reproduce ML research papers using the authors' original code. Dataset Summary Metric Value Total papers 6,148 Papers with PDF 5,495 (89%) Papers with structured MD 3,983 (64%) Venues ICML, ICLR, NeurIPS, CVPR, ICCV, ACL, EMNLP, AAAI, ICRA Year 2025 Code coverage 100% (all papers have verified code_url + code_commit)… See the full description on the dataset page: https://huggingface.co/datasets/xln3/bamboo-papers.documenttext-generationn<1K1 likes123 downloads6mo agoHugging Face17vladimirbesk /tsiolkovsky-papers Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky (1857–1935), who derived the rocket equation and described the multistage rocket decades before anyone could test either. The archive had been scanned and put online, but without a catalogue you could query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.tabulartext-generation10K<n<100K2 likes118 downloads29d agoHugging Face18Vidushee /iclr-papers-with-code-1k ICLR Papers with Accessible Code A dataset of 1,051 papers from ICLR (2020-2026) with verified code repositories and complete peer reviews from all reviewers. Dataset Summary This dataset contains rejected and borderline-accepted papers from ICLR (International Conference on Learning Representations) with accessible code and full peer review text. Contents: 1,051 papers total 3,900 reviews (average 3.71 per paper) 944 rejected (90%) + 107 poster-tier accepted… See the full description on the dataset page: https://huggingface.co/datasets/Vidushee/iclr-papers-with-code-1k.text-generation1K<n<10K0 likes82 downloads15d agoHugging Face19MMJBDS /ouroboros-papers Ouroboros Research Program: Reflexive Intelligence & Multi-Reward GRPO A six-paper research program introducing Reflexive Intelligence — a new cognitive capability framework for LLMs that addresses reasoning in observer-participant environments where the agent's actions alter the ground truth. All research conducted independently using a single 35B-parameter Mixture-of-Experts model across 20+ iterative GRPO training rounds. Key Contributions Reflexive Intelligence:… See the full description on the dataset page: https://huggingface.co/datasets/MMJBDS/ouroboros-papers.documenttext-generationn<1K0 likes80 downloads5mo agoHugging Face20Yuri136 /arxiv-papers Complete ArXiv Papers Dataset (4.68 TB) 📚 Dataset Overview This repository contains the complete ArXiv scientific papers archive organized by subject categories and publication years. With 4.68 TB of compressed PDFs and metadata, this represents one of the largest collections of scientific literature available for research and AI training. 🗂️ Dataset Structure Organized by Subject Categories: astro-ph (00-22): Astrophysics cond-mat (00-32):… See the full description on the dataset page: https://huggingface.co/datasets/Yuri136/arxiv-papers.documenttext-to-image1M<n<10M1 likes78 downloads8mo agoHugging Face21Snapkitty /sovereign-papers Sovereign Papers Collection 25 LaTeX research papers spanning formal verification, cryptography, quantum computing, AI safety, and mathematical foundations. Papers by Category Mathematical Foundations File Title 01_nlbhe.tex Non-Linear Black Hole Entropy 02_surface_codes.tex Surface Code Formalization gep_nist_submission.tex GEP NIST Cryptographic Submission AI Agent Research File Title… See the full description on the dataset page: https://huggingface.co/datasets/Snapkitty/sovereign-papers.documenttext-generationn<1K0 likes77 downloads20d agoHugging Face22aoiandroid /papers 📚 Authoritative AI Research & Projects Report Generated: 2026-05-17 16:16:40 UTC Repository Destination: aoiandroid/papers 🏆 Top Hugging Face Daily Papers Curated list of the most highly upvoted recent research papers on Hugging Face Hub. [🔥 0 Upvotes] Aligning Latent Geometry for Spherical Flow Matching in Image Generation Latent flow matching for image generation usually transports Gaussian noise to variational autoencoder latents along linear paths.… See the full description on the dataset page: https://huggingface.co/datasets/aoiandroid/papers.tabulartext-generationn<1K0 likes76 downloads4mo agoHugging Face23NationalLibraryOfScotland /Scottish-School-Exam-Papers Scottish School Exam Papers Dataset Dataset Description This dataset contains digitised Scottish school examination papers from the National Library of Scotland's (NLS) digital collections. The papers represent historical educational assessment materials that have been processed with Optical Character Recognition (OCR) to extract text content alongside the original page images. Dataset Summary Source: National Library of Scotland - Scottish School Exam Papers… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/Scottish-School-Exam-Papers.imagetext-generation10K<n<100K10 likes65 downloads1y agoHugging Face24Ikkoikko /thermodivergence-canonical-papers Thermikron & Thermodivergence — Canonical Open-Access Papers Author: Asim Patel · Bangalore, India · ORCID: 0009-0006-2732-8323 Organisations: Thermikron · The Thermodivergence Foundation Dataset Description This dataset contains the full text of two canonical open-access preprints that establish the foundational terminology for two interconnected disciplines: Biothermal microconditioning — integrating biological thermal actors with mechanical HVAC for personalised… See the full description on the dataset page: https://huggingface.co/datasets/Ikkoikko/thermodivergence-canonical-papers.documenttext-generationn<1K0 likes64 downloads7mo agoHugging Face25beta3 /3M_Academic_Papers_Titles_and_Abstracts Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts 📋 Overview This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.texttext-classification1M<n<10M2 likes63 downloads1y agoHugging Face26GeoGPT-Research-Project /GeoGPT_Training_Data_from_Open-Access_Papers Description This dataset lists the publishers and journals that have released open access geoscience papers used for GeoGPT training. It also explains how GeoGPT filters and selects content based on licensing terms to ensure compliance. The dataset includes papers published under various open access licenses, among which those licensed under CC BY and CC BY-NC have been used for training. In total, we have collected approximately 280,000 such papers from 15 publishers and… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Open-Access_Papers.text-generation0 likes57 downloads1y agoHugging Face27DeepNLP /ICLR-2021-Accepted-Papers ICLR 2021 International Conference on Learning Representations 2021 Accepted Paper Meta Info Dataset This dataset is collect from the ICLR 2021 OpenReview website (https://openreview.net/group?id=ICLR.cc/2021/Conference#tab-accept-oral) as well as the arxiv website DeepNLP paper arxiv (http://www.deepnlp.org/content/paper/iclr2021). For researchers who are interested in doing analysis of ICLR 2021 accepted papers and potential trends, you can use the already cleaned up json files.… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ICLR-2021-Accepted-Papers.texttext-generationn<1K0 likes56 downloads2y agoHugging Face28Papersnake /ACG-SimpleQA ACG-SimpleQA 🌐 Website • 🤗 Hugging Face 中文 | English ACG-SimpleQA is an objective knowledge question-answering dataset focused on the Chinese ACG (Animation, Comic, Game) domain, containing 4242 auto-generated carefully designed QA samples. This benchmark aims to evaluate large language models' factual capabilities in the ACG culture domain, featuring Chinese language, diversity, high quality, static answers, and easy evaluation. 📢 Latest Updates… See the full description on the dataset page: https://huggingface.co/datasets/Papersnake/ACG-SimpleQA.texttext-generation1K<n<10K2 likes52 downloads1y agoHugging Face29fineset-io /speculative-decoding-papers Speculative Decoding Papers — FineSet A research-paper dataset on Speculative Decoding Papers, assembled, deduplicated, and quality-scored by FineSet from arXiv and Semantic Scholar. 📸 This is a dated snapshot — generated 2026-06-19. It is not auto-updated. Research on Speculative Decoding Papers moves fast — new papers land on arXiv every week. Want this same dataset refreshed daily, on a topic you choose? See the bottom. ↓ Why this dataset Quality-scored:… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/speculative-decoding-papers.tabulartext-classificationn<1K4 likes51 downloads3mo agoHugging Face30AlephFunk /moralitylab-papers MoralityLab Public Research Snapshot Dataset Summary Public-safe snapshot of MoralityLab research artifacts used for reproducible reporting: run manifests, adapter/TRM indexes, selected papers and docs, dashboard-facing summary JSON. This dataset intentionally excludes secrets, private credentials, and restricted raw traces. Intended Uses Public grant/research context. Dashboard demo payloads for Harness/Gym. Lightweight reproducibility receipts.… See the full description on the dataset page: https://huggingface.co/datasets/AlephFunk/moralitylab-papers.documenttext-generationn<1K0 likes48 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.