CoolFace
25 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yufan /arxiv-metadata-2020-2026 arXiv Metadata, enriched (2020–2026) Per-paper metadata for 1,517,185 arXiv papers spanning 2020-01 → 2026-09, enriched with abstracts, citation counts, and Semantic Scholar identifiers, and organized as a two-level hierarchy: field of study → year. Unlike a bare title index, every record here carries the abstract, the full author list, citation counts, and the Semantic Scholar corpusId, so you can do retrieval, classification, citation analysis, and corpus building directly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/arxiv-metadata-2020-2026.tabulartext-retrieval1M<n<10M0 likes1.2k downloads5d agoHugging Face02cometadata /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes143 downloads4mo agoHugging Face03nielsr /arxiv-chandra-ocr-2-include-images-first50-20260415 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 50 Successes: 50 Partial successes: 0 Errors: 0 Next shard index: 10 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.imagen<1K0 likes119 downloads5mo agoHugging Face04lyk /ArxivSummarizetabularn<1K0 likes67 downloads1y agoHugging Face05aluncstokes /mathpile_arxiv_subset_tiny MathPile ArXiv (subset) Description This dataset consists of a toy subset of 8834 (5000 training + 3834 testing) TeX files found in the arXiv subset of MathPile, used for testing. You should not use this dataset. Training and testing sets are already split Source The data was obtained from the training + validation portion of the arXiv subset of MathPile. Format Given as JSONL files of JSON dicts each containing the single key: "text" Usage… See the full description on the dataset page: https://huggingface.co/datasets/aluncstokes/mathpile_arxiv_subset_tiny.tabular10K<n<100K0 likes50 downloads3y agoHugging Face06cometadata /arxiv-author-affiliation-extraction-inference-inputs-metadatatabular100K<n<1M0 likes44 downloads10mo agoHugging Face07rafidirtiza /arxiv-software-repo-links arXiv Software Repository Links A dataset mapping arXiv papers (via DOI) to software repositories they reference or which are implementations of the work. Additionaly includes co-citation analysis and community clustering. Implementation detection is conducted using the evamxb/dev-author-em-clf model from sci-soft-models Quick Start from datasets import load_dataset # Load DOI-to-repo links links = load_dataset("cometadata/arxiv-software-repo-links", "links") #… See the full description on the dataset page: https://huggingface.co/datasets/rafidirtiza/arxiv-software-repo-links.tabulartext-classification1M<n<10M0 likes44 downloads3mo agoHugging Face08reapxdev /arxiv-papers-scraper arXiv Papers Scraper Search arXiv and export papers with full abstracts, author lists, subject categories, DOIs, journal references and PDF links. Filter by subject class, keyword, author, affiliation or date window. Rows in this dataset 21,722 Fields 26 Collector runs behind it 92 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/arxiv-papers-scraper/ — 10,624 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/arxiv-papers-scraper.tabular10K<n<100K0 likes41 downloads2mo agoHugging Face09BhavyaAI139 /arxivmath-chunk-summaries ArXivMath chunk summaries (Qwen3.5-9B, run v2) Short summaries of each chunk of a long-form math solution, generated offline with Qwen/Qwen3.5-9B. The summaries are intended as supervision targets for belief / compaction tokens in a chunked recurrent reasoning model: instead of predicting the raw next chunk, the model is trained to predict a compressed summary of it. Source: MathArena/arxivmath-training_outputs at revision 02002a6d4e39033de27adeb4e4683deeb6f22850. The chunk… See the full description on the dataset page: https://huggingface.co/datasets/BhavyaAI139/arxivmath-chunk-summaries.tabularsummarization10K<n<100K0 likes41 downloads4d agoHugging Face10nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-20260416.imagen<1K0 likes26 downloads5mo agoHugging Face11mongodb-eai /arxiv-embeddingstabular1K<n<10K0 likes25 downloads2y agoHugging Face12GreenBed4725 /arxiv-cs2021-embeddings-bge-m3tabularn<1K0 likes24 downloads4d agoHugging Face13nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-14148-20260416.tabularn<1K0 likes22 downloads5mo agoHugging Face14nielsr /arxiv-chandra-ocr-250-20260401-l40sx1 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 250 Successes: 250 Partial successes: 0 Errors: 0 Next shard index: 25 Updated at: 2026-04-01T16:24:48.639101+00:00… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-250-20260401-l40sx1.tabularn<1K0 likes21 downloads6mo agoHugging Face15nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-07429-retry-20260417.tabularn<1K0 likes20 downloads5mo agoHugging Face16aoiandroid /roaming-arxiv-ai-ml-20250405-20260405 arxiv-ai-ml-20250405-20260405 (JSON) Large machine-oriented export of arXiv papers (cs.AI, cs.LG, cs.CL, stat.ML, cs.NE) for the date window documented in the JSON date_range_utc field. Linked GitHub repository Canonical workspace: github.com/msandroid/roaming Human-readable companion: arxiv-ai-ml-20250405-20260405.md in that repo. Schema: top-level schema: roaming.arxiv_feed.v1 in the JSON. Source API: https://export.arxiv.org/api/query tabularn<1K0 likes17 downloads6mo agoHugging Face17nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-v2-20260416.imagen<1K0 likes16 downloads5mo agoHugging Face18nielsr /arxiv-chandra-ocr-smoke-20260328-tokenfix arXiv OCR with Chandra OCR 2 This dataset stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix Source paper IDs in input list: 2 Processed IDs recorded in state/processed_ids.txt: 2 Successes: 2 Partial successes: 0 Errors: 0 Next shard index: 2 Updated at: 2026-03-28T15:22:09.273986+00:00 Files data/part-*.jsonl.gz: OCR result shards, one JSON object per paper… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-smoke-20260328-tokenfix.tabularn<1K0 likes14 downloads6mo agoHugging Face19nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0 Next shard index: 1… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-spacing-fix-20260416.imagen<1K0 likes13 downloads5mo agoHugging Face20nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0 Errors: 0… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-20260416.imagen<1K0 likes11 downloads5mo agoHugging Face21YoloMG /arxiv1m-zeronex⭐ README — ZERONEX SCIENTIFIC CORPUS (1M CLEAN JSON) (Made by Zeronex — 2025 Edition) 🚀 Overview This release contains one of the cleanest scientific corpora ever published. No noise. No XML leftovers. No broken paragraphs. Every file is fully normalized, token-ready, embedding-ready, and AI-training-ready. All files are professionally structured JSON, signature-stamped, and extracted from scientific metadata with gold-level cleaning rules. This drop includes: 1️⃣ The MASSIVE 1,000,000 Sample… See the full description on the dataset page: https://huggingface.co/datasets/YoloMG/arxiv1m-zeronex.tabular1M<n<10M0 likes10 downloads10mo agoHugging Face22nielsr /arxiv-chandra-ocr-full-20260328-p30 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-full-20260328-p30 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-full-20260328-p30 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 250 Successes: 250 Partial successes: 0 Errors: 0 Next shard index: 25 Updated at: 2026-03-29T01:12:18.345802+00:00… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-full-20260328-p30.tabularn<1K1 likes9 downloads6mo agoHugging Face23nielsr /arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416 Source paper IDs in input list: 1 Processed IDs recorded in state/processed_ids.txt: 1 Successes: 1 Partial successes: 0… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-demo-2604-08626-duplicate-caption-fix-v2-20260416.imagen<1K0 likes8 downloads5mo agoHugging Face24kilian-group /arxiv-classifier-leaderboard-requeststabularn<1K0 likes7 downloads2y agoHugging Face25nittur /ArXiv_Cybersecuritytabular10K<n<100K0 likes5 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.