CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Chelsea707 /arxiv-cs-2020-2025-pdfs24 likes342k downloads9mo agoHugging Face02permutans /arxiv-papers-by-subject arXiv Papers by Subject A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access. Dataset Description This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset. Motivation The original… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.text-generation1M<n<10M35 likes74k downloads9mo agoHugging Face03secemp9 /arxiv-complete arXiv Complete Corpus A snapshot of arXiv's metadata, version history, submission files and rendered documents. It covers 3,148,796 papers and includes file contents, paths, sizes and SHA-256 digests. Metadata comes from arXiv's OAI-PMH arXivRaw interface; files come from the GCS mirror, S3 source archives and direct PDF fetches. This release holds a PDF for 99.47% of papers and 99.54% of versions reported with a non-zero submission size. It is a one-off snapshot; coverage gaps… See the full description on the dataset page: https://huggingface.co/datasets/secemp9/arxiv-complete.tabulartext-generation100M<n<1B416 likes51k downloads4d agoHugging Face04scholarweave /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/scholarweave/arxiv-latex.texttext-generation1M<n<10M147 likes21k downloads6d agoHugging Face05MMInstruction /ArxivCap Dataset Card for ArxivCap Data Instances Example-1 of single (image, caption) pairs "......" stands for omitted parts. { 'src': 'arXiv_src_2112_060/2112.08947', 'meta': { 'meta_from_kaggle': { 'journey': '', 'license': 'http://arxiv.org/licenses/nonexclusive-distrib/1.0/', 'categories': 'cs.ET' }, 'meta_from_s2': { 'citationCount': 8… See the full description on the dataset page: https://huggingface.co/datasets/MMInstruction/ArxivCap.imageimage-to-text100K<n<1M58 likes17k downloads2y agoHugging Face06obswork /arxiv-ai-ml-100k-papers license: other tags: - arxiv - ocr - machine-learning --- # obswork/arxiv-ai-ml-100k A 99,999-paper stratified subset of [`Rendra8631/arxiv-papers`](https://huggingface.co/datasets/Rendra8631/arxiv-papers) at revision `a2c6afb51332d2744b46308df6917697582f8cd4`, filtered to the primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. Only papers submitted in `2023`-`2025` are included. This dataset is a build artifact of the OCR… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-papers.1 likes14k downloads5mo agoHugging Face07SlowGuess /Arxiv_2025_OCR Arxiv_2025_OCR OCR Data. 0 likes12k downloads8mo agoHugging Face08taesiri /arxiv_db4 likes8.9k downloads2y agoHugging Face09SlowGuess /Arxiv_2024_OCR Arxiv_2024_OCR OCR Data. 0 likes8.6k downloads8mo agoHugging Face10ccdv /arxiv-summarization Arxiv dataset for summarization Dataset for summarization of long documents.Adapted from this repo.Note that original data are pre-tokenized so this dataset returns " ".join(text) and add "\n" for paragraphs. This dataset is compatible with the run_summarization.py script from Transformers if you add this line to the summarization_name_mapping variable: "ccdv/arxiv-summarization": ("article", "abstract") Data Fields id: paper id article: a string containing the body of… See the full description on the dataset page: https://huggingface.co/datasets/ccdv/arxiv-summarization.textsummarization100K<n<1M136 likes8.2k downloads2y agoHugging Face11jamescalam /ai-arxiv2-chunkstext100K<n<1M4 likes7.9k downloads3y agoHugging Face12taesiri /arxiv_qa25 likes7.7k downloads2y agoHugging Face13mlfoundations /MINT-1T-ArXiv 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-ArXiv.imageimage-to-text1M<n<10M61 likes7.5k downloads2y agoHugging Face14SlowGuess /Arxiv_2023_OCR Arxiv_2023_OCR OCR Data. 0 likes7.5k downloads8mo agoHugging Face15SlowGuess /Arxiv_2021_OCR Arxiv_2021_OCR OCR Data. 0 likes7.4k downloads9mo agoHugging Face16SlowGuess /Arxiv_2020_OCR Arxiv_2020_OCR OCR Data. 0 likes7.1k downloads8mo agoHugging Face17TIGER-Lab /arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot. image10M<n<100M5 likes5.8k downloads1y agoHugging Face18obswork /arxiv-ai-ml-100k-pages license: other tags: - arxiv - ocr - machine-learning --- # obswork/arxiv-ai-ml-100k-pages A **page-bounded** stratified subset of the raw pool dataset [`obswork/arxiv-ai-ml-100k`](https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k), filtered to primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. The raw pool is itself a 100k-paper stratified sample from… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-pages.0 likes5.6k downloads5mo agoHugging Face19nick007x /arxiv-papersdocument1M<n<10M202 likes5.5k downloads6mo agoHugging Face20ChristophSchuhmann /1-sentence-level-gutenberg-en_arxiv_pubmed_sodatext100M<n<1B1 likes5.3k downloads3y agoHugging Face21kalomaze /alphabetic-arxiv-authors-it1text100K<n<1M0 likes5k downloads1y agoHugging Face22librarian-bots /arxiv-metadata-snapshot Dataset Card for "arxiv-metadata-oai-snapshot" More Information needed This is a mirror of the metadata portion of the arXiv dataset. The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset. Metadata This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing: id: ArXiv ID (can be used to access the paper, see below) submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.texttext-generation1M<n<10M22 likes4.8k downloads3d agoHugging Face23Xia-2004 /arx-left-cube ARX Left Cube YuHai HDF5 Dataset This dataset contains ARX-X5 left-arm single-arm teleoperation episodes for a cube manipulation task. Each episode is stored as one HDF5 file. https://huggingface.co/datasets/Xia-2004/arx-left-cube Files episode_000000.hdf5 ... episode_000200.hdf5 201 episodes 63,619 total frames RGB images are stored as uint8 Actions are stored as float32 HDF5 Format Each episode_*.hdf5 contains: Key Shape Dtype Meaning action (T… See the full description on the dataset page: https://huggingface.co/datasets/Xia-2004/arx-left-cube.roboticsn<1K0 likes4.4k downloads4mo agoHugging Face24mteb /arxiv-clustering-s2s ArXivHierarchicalClusteringS2S An MTEB dataset Massive Text Embedding Benchmark Clustering of titles from arxiv. Clustering of 30 sets, either on the main or secondary category Task category t2c Domains Academic, Written Reference https://www.kaggle.com/Cornell-University/arxiv How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["ArXivHierarchicalClusteringS2S"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arxiv-clustering-s2s.texttext-classificationn<1K1 likes4.1k downloads7mo agoHugging Face25CShorten /ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning. The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering. The dataset is maintained by with requests to the ArXiv API. The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.tabular100K<n<1M72 likes4k downloads4y agoHugging Face26common-pile /arxiv_papers ArXiv Papers Description ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more. When uploading papers, authors can choose from a variety of licenses. This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline: first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.texttext-generation100K<n<1M17 likes3.7k downloads1y agoHugging Face27mteb /arxiv-clustering-p2p ArXivHierarchicalClusteringP2P An MTEB dataset Massive Text Embedding Benchmark Clustering of titles+abstract from arxiv. Clustering of 30 sets, either on the main or secondary category Task category t2c Domains Academic, Written Reference https://www.kaggle.com/Cornell-University/arxiv How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/arxiv-clustering-p2p.texttext-classificationn<1K3 likes3.6k downloads7mo agoHugging Face28gfissore /arxiv-abstracts-2021 Dataset Card for arxiv-abstracts-2021 Dataset Summary A dataset of metadata including title and abstract for all arXiv articles up to the end of 2021 (~2 million papers). Possible applications include trend analysis, paper recommender engines, category prediction, knowledge graph construction and semantic search interfaces. In contrast to arxiv_dataset, this dataset doesn't include papers submitted to arXiv after 2021 and it doesn't require any external download.… See the full description on the dataset page: https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021.textsummarization1M<n<10M41 likes3.5k downloads4y agoHugging Face29MathArena /arxivmath Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary This dataset contains the questions from ArXivMath used for the MathArena Leaderboard Data Fields The dataset contains the following fields: problem_idx (int64): Problem index within the corresponding MathArena benchmark. answer (string): Gold final answer. problem_type (list[string]): Problem type/category labels. source (float64): arXiv… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/arxivmath.tabularn<1K0 likes3.1k downloads4mo agoHugging Face30AlgorithmicResearchGroup /arxiv_s2orc_parsed Dataset Card for "ArtifactAI/arxiv_s2orc_parsed" Dataset Description https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_s2orc_parsed Dataset Summary AlgorithmicResearchGroup/arxiv_s2orc_parsed is a subset of the AllenAI S2ORC dataset, a general-purpose corpus for NLP and text mining research over scientific papers, The dataset is filtered strictly for ArXiv papers, including the full text for each paper. Github links have been extracted… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_s2orc_parsed.texttext-generation1M<n<10M28 likes3k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.