CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01permutans /arxiv-papers-by-subject arXiv Papers by Subject A reorganised version of the nick007x/arxiv-papers dataset, partitioned by subject code, year, and month for efficient selective access. Dataset Description This dataset contains metadata for over 2.5 million arXiv papers, organised into a hierarchical directory structure that allows users to download only the specific subjects and time periods they need, rather than the entire dataset. Motivation The original… See the full description on the dataset page: https://huggingface.co/datasets/permutans/arxiv-papers-by-subject.text-generation1M<n<10M35 likes88k downloads9mo agoHugging Face02tegridydev /research-papers research-papers Dataset Overview The Research Papers Dataset is a collection of academic research documents categorized by their primary research topic. This dataset is designed for tasks such as model finetuning, document classification, optical character recognition (OCR) testing and multimodal document understanding (Feel free to use it however you see fit!). Curated by: tegridy Language: English Format: PDF | MD Repo Structure The dataset… See the full description on the dataset page: https://huggingface.co/datasets/tegridydev/research-papers.documentimage-classificationn<1K19 likes15k downloads4mo agoHugging Face03hysts-bot-data /daily-papers-embeddingstext10K<n<100K8 likes14k downloads5h agoHugging Face04obswork /arxiv-ai-ml-100k-papers license: other tags: - arxiv - ocr - machine-learning --- # obswork/arxiv-ai-ml-100k A 99,999-paper stratified subset of [`Rendra8631/arxiv-papers`](https://huggingface.co/datasets/Rendra8631/arxiv-papers) at revision `a2c6afb51332d2744b46308df6917697582f8cd4`, filtered to the primary subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. Only papers submitted in `2023`-`2025` are included. This dataset is a build artifact of the OCR… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-papers.1 likes12k downloads5mo agoHugging Face05paperswithbacktest /Stocks-Daily-Price Stocks Daily Price This dataset includes daily price data for various stocks. 25,986,919 rows over 7,764 symbols, 8 columns, covering 1962-01-02 to 2026-08-05. Refreshed monthly. Strategies Built on This Data 2,401 papers in the Papers With Backtest catalogue declare this dataset as an input. 2,238 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.35, and 45% clear a t-statistic of 1.96 on their own sample… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Daily-Price.tabulartime-series-forecasting10M<n<100M65 likes11k downloads20d agoHugging Face06nick007x /arxiv-papersdocument1M<n<10M202 likes5.9k downloads6mo agoHugging Face07scientifi-papers /scientific-papers Scientific Papers - Raw Full Text ~57 million scientific papers with full text, extracted from multiple large-scale academic paper collections. This dataset provides raw full text suitable for pre-training, fine-tuning, or building search indices over scientific literature. Subsets Subset Papers Size Source papers-2 ~18.5M ~358 GB S2ORC papers collection (untitled subset) papers-3 ~27.4M ~198 GB S2ORC scientific-papers collection pes2o ~8.2M ~106 GB… See the full description on the dataset page: https://huggingface.co/datasets/scientifi-papers/scientific-papers.texttext-generation10M<n<100M2 likes5.5k downloads4mo agoHugging Face08paperswithbacktest /Universe-Daily-Pricegated Universe Daily Price The index of every symbol carried by the public daily price datasets, with the repository that holds it. 21,693 rows over 20,895 symbols, 4 columns. Updated by Papers With Backtest. Why It Matters This is the lookup table the other price datasets need: Routing: A symbol on its own does not say which file holds it. repo_id answers that in one join, so a strategy that mixes equities, futures and rates loads from the right place without… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Universe-Daily-Price.tabulartime-series-forecasting10K<n<100K0 likes5.5k downloads20d agoHugging Face09sci-papers /scientific-paperstext10M<n<100M0 likes5.4k downloads1y agoHugging Face10CShorten /ML-ArXiv-PapersThis dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about Machine Learning. The core dataset is filtered from the full ArXiv dataset hosted on Kaggle: https://www.kaggle.com/datasets/Cornell-University/arxiv. The original dataset contains roughly 2 million papers. This dataset contains roughly 100,000 papers following the category filtering. The dataset is maintained by with requests to the ArXiv API. The current iteration of the dataset only contains… See the full description on the dataset page: https://huggingface.co/datasets/CShorten/ML-ArXiv-Papers.tabular100K<n<1M72 likes4k downloads4y agoHugging Face11common-pile /arxiv_papers ArXiv Papers Description ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more. When uploading papers, authors can choose from a variety of licenses. This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline: first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers.texttext-generation100K<n<1M17 likes3.5k downloads1y agoHugging Face12armanc /scientific_papersScientific papers datasets contains two sets of long and structured documents. The datasets are obtained from ArXiv and PubMed OpenAccess repositories. Both "arxiv" and "pubmed" have two features: - article: the body of the document, pagragraphs seperated by "/n". - abstract: the abstract of the document, pagragraphs seperated by "/n". - section_names: titles of sections, seperated by "/n".summarization100K<n<1M177 likes3.4k downloads3y agoHugging Face13paperswithbacktest /Forex-Daily-Pricegated Forex Daily Price This dataset includes daily price data for various FX pairs. 437,430 rows over 166 symbols, 7 columns, covering 1970-01-04 to 2026-07-31. Refreshed monthly. Strategies Built on This Data 292 papers in the Papers With Backtest catalogue declare this dataset as an input. 260 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.21, and 34% clear a t-statistic of 1.96 on their own sample, against 48%… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Forex-Daily-Price.tabulartime-series-forecasting100K<n<1M5 likes3.1k downloads20d agoHugging Face14hysts-bot-data /daily-paperstext10K<n<100K17 likes3k downloads5h agoHugging Face15paperswithbacktest /Indices-Daily-Pricegated Indices Daily Price This dataset includes daily price data for various indices. 815,441 rows over 113 symbols, 8 columns, covering 1927-12-30 to 2026-08-03. Refreshed monthly. Strategies Built on This Data 1,324 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,226 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.49, and 62% clear a t-statistic of 1.96 on their own sample, against… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Indices-Daily-Price.tabulartime-series-forecasting100K<n<1M2 likes2.8k downloads20d agoHugging Face16hysts-bot-data /daily-papers-stats3 likes2.7k downloads5m agoHugging Face17LqJia /Papersdocument1K<n<10K1 likes2.5k downloads1mo agoHugging Face18common-pile /arxiv_papers_filtered ArXiv Papers Description ArXiv is an online open-access repository of over 2.4 million scholarly papers covering fields such as computer science, mathematics, physics, quantitative biology, economics, and more. When uploading papers, authors can choose from a variety of licenses. This dataset includes text from all papers uploaded under CC BY, CC BY-SA, and CC0 licenses through a three-step pipeline: first, the latex source files for openly licensed papers were… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_papers_filtered.text100K<n<1M13 likes2.3k downloads1y agoHugging Face19ashishkgpian /arxiv_papers_0_50_190 likes2k downloads1y agoHugging Face20MongoDB /subset_arxiv_papers_with_embeddingsThis dataset is a curated subset of the original arXiv dataset, each entry enriched with a 256-dimensional embedding vector. The embeddings are generated using OpenAI's "text-embedding-3-small" model. For each data point, the embedding is created by concatenating the text of the title, author(s), and abstract into a single string, which is then processed by the embedding model. This approach captures the semantic essence of each document, facilitating tasks such as similarity search… See the full description on the dataset page: https://huggingface.co/datasets/MongoDB/subset_arxiv_papers_with_embeddings.text10K<n<100K2 likes2k downloads2y agoHugging Face21GenAI4ELab /papercli-papers-acldocument10K<n<100K1 likes2k downloads3mo agoHugging Face22ashishkgpian /0_50_arxiv_papers_17429083840 likes1.9k downloads1y agoHugging Face23uynitsuj /paper-sim-n512-traces paper-sim-n512-traces Rollout traces for the n=512 simulated bottle-in-bin evaluation reported in Table I of the WARP-RM CoRL rebuttal. Two arms, 512 paired scenes each: directory prefix arm bottles/scene throughput all-6 fullhz_vanilla_sh00..15 Vanilla BC (100% of data) 3.885 237/hr 9.4% fullhz_paperwarp512_sh00..15 WARP-BC (31.5% kept) 4.533 290/hr 25.0% 512 traces per arm = 16 shards x 32 worlds. Each qpos_trace_NNN_sSEED.npz holds the full qpos trajectory… See the full description on the dataset page: https://huggingface.co/datasets/uynitsuj/paper-sim-n512-traces.robotics0 likes1.9k downloads1mo agoHugging Face24ashishkgpian /0_50_arxiv_papers_17428993580 likes1.8k downloads1y agoHugging Face25ashishkgpian /arxiv_papers_0_50_220 likes1.8k downloads1y agoHugging Face26paperswithbacktest /ETFs-Daily-Pricegated ETFs Daily Price This dataset includes daily price data for various ETFs. 10,813,800 rows over 5,802 symbols, 8 columns, covering 1978-01-03 to 2026-08-03. Refreshed monthly. Strategies Built on This Data 502 papers in the Papers With Backtest catalogue declare this dataset as an input. 476 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.55, and 70% clear a t-statistic of 1.96 on their own sample, against 48%… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/ETFs-Daily-Price.tabulartime-series-forecasting10M<n<100M3 likes1.7k downloads20d agoHugging Face27yufan /top-conference-papers Top Conference Papers (2024-2026) Paper metadata from 10 top computer-science conferences (ML / NLP / IR / Data Mining / RecSys), covering 2024-2026. 38,886 papers in total. Structure config = conference: ACL, CIKM, ICLR, ICML, KDD, NeurIPS, RecSys, SIGIR, WSDM, WWW split = year: 2024, 2025, and 2026 where available from datasets import load_dataset ds = load_dataset("yufan/top-conference-papers", "ICLR") # one conference ds["2024"]… See the full description on the dataset page: https://huggingface.co/datasets/yufan/top-conference-papers.document10K<n<100K1 likes1.7k downloads25d agoHugging Face28ashishkgpian /arxiv_papers_50_100_190 likes1.6k downloads1y agoHugging Face29paperswithbacktest /Stocks-Quarterly-BalanceSheetgated Stocks Quarterly Balance Sheet This dataset includes quarterly balance sheet data for various stocks. 342,800 rows over 6,893 symbols, 39 columns, covering 1983-06-30 to 2026-06-30. Refreshed monthly. Strategies Built on This Data 1,067 papers in the Papers With Backtest catalogue declare this dataset as an input. 1,012 of them have been coded and run over their own full history. The median replicated Sharpe ratio is +0.27, and 37% clear a t-statistic of 1.96 on… See the full description on the dataset page: https://huggingface.co/datasets/paperswithbacktest/Stocks-Quarterly-BalanceSheet.time-series-forecasting100K<n<1M2 likes1.6k downloads20d agoHugging Face30ashishkgpian /0_50_arxiv_papers_17429066670 likes1.6k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.