CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /Wikipedia-AbstractWikipedia Abstract Introducing Wikipedia Abstract, a comprehensive dataset encompassing abstracts, complete articles, and a popularity score index for both widely spoken and lesser-known Wikipedia subsets. Our dedication to Wikipedia-X ensures a centralized Wikipedia dataset that undergoes regular updates and adheres to the highest standards. A central focus of our efforts was to include exotic languages that often lack up-to-date Wikipedia dumps or may not have any dumps at all.… See the full description on the dataset page: https://huggingface.co/datasets/laion/Wikipedia-Abstract.texttext-classification10M<n<100M9 likes4.3k downloads2y agoHugging Face02aisingapore /NLG-Abstractive-Summarizationgated SEA Abstractive Summarization SEA Abstractive Summarization evaluates a model's ability to read a document, identify the key points within, and summarize them into a coherent and fluent text while paraphrasing the document. It is sampled from XL-Sum for Indonesian, Tamil, Thai, and Vietnamese. Supported Tasks and Leaderboards SEA Abstractive Summarization is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLG-Abstractive-Summarization.texttext-generationn<1K0 likes2.3k downloads9mo agoHugging Face03UniverseTBD /arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.texttext-generation1M<n<10M7 likes1.7k downloads3y agoHugging Face04common-pile /arxiv_abstracts ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts.texttext-generation1M<n<10M13 likes786 downloads1y agoHugging Face05common-pile /arxiv_abstracts_filtered ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts_filtered.texttext-generation1M<n<10M9 likes589 downloads10mo agoHugging Face06AbstractPhil /human-templated-captions-1bcsv delimiter is = ".,|,." apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading. This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon. Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.texttext-generation100M<n<1B1 likes325 downloads1y agoHugging Face07AbstractPhil /wordnet-lexical-topology WordNet Lexical Topology Dataset Dataset Summary The WordNet Lexical Topology Dataset provides comprehensive n-gram frequency analysis from multiple sources: NLTK WordNet: Original Princeton WordNet with 117,659 synsets HF WordNet: Frequency-weighted definitions from 864,894 entries with cardinality data Unicode: Character names from 143,041 Unicode codepoints This dataset preserves sequential information crucial for language modeling and text generation, with over 12… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-lexical-topology.tabulartext-generation10M<n<100M2 likes219 downloads1y agoHugging Face08Lots-of-LoRAs /task664_mmmlu_answer_generation_abstract_algebra Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task664_mmmlu_answer_generation_abstract_algebra Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task664_mmmlu_answer_generation_abstract_algebra.texttext-generationn<1K0 likes119 downloads2y agoHugging Face09datajuicer /the-pile-pubmed-abstracts-refined-by-data-juicer The Pile -- PubMed Abstracts (refined by Data-Juicer) A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.texttext-generationn<1K3 likes116 downloads3y agoHugging Face10AbstractPhil /random-captions-10mRandomly generated captions using tokenization templates and lists. .,|,. is the caption delimiter, so split accordingly. texttext-generationn<1K0 likes105 downloads1y agoHugging Face11AbstractPhil /json-coco-format JSON COCO Format — task-differentiated SFT data A multi-task supervised fine-tuning dataset that teaches a model to convert image-synthesis caption prompts into JSON whose structure varies by task. Built from MS-COCO captions (Karpathy split) with Claude Sonnet 4.6 as the teacher; designed for training per-task LoRAs on Qwen/Qwen3.5-0.8B. Each row is in the Qwen3.5-native tool-call shape: a messages array with an assistant turn whose tool_calls[0].function.arguments is a dict… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/json-coco-format.texttext-generation100K<n<1M0 likes80 downloads4mo agoHugging Face12Lots-of-LoRAs /task619_ohsumed_abstract_title_generation Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task619_ohsumed_abstract_title_generation Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task619_ohsumed_abstract_title_generation.texttext-generation1K<n<10K0 likes72 downloads2y agoHugging Face13beta3 /3M_Academic_Papers_Titles_and_Abstracts Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts 📋 Overview This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.texttext-classification1M<n<10M2 likes71 downloads1y agoHugging Face14fromziro /arxiv-abstracts-2004 ArXiv Abstracts 2004 Original Dataset: common-pile/arxiv_abstracts ArXiv-Abstracts-2004 is a filtered collection of abstracts from the Common-Pile ArXiv dataset containing works created on or before 2004. Stats Size (MB) Lines 351MB 303,761 Note: The lines, in the .jsonl file, are ordered from oldest to newest. Notice We do not claim ownership of or credit for any prior work done by the Common-Pile team. This dataset is only a… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/arxiv-abstracts-2004.texttext-generation100K<n<1M1 likes49 downloads2mo agoHugging Face15PursuitOfDataScience /arxiv-llama4-maverick-abstract arXiv Abstract Dataset (Llama-4-Maverick-17B-128E-Instruct-FP8) Dataset Description This dataset contains high-quality abstracts for scientific papers from the arXiv repository, generated using the Llama-4-Maverick-17B-128E-Instruct-FP8 model. Each abstract provides a concise, accurate overview of the research paper while preserving key technical details and contributions. Dataset Features High-quality abstracts: Generated using… See the full description on the dataset page: https://huggingface.co/datasets/PursuitOfDataScience/arxiv-llama4-maverick-abstract.textsummarization100K<n<1M1 likes45 downloads1y agoHugging Face16ufal /bilingual-abstracts-corpus ÚFAL Bilingual Abstracts Corpus This is a parallel (bilingual) corpus of Czech and mostly English abstracts of scientific papers and presentations published by authors from the Institute of Formal and Applied Linguistics, Charles University in Prague. For each publication record, the authors are obliged to provide both the original abstract (in Czech or English), and its translation (English or Czech) in the internal Biblio system. The data was filtered for duplicates and missing… See the full description on the dataset page: https://huggingface.co/datasets/ufal/bilingual-abstracts-corpus.texttranslation1K<n<10K4 likes37 downloads3y agoHugging Face17AbstractPhil /cc-prompts-sharded Conceptual Captions — sharded for the Qwen task-conversion pipeline Source: Conceptual Captions (conceptual_captions), deduplicated, minimum 4 words, split into three roughly-equal shards plus a "long" shard for captions over 50 words (reserved for later high-capacity model processing). Each row: {"id": "cc_00000123", "caption": "<text>", "n_words": <int>} id is the position of the row in the CC stream, zero-padded to 8 digits. Stable across re-runs. Used by the downstream… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/cc-prompts-sharded.texttext-generation1M<n<10M0 likes31 downloads4mo agoHugging Face18alexshengzhili /Abstract2Appendix_v1_10k Dataset Card: Abstract2Appendix v1 Dataset Description The Abstract2Appendix v1 dataset is a high-quality collection of academic peer reviews and their associated research paper metadata. This dataset combines reviews from four premier machine learning and AI conferences: NeurIPS 2023, EMNLP 2023, TMLR, and ICLR 2023, shuffled into a unified corpus. It is designed to enhance long-context capabilities in Large Language Models (LLMs) and supports tasks such as fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/alexshengzhili/Abstract2Appendix_v1_10k.tabulartext-generation1K<n<10K3 likes29 downloads2y agoHugging Face19onurkeles /econ_paper_abstracts Dataset Card for Economics Paper Dataset Dataset Summary The Economics Research Paper Dataset was designed to support the development of the LLaMA-2-Econ models, with a focus on Title Generation, Abstract Classification, and Question & Answer (Q&A) tasks. It comprises abstracts and titles of economics research papers, along with synthetic Q&A pairs derived from the abstracts, to facilitate training of large language models for economics-specific applications.… See the full description on the dataset page: https://huggingface.co/datasets/onurkeles/econ_paper_abstracts.texttext-classification1K<n<10K0 likes25 downloads3y agoHugging Face20AbstractPhil /wordnet-definitions WordNet Multiple Definitions - Columnar Format Overview This dataset is an optimized columnar version of WordNet multiple definitions, designed for high-performance queries and rapid extraction. Each definition was sourced by GPT-5 Nano. I may update this to include additional definitions in the future, but I will not break the format. The original dataset has a more unabridged and noisy set of data; so I'm definitely going to leave it intact. Noisy training is important… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/wordnet-definitions.tabulartext-generation100K<n<1M1 likes23 downloads1y agoHugging Face21ClarusC64 /abstraction-level-stability-v01Cardinal Meta Dataset 3.1Abstraction Level Stability Purpose Test whether claims stay at the correct abstraction level Test whether level changes are named and justified Test whether concrete cases are not inflated into general truths Central question What level is this claim operating at What this dataset catches Instance to general jumps Proxy to property inflation Model to reality reification Short term change treated as long term trend Principle treated as effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/abstraction-level-stability-v01.texttext-generationn<1K0 likes18 downloads8mo agoHugging Face22yashm /pubmed-bioinformatics-abstractsgated PubMed Bioinformatics Abstracts (2020–2026) A curated collection of 174,154 bioinformatics journal article abstracts from PubMed, covering publications from 2020 to 2026. Curated by: Dr. Yash M Gupta Intended Use Fine-tuning large language models (LLMs) on biomedical / bioinformatics domain text. Dataset Structure Column Type Description pmid string PubMed ID title string Article title abstract string Full abstract text year string Publication… See the full description on the dataset page: https://huggingface.co/datasets/yashm/pubmed-bioinformatics-abstracts.texttext-generation100K<n<1M2 likes7 downloads5mo agoHugging Face23yashm /pubmed-bioinformatics-abstracts_all_years PubMed Bioinformatics Abstracts — All Years A curated dataset of bioinformatics research abstracts fetched from PubMed via the NCBI Entrez API, covering publications from 1990 through 2026. Designed for LLM fine-tuning, biomedical NLP, and scientific text analysis. Curated by: Dr. Yash M Gupta Pipeline: yashmgupta/pubmed-bioinformatics-abstracts Intended Use Fine-tuning large language models (LLMs) on biomedical / bioinformatics domain text. Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/yashm/pubmed-bioinformatics-abstracts_all_years.texttext-generation100K<n<1M2 likes7 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.