CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UniverseTBD /arxiv-abstracts-largeThe arXiv Dataset is a comprehensive knowledge repository of 1.7 million scholarly articles drawn from the vast domains of physics, computer science, statistics, electrical engineering, quantitative biology, and economics among others. It provides open access to vital features such as article titles, authors, categories, abstracts, full text PDFs, and more. The dataset offers immense depth, allowing for exploration into various subdisciplines and interconnections between them. It serves as a… See the full description on the dataset page: https://huggingface.co/datasets/UniverseTBD/arxiv-abstracts-large.texttext-generation1M<n<10M7 likes1.7k downloads3y agoHugging Face02common-pile /arxiv_abstracts ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts.texttext-generation1M<n<10M13 likes779 downloads1y agoHugging Face03common-pile /arxiv_abstracts_filtered ArXiv Abstracts Description Each paper uploaded to ArXiv includes structured metadata fields, including an abstract summarizing the paper’s findings and contributions. According to ArXiv’s licensing policy, the metadata for any paper submitted to ArXiv is distributed under the CC0 license, regardless of the license of the paper itself. Thus, this dataset contains the abstract for every paper submitted to ArXiv through late 2024. We source the abstracts from ArXiv’s API… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/arxiv_abstracts_filtered.texttext-generation1M<n<10M9 likes577 downloads10mo agoHugging Face04datajuicer /the-pile-pubmed-abstracts-refined-by-data-juicer The Pile -- PubMed Abstracts (refined by Data-Juicer) A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.texttext-generationn<1K3 likes116 downloads3y agoHugging Face05fromziro /arxiv-abstracts-2004 ArXiv Abstracts 2004 Original Dataset: common-pile/arxiv_abstracts ArXiv-Abstracts-2004 is a filtered collection of abstracts from the Common-Pile ArXiv dataset containing works created on or before 2004. Stats Size (MB) Lines 351MB 303,761 Note: The lines, in the .jsonl file, are ordered from oldest to newest. Notice We do not claim ownership of or credit for any prior work done by the Common-Pile team. This dataset is only a… See the full description on the dataset page: https://huggingface.co/datasets/fromziro/arxiv-abstracts-2004.texttext-generation100K<n<1M1 likes47 downloads2mo agoHugging Face06ufal /bilingual-abstracts-corpus ÚFAL Bilingual Abstracts Corpus This is a parallel (bilingual) corpus of Czech and mostly English abstracts of scientific papers and presentations published by authors from the Institute of Formal and Applied Linguistics, Charles University in Prague. For each publication record, the authors are obliged to provide both the original abstract (in Czech or English), and its translation (English or Czech) in the internal Biblio system. The data was filtered for duplicates and missing… See the full description on the dataset page: https://huggingface.co/datasets/ufal/bilingual-abstracts-corpus.texttranslation1K<n<10K4 likes35 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.