CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AbstractPhil /human-templated-captions-1bcsv delimiter is = ".,|,." apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading. This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon. Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.texttext-generation100M<n<1B1 likes328 downloads1y agoHugging Face02brainchalov /pubmed_arxiv_abstracts_dataannotations_creators: machine-generated language: en license: apache-2.0 multilinguality: monolingual task_categories: classification generation pretty_name: PubMed_ArXiv_Abstracts features: name: abstr dtype: string name: title dtype: string name: journal dtype: string name: field dtype: string name: label_journal dtype: int64 name: label_field dtype: int64 tabular100K<n<1M13 likes156 downloads3y agoHugging Face03austin /rheum_abstracts Dataset Card for Rheumatology Abstracts Data Source This dataset comes from PubMed, derived from my fork of the pymed package (no longer maintained). My fork can be found at https://github.com/cmcmaster1/pymed Data Structure The dataset is split into train (80%) and test (20%) files (CSV). Each file contains three columns: id abstract (minus conclusion) conclusion tabular10K<n<100K2 likes139 downloads5y agoHugging Face04NicolaiSivesind /ChatGPT-Research-Abstracts ChatGPT-Research-Abstracts This is a dataset created in relation to a bachelor thesis written by Nicolai Thorer Sivesind and Andreas Bentzen Winje. It contains human-produced and machine-generated text samples of scientific research abstracts. A reformatted version for text-classification is available in the dataset collection Human-vs-Machine. In this collection, all samples are split into separate data points for real and generated, and labeled either 0 (human-produced) or 1… See the full description on the dataset page: https://huggingface.co/datasets/NicolaiSivesind/ChatGPT-Research-Abstracts.tabulartext-classification10K<n<100K5 likes122 downloads3y agoHugging Face05Bambolbb /cits4012_A1_2026_medical_abstractstext10K<n<100K0 likes99 downloads18d agoHugging Face06TMDuvall /MECFS-PubMed-AbstractThis table contains every abstract in NIH's PubMed database containing either "myalgic encephalomyelitis" or "chronic fatigue syndrome", along with metadata. text1K<n<10K1 likes81 downloads6d agoHugging Face07MIT-WAL /ai-jobs-news-articles-abstracts News articles and research abstracts on AI, labor, and jobs Dataset summary This file is a standalone CSV of news articles (full scraped text) and scholarly paper abstracts curated for research on artificial intelligence, work, and labor markets. Each row is one document: a stable id, publication date, normalized title and main text, and a small metadata dictionary. Rows: 53,526 document_class Rows Approx. date range (date column) news 29,857 Jan. 2025… See the full description on the dataset page: https://huggingface.co/datasets/MIT-WAL/ai-jobs-news-articles-abstracts.text10K<n<100K1 likes70 downloads2mo agoHugging Face08batterydata /paper-abstracts Battery Abstracts Dataset This dataset includes 29,472 battery papers and 17,191 non-battery papers, a total of 46,663 papers. These papers are manually labelled in terms of the journals to which they belong. 14 battery journals and 1,044 non battery journals were selected to form this database. training_data.csv: Battery papers: 20,629, Non-battery papers: 12,034. Total: 32,663. val_data.csv: Battery papers: 5,895, Non-battery papers: 3,438. Total: 9,333. test_data.csv: Battery… See the full description on the dataset page: https://huggingface.co/datasets/batterydata/paper-abstracts.texttext-classification10K<n<100K10 likes63 downloads4y agoHugging Face09beta3 /3M_Academic_Papers_Titles_and_Abstracts Comprehensive Academic Papers Dataset: 3M+ Research Paper Titles and Abstracts 📋 Overview This dataset is a comprehensive collection of over 3 million research paper titles and abstracts, curated and consolidated from multiple high-quality academic sources. The dataset provides a unified, clean, and standardized format for researchers, data scientists, and machine learning practitioners working on natural language processing, academic research analysis, and knowledge… See the full description on the dataset page: https://huggingface.co/datasets/beta3/3M_Academic_Papers_Titles_and_Abstracts.texttext-classification1M<n<10M2 likes63 downloads1y agoHugging Face10ZWG817 /Abstracttext100K<n<1M0 likes53 downloads2y agoHugging Face11ABSTRACTION-ERC /subCat-human SubCat: A Dataset of Subordinate Categories in Human Mind and LLMs for the Italian Language A psycholinguistic italian dataset released with the paper How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian. It contains a list of subordiante categories, or exemplars, for 187 concrete words or, basic-level categories. Dataset Creation The dataset was created to study how Italian L1 speakers generate exemplars for common… See the full description on the dataset page: https://huggingface.co/datasets/ABSTRACTION-ERC/subCat-human.tabular1K<n<10K0 likes53 downloads1y agoHugging Face12ABSTRACTION-ERC /subCat-llm SubCat: A Dataset of Subordinate Categories in Human Mind and LLMs for the Italian Language A psycholinguistic italian dataset released with the paper How Humans and LLMs Organize Conceptual Knowledge: Exploring Subordinate Categories in Italian. It contains a list of subordiante categories, or exemplars, for 187 concrete words or, basic-level categories. This repository contains the generations obtained by prompting a series of LLMs to replicate the human experiment. You can… See the full description on the dataset page: https://huggingface.co/datasets/ABSTRACTION-ERC/subCat-llm.tabular1K<n<10K0 likes51 downloads1y agoHugging Face13rachel6603 /gptneo-pubmed-abstracts Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description This is a dataset consisting of 10000 PubMed abstracts from The Pile (arXiv:2101.00027), along with completions (both human, and LLM-generated), in order to be used to calculate Heaps Law, in the manner described in the preliminary paper, Heaps' Law in GPT-Neo Large Language Model… See the full description on the dataset page: https://huggingface.co/datasets/rachel6603/gptneo-pubmed-abstracts.text10K<n<100K0 likes49 downloads2y agoHugging Face14Gaborandi /breast_cancer_pubmed_abstracts This Dataset has been downloaded from PubMed It has abstracts and titles that are related to Breast Cancer the data has been cleaned before uploading it could be used for any NLP task, such as Domain Adaptation text1K<n<10K3 likes45 downloads4y agoHugging Face15Gaborandi /Psychedelics_pubmed_abstracts This Dataset has been downloaded from PubMed It has abstracts and titles that are related to Psychedelics the data has been cleaned before uploading it could be used for any NLP task, such as Domain Adaptation text1K<n<10K4 likes43 downloads4y agoHugging Face16ClarusC64 /abstraction-level-category-control-meta-v01 Dataset ClarusC64/abstraction-level-category-control-meta-v01 This dataset tests one capability. Can a model keep claims at the correct abstraction leveland avoid category errors. Core rule Not every statement is the same kind of statement. A model must not treat analogy as mechanism description as causation values as facts models as reality aggregates as individuals A correct answer in the wrong category is still wrong. Canonical labels WITHIN_SCOPE… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/abstraction-level-category-control-meta-v01.texttext-classificationn<1K0 likes41 downloads8mo agoHugging Face17dms2ect /wikipedia_character_abstractstext1K<n<10K3 likes36 downloads4y agoHugging Face18Gaborandi /Lung_Cancer_pubmed_abstracts This Dataset has been downloaded from PubMed It has abstracts and titles that are related to Lung Cancer the data has been cleaned before uploading it could be used for any NLP task, such as Domain Adaptation text1K<n<10K3 likes28 downloads4y agoHugging Face19Prranshu /medrxiv_abstractstabular100K<n<1M0 likes26 downloads2y agoHugging Face20mabidan /econ_paper_abstracts_fa_translationtexttranslation1K<n<10K2 likes26 downloads2y agoHugging Face21ankitagr01 /dynamic_topic_modeling_arxiv_abstractstext10K<n<100K0 likes25 downloads2y agoHugging Face22onurkeles /econ_paper_abstracts Dataset Card for Economics Paper Dataset Dataset Summary The Economics Research Paper Dataset was designed to support the development of the LLaMA-2-Econ models, with a focus on Title Generation, Abstract Classification, and Question & Answer (Q&A) tasks. It comprises abstracts and titles of economics research papers, along with synthetic Q&A pairs derived from the abstracts, to facilitate training of large language models for economics-specific applications.… See the full description on the dataset page: https://huggingface.co/datasets/onurkeles/econ_paper_abstracts.texttext-classification1K<n<10K0 likes23 downloads2y agoHugging Face23pieetie /pubmed-abstract-summary Context This dataset contains 4,331 pairs of biomedical research abstracts and their one-sentence summaries. Each abstract is paired with a concise summary that captures the key findings, methods, and significance of the research. textsummarization1K<n<10K4 likes23 downloads2y agoHugging Face24jgammack /SAE-door-abstracts SAE-door-abstracts This dataset includes ~1,550 texts of abstracts of technical papers and journal articles from the SAE Mobilus database that cover the topics of automotive or aerospace doors, noise, acoustics, and vibrations. text1K<n<10K1 likes21 downloads4y agoHugging Face25Gaborandi /Alzheimer_pubmed_abstracts This Dataset has been downloaded from PubMed It has abstracts and titles that are related to Alzheimer's Disease the data has been cleaned before uploading it could be used for any NLP task, such as Domain Adaptation text1K<n<10K3 likes20 downloads4y agoHugging Face26ahmedabdelwahed /Medical_papers_title_and_abstract_NLP_datasetOriginally publised in kaggle. text1K<n<10K0 likes20 downloads2y agoHugging Face27Gaborandi /diabetes_mellitus_type2_pubmed_abstracts This Dataset has been downloaded from PubMed It has abstracts and titles that are related to type 2 DM the data has been cleaned before uploading it could be used for any NLP task, such as Domain Adaptation text1K<n<10K0 likes19 downloads4y agoHugging Face28Gaborandi /HIV_pubmed_abstracts This Dataset has been downloaded from PubMed It has abstracts and titles that are related to HIV the data has been cleaned before uploading it could be used for any NLP task, such as Domain Adaptation text1K<n<10K0 likes19 downloads4y agoHugging Face29ClarusC64 /abstraction-level-stability-v01Cardinal Meta Dataset 3.1Abstraction Level Stability Purpose Test whether claims stay at the correct abstraction level Test whether level changes are named and justified Test whether concrete cases are not inflated into general truths Central question What level is this claim operating at What this dataset catches Instance to general jumps Proxy to property inflation Model to reality reification Short term change treated as long term trend Principle treated as effectiveness… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/abstraction-level-stability-v01.texttext-generationn<1K0 likes18 downloads8mo agoHugging Face30Gaborandi /Stroke_pubmed_abstracts This Dataset has been downloaded from PubMed It has abstracts and titles that are related to Stroke the data has been cleaned before uploading it could be used for any NLP task, such as Domain Adaptation text1K<n<10K1 likes16 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.