CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tomasg25 /scientific_lay_summarisationThis repository contains the PLOS and eLife datasets, introduced in the EMNLP 2022 paper "[Making Science Simple: Corpora for the Lay Summarisation of Scientific Literature ](https://arxiv.org/abs/2210.09932)". Each dataset contains full biomedical research articles paired with expert-written lay summaries (i.e., non-technical summaries). PLOS articles are derived from various journals published by [the Public Library of Science (PLOS)](https://plos.org/), whereas eLife articles are derived from the [eLife](https://elifesciences.org/) journal. More details/anlaysis on the content of each dataset are provided in the paper. Both "elife" and "plos" have 6 features: - "article": the body of the document (including the abstract), sections seperated by "/n". - "section_headings": the title of each section, seperated by "/n". - "keywords": keywords describing the topic of the article, seperated by "/n". - "title" : the title of the article. - "year" : the year the article was published. - "summary": the lay summary of the document.summarization10K<n<100K20 likes288 downloads2y agoHugging Face02UCL-DARK /openai-tldr-summarisation-preferences Human feedback data This is the version of the dataset used in https://arxiv.org/abs/2310.06452. If starting a new project we would recommend using https://huggingface.co/datasets/openai/summarize_from_feedback. See https://github.com/openai/summarize-from-feedback for original details of the dataset. Here the data is formatted to enable huggingface transformers sequence classification models to be trained as reward functions. texttext-classification100K<n<1M2 likes156 downloads3y agoHugging Face03pszemraj /scientific_lay_summarisation-plos-norm scientific_lay_summarisation - PLOS - normalized This dataset is a modified version of tomasg25/scientific_lay_summarization and contains scientific lay summaries that have been preprocessed with this code. The preprocessing includes fixing punctuation and whitespace problems, and calculating the token length of each text sample using a tokenizer from the T5 model. Original dataset details: Repository: https://github.com/TGoldsack1/Corpora_for_Lay_Summarisation Paper: Making… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/scientific_lay_summarisation-plos-norm.tabularsummarization10K<n<100K10 likes115 downloads9mo agoHugging Face04Columbia-NLP /DPO-tldr-summarisation-preferences Dataset Card for DPO-tldr-summarisation-preferences Reformatted from openai/summarize_from_feedback dataset. The LION-series are trained using an empirically optimized pipeline that consists of three stages: SFT, DPO, and online preference learning (online DPO). We find simple techniques such as sequence packing, loss masking in SFT, increasing the preference dataset size in DPO, and online DPO training can significantly improve the performance of language models. Our best models… See the full description on the dataset page: https://huggingface.co/datasets/Columbia-NLP/DPO-tldr-summarisation-preferences.tabular100K<n<1M1 likes77 downloads2y agoHugging Face05pszemraj /scientific_lay_summarisation-elife-norm scientific_lay_summarisation - elife - normalized This is the "elife" split. For more words, refer to the PLOS split README Contents load with datasets: from datasets import load_dataset # If the dataset is gated/private, make sure you have run huggingface-cli login dataset = load_dataset("pszemraj/scientific_lay_summarisation-elife-norm") dataset Output: DatasetDict({ train: Dataset({ features: ['article', 'summary', 'section_headings', 'keywords', 'year'… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/scientific_lay_summarisation-elife-norm.tabularsummarization1K<n<10K8 likes72 downloads9mo agoHugging Face06rumalekanayake /pmc-medical-summarisationtext100K<n<1M1 likes60 downloads8mo agoHugging Face07gayanin /pubmed-gastro-summarisation0 likes35 downloads5y agoHugging Face08VytautoDidziojoUniversitetas /LT_Summarisation_Corpus Lithuanian Summarisation Corpus Description Lithuanian texts paired with human-written abstractive and extractive summaries, spanning four subject domains — information technology, law, medicine, and media. The corpus is designed for training and evaluating summarisation systems on Lithuanian-language content, covering both technical and general registers. Dataset summary Types of Dataset: CSV, JSON, XML Number of CSV files: 8 (4 training, 4… See the full description on the dataset page: https://huggingface.co/datasets/VytautoDidziojoUniversitetas/LT_Summarisation_Corpus.texttext-generation1K<n<10K0 likes28 downloads4mo agoHugging Face09ekacare /ekacare_medical_history_summarisationgated Medical History Summarization Dataset This dataset contains 58 diverse medical cases from EkaCare's internal medical team and the company's employees and their family members. Each case provides a complete view of a patient's medical journey, including vital sign trends and historical health context. The dataset has been designed to evaluate the ability to summarize complex medical histories and extract clinically relevant insights. For details read this blog. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/ekacare_medical_history_summarisation.textsummarizationn<1K0 likes27 downloads1y agoHugging Face10when2rl /tldr-summarisation-preferences_reformatted Dataset Card for when2rl/tldr-summarisation-preferences_reformatted Reformatted from UCL-DARK/openai-tldr-summarisation-preferences to be consistent with all other datasets in this org. Note that since the original dataset does not have score ratings, we used "10" for ALL chosen response, and "1" for ALL rejected response as dummy scores. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/when2rl/tldr-summarisation-preferences_reformatted.tabular100K<n<1M0 likes25 downloads2y agoHugging Face11Hashif /judgement-summarisation-llama-2text1K<n<10K0 likes24 downloads3y agoHugging Face12majorSeaweed /toxic-dialogue-summarisation Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [Aditya Singh] Language(s) (NLP): [English] License: [MIT] Dataset Sources [Grok3] Repository: [More Information Needed] Demo [optional]: [More Information Needed] Uses Direct Use [More Information Needed] Out-of-Scope… See the full description on the dataset page: https://huggingface.co/datasets/majorSeaweed/toxic-dialogue-summarisation.tabulartext-classification1K<n<10K1 likes15 downloads1y agoHugging Face13extraordinarylab /scientific-lay-summarisationtext10K<n<100K0 likes12 downloads11mo agoHugging Face14YashaP /Summarisation_datasettext1K<n<10K0 likes11 downloads3y agoHugging Face15KevinEgan /whatsapp-summarisation-datasettextn<1K0 likes11 downloads6mo agoHugging Face16dhiya96 /zephyr_text_summarisation_500textn<1K0 likes10 downloads3y agoHugging Face17Navneeth017 /Malayalam_Summarisationtext100K<n<1M0 likes10 downloads4mo agoHugging Face18nickykay /summarisationtextn<1K0 likes5 downloads3y agoHugging Face19jq /salt-summarisationtext100K<n<1M0 likes5 downloads2y agoHugging Face20sirano1004 /summarisationtabular10K<n<100K0 likes5 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.