CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CarperAI /openai_summarize_tldr Dataset Card for "openai_summarize_tldr" More Information needed text100K<n<1M32 likes5.4k downloads4y agoHugging Face02SabaPivot /KMMLU-Summarized-Chain_of_Thought Dataset Card for Condensed Chain-of-Thought KMMLU Dataset This dataset card provides detailed information about the condensed KMMLU dataset. The dataset has been summarized using Upstage's LLM: Solar-Pro to condense the original KMMLU training and development data while preserving its quality and usability. Additionally, a new column, 'chain_of_thought', has been introduced to align with the reasoning approach outlined in the paper "Chain-of-Thought Prompting Elicits Reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/SabaPivot/KMMLU-Summarized-Chain_of_Thought.tabular100K<n<1M1 likes3.5k downloads2y agoHugging Face03CarperAI /openai_summarize_comparisonstext100K<n<1M44 likes3.4k downloads4y agoHugging Face04openai /summarize_from_feedbackSummarize from Feedback contains the human feedback data released by the "Learning to summarize from human feedback" paper.text100K<n<1M221 likes2.3k downloads4y agoHugging Face05nyu-dice-lab /wavepulse-radio-summarized-transcripts WavePulse Radio Summarized Transcripts Dataset Summary WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions. The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.texttext-generation100K<n<1M1 likes1.4k downloads2y agoHugging Face06ai2-adapt-dev /summarize_from_feedback_smalltext1K<n<10K0 likes1.1k downloads2y agoHugging Face07vwxyzjn /summarize_from_feedback_tldr_3_filteredThis is the query dataset taken directly from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset textsummarization100K<n<1M1 likes515 downloads3y agoHugging Face08llama-duo /gemma7b-summarize-eval-by-gemini15flashtabular1K<n<10K1 likes319 downloads2y agoHugging Face09tmskss /linux-man-pages-tldr-summarized Dataset Card for linux-man-pages-tldr-summarized Dataset Summary This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages. Supported Tasks This dataset should be used to fine-tune language models for summarization tasks. textsummarizationn<1K9 likes311 downloads3y agoHugging Face10llama-duo /gemma7b-summarize-eval-by-claude3sonnettabular1K<n<10K0 likes310 downloads2y agoHugging Face11llama-duo /synth_summarize_datasettext100K<n<1M5 likes233 downloads2y agoHugging Face12vwxyzjn /summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144 TL;DR SFT Dataset for OpenAI's Summarize from Feedback task The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset These columns are taken directly from the aforementioned dataset: id: unique identifier for the post subreddit: subreddit the post was taken from title: title of the post post: body of the post summary: summary of the post reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.tabular100K<n<1M0 likes225 downloads3y agoHugging Face13chansung /synth_summarize_datasettext100K<n<1M2 likes180 downloads2y agoHugging Face14ichanchiu /Summarized_10K-MDA Summarized 10-K MD&A Dataset Description The Summarized 10-K MD&A dataset provides concise, machine-generated summaries of 10-K filings for publicly traded companies. These filings are sourced from the SEC EDGAR database, and the dataset is designed to facilitate financial text analysis, such as summarization, sentiment analysis, and financial disclosure studies. Key Features Language: English Dataset Size: 98,100 rows License: MIT License Source: SEC EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/ichanchiu/Summarized_10K-MDA.text10K<n<100K0 likes174 downloads2y agoHugging Face15chansung /new_summarize_synth_ds3text100K<n<1M0 likes148 downloads2y agoHugging Face16llama-duo /synth_summarize_dataset_deduptext100K<n<1M2 likes139 downloads2y agoHugging Face17chansung /new_summarize_synth_ds2text100K<n<1M0 likes126 downloads2y agoHugging Face18farzad0514 /summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1text100K<n<1M0 likes123 downloads1y agoHugging Face19npc-engine /light-batch-summarize-dialogue Light dataset Dialogues are preprocessed into a form: <Character name>: <character line> ... <Character name>: <character line> Summarize the document text10K<n<100K9 likes117 downloads4y agoHugging Face20vwxyzjn /summarize_from_feedback_oai_preprocessing_pythia-160m_48 Dataset Card for "summarize_from_feedback_oai_preprocessing_pythia-160m_48" More Information needed tabular100K<n<1M0 likes99 downloads3y agoHugging Face21yguooo /summarize_from_feedback_oai_preprocessing_pythia_scene0_1incontexttabular100K<n<1M0 likes96 downloads2y agoHugging Face22luckeciano /learning-to-summarizetext100K<n<1M1 likes95 downloads3y agoHugging Face23vwxyzjn /summarize_from_feedback_oai_preprocessing_1706381144 Dataset Card for "summarize_from_feedback_oai_preprocessing_1706381144" More Information needed tabular100K<n<1M2 likes89 downloads3y agoHugging Face24yguooo /summarize_from_feedback_oai_preprocessing_llama3_scene1tabular100K<n<1M0 likes84 downloads2y agoHugging Face25HuggingFaceH4 /summarize-from-feedback Dataset Card for "summarize-from-feedback" More Information needed text100K<n<1M4 likes83 downloads4y agoHugging Face26cleanrl /summarize_from_feedback_oai_preprocessing_1705009345 Dataset Card for "summarize_from_feedback_oai_preprocessing_1705009345" More Information needed tabular100K<n<1M0 likes77 downloads3y agoHugging Face27iamjoon /finance_news_summarizertextn<1K0 likes77 downloads1y agoHugging Face28vwxyzjn /summarize_from_feedback_oai_preprocessing_gpt2_153 Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_153" More Information needed tabular100K<n<1M0 likes76 downloads3y agoHugging Face29yguooo /summarize_from_feedback_oai_preprocessing_llama3tabular100K<n<1M0 likes76 downloads2y agoHugging Face30AmareshHebbar /clinical-summarizer-sft Clinical Note Summarizer (SOAP Format) Part of the AxisMapper Medical AI Suite — 16 domain-specific SFT datasets for fine-tuning medical LLMs. Built by AmareshHebbar | Studio Ilios / Humanova Minds What this dataset does Long clinical notes → structured SOAP summaries Why download this Automate clinical documentation. Reduce physician burnout by summarizing visit notes into Subjective / Objective / Assessment / Plan format. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AmareshHebbar/clinical-summarizer-sft.texttext-generation10K<n<100K0 likes74 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.