datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openai_summarize_tldr
Dataset Card for "openai_summarize_tldr"
More Information needed
KMMLU-Summarized-Chain_of_Thought
Dataset Card for Condensed Chain-of-Thought KMMLU Dataset
This dataset card provides detailed information about the condensed KMMLU dataset. The dataset has been summarized using Upstage's LLM: Solar-Pro to condense the original KMMLU training and development data while preserving its quality and usability. Additionally, a new column, 'chain_of_thought', has been introduced to align with the reasoning approach outlined in the paper "Chain-of-Thought Prompting Elicits Reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/SabaPivot/KMMLU-Summarized-Chain_of_Thought.openai_summarize_comparisonssummarize_from_feedbackSummarize from Feedback contains the human feedback data released by the "Learning to summarize from human feedback" paper.wavepulse-radio-summarized-transcripts
WavePulse Radio Summarized Transcripts
Dataset Summary
WavePulse Radio Summarized Transcripts is a large-scale dataset containing summarized transcripts from 396 radio stations across the United States, collected between June 26, 2024, and October 3, 2024. The dataset comprises approximately 1.5 million summaries derived from 485,090 hours of radio broadcasts, primarily covering news, talk shows, and political discussions.
The raw version of the transcripts is available… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/wavepulse-radio-summarized-transcripts.summarize_from_feedback_smalllatin-summarizer-dataset
✨ LatinSummarizer Dataset ✨
Note: If Dataset Viewer is not available, see samples of dataset for samples from the dataset.The LatinSummarizer Dataset is a comprehensive collection of Latin texts designed to support natural language processing research for a low-resource language. It provides parallel data for various tasks, including translation (Latin-to-English) and summarization (extractive and abstractive).
This dataset was created for a… See the full description on the dataset page: https://huggingface.co/datasets/LatinNLP/latin-summarizer-dataset.summarize_from_feedback_tldr_3_filteredThis is the query dataset taken directly from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
gemma2b-summarize-eval-by-claude3sonnetgemma2b-summarize-eval-by-gemini15flashgemma7b-summarize-eval-by-gemini15flashlinux-man-pages-tldr-summarized
Dataset Card for linux-man-pages-tldr-summarized
Dataset Summary
This dataset contains linux man pages downloaded from man7, with a prefix: 'summarize: ', and the corresponding summarization downloaded from TLDR-pages.
Supported Tasks
This dataset should be used to fine-tune language models for summarization tasks.
gemma7b-summarize-eval-by-claude3sonnetsynth_summarize_datasetsummarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.synth_summarize_datasetSummarized_10K-MDA
Summarized 10-K MD&A
Dataset Description
The Summarized 10-K MD&A dataset provides concise, machine-generated summaries of 10-K filings for publicly traded companies. These filings are sourced from the SEC EDGAR database, and the dataset is designed to facilitate financial text analysis, such as summarization, sentiment analysis, and financial disclosure studies.
Key Features
Language: English
Dataset Size: 98,100 rows
License: MIT License
Source: SEC EDGAR… See the full description on the dataset page: https://huggingface.co/datasets/ichanchiu/Summarized_10K-MDA.new_summarize_synth_ds3summarized-hyperpartisan-news-by-facebook-bart-large-cnn-v1light-batch-summarize-dialogue
Light dataset
Dialogues are preprocessed into a form:
<Character name>: <character line>
...
<Character name>: <character line>
Summarize the document
new_summarize_synth_ds2synth_summarize_dataset_dedupsummarize_from_feedback_oai_preprocessing_pythia-160m_48
Dataset Card for "summarize_from_feedback_oai_preprocessing_pythia-160m_48"
More Information needed
summarize_from_feedback_oai_preprocessing_pythia_scene0_1incontextlearning-to-summarizesummarize_from_feedback_oai_preprocessing_1706381144
Dataset Card for "summarize_from_feedback_oai_preprocessing_1706381144"
More Information needed
gemma2b-summarize-locallm-responsesummarize-from-feedback
Dataset Card for "summarize-from-feedback"
More Information needed
summarize_from_feedback_oai_preprocessing_1705009345
Dataset Card for "summarize_from_feedback_oai_preprocessing_1705009345"
More Information needed
summarize_from_feedback_oai_preprocessing_llama3_scene1
