datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AI_HUB_DATASET_after_preprocessingWhisper_FineTuning_Su_preprocessingsummarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.ode-preprocessing-hy15-testdata-preprocessing-automl-benchmarks
Data Preprocessing AutoML Benchmarks
This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML.
Usage
Load a specific dataset configuration like this:
from datasets import load_dataset
# Example for loading the TREC dataset
dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec")
Available Datasets
Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.brats_preprocessingsummarize_from_feedback_oai_preprocessing_pythia-160m_48
Dataset Card for "summarize_from_feedback_oai_preprocessing_pythia-160m_48"
More Information needed
summarize_from_feedback_oai_preprocessing_pythia_scene0_1incontextsummarize_from_feedback_oai_preprocessing_1706381144
Dataset Card for "summarize_from_feedback_oai_preprocessing_1706381144"
More Information needed
RoboSteer-Preprocessing
RoboSteer Preprocessing
Reusable intermediate preprocessing assets for RoboSteer. This initial release contains Level 1 processed text instructions and image-conditioned static videos. Other levels and preprocessing stages can be added under their own directories.
Release v1.0.0
213,044 text records in nine lossless UTF-8 Parquet tables.
28,666 original MP4 files: 14,333 IMG_TXT_HUMAN and 14,333 IMG_TXT_SKEL.
Independent, uncompressed tar shards contain MP4 files… See the full description on the dataset page: https://huggingface.co/datasets/PhoebeCC/RoboSteer-Preprocessing.summarize_from_feedback_oai_preprocessing_1705009345
Dataset Card for "summarize_from_feedback_oai_preprocessing_1705009345"
More Information needed
summarize_from_feedback_oai_preprocessing_llama3_scene1summarize_from_feedback_oai_preprocessing_gpt2_153
Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_153"
More Information needed
summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_48
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_48.summarize_from_feedback_oai_preprocessing_llama3_scene2summarize_from_feedback_oai_preprocessing_pythia_scene4summarize_from_feedback_oai_preprocessing_1711138793
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138793"
More Information needed
summarize_from_feedback_oai_preprocessing_llama3_scene0summarize_from_feedback_oai_preprocessing_llama3_scene3summarize_from_feedback_oai_preprocessing_pythia_scene0summarize_from_feedback_oai_preprocessing
Dataset Card for "summarize_from_feedback_oai_preprocessing"
More Information needed
summarize_from_feedback_oai_preprocessing_1711138084
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138084"
More Information needed
summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_53
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_53.summarize_from_feedback_oai_preprocessing_pythia_scene2summarize_from_feedback_oai_preprocessing_1711138537
Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138537"
More Information needed
summarize_from_feedback_oai_preprocessing_pythia_scene0_thesummarize_from_feedback_tldr_3_filtered_oai_preprocessing_gpt2_48
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_gpt2_48.summarize_from_feedback_oai_preprocessing_pythia-160m_169
Dataset Card for "summarize_from_feedback_oai_preprocessing_pythia-160m_169"
More Information needed
summarize_from_feedback_oai_preprocessing_llama3_scene4summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia_scene0_1incontext
TL;DR SFT Dataset for OpenAI's Summarize from Feedback task
The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset
These columns are taken directly from the aforementioned dataset:
id: unique identifier for the post
subreddit: subreddit the post was taken from
title: title of the post
post: body of the post
summary: summary of the post
reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia_scene0_1incontext.
