CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vwxyzjn /summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144 TL;DR SFT Dataset for OpenAI's Summarize from Feedback task The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset These columns are taken directly from the aforementioned dataset: id: unique identifier for the post subreddit: subreddit the post was taken from title: title of the post post: body of the post summary: summary of the post reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.tabular100K<n<1M0 likes216 downloads3y agoHugging Face02wlsaidhi /ode-preprocessing-hy15-testtextn<1K0 likes123 downloads9mo agoHugging Face03MothMalone /data-preprocessing-automl-benchmarks Data Preprocessing AutoML Benchmarks This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML. Usage Load a specific dataset configuration like this: from datasets import load_dataset # Example for loading the TREC dataset dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec") Available Datasets Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.texttext-classification100K<n<1M0 likes112 downloads1y agoHugging Face04vwxyzjn /summarize_from_feedback_oai_preprocessing_pythia-160m_48 Dataset Card for "summarize_from_feedback_oai_preprocessing_pythia-160m_48" More Information needed tabular100K<n<1M0 likes99 downloads3y agoHugging Face05yguooo /summarize_from_feedback_oai_preprocessing_pythia_scene0_1incontexttabular100K<n<1M0 likes97 downloads2y agoHugging Face06vwxyzjn /summarize_from_feedback_oai_preprocessing_1706381144 Dataset Card for "summarize_from_feedback_oai_preprocessing_1706381144" More Information needed tabular100K<n<1M2 likes93 downloads3y agoHugging Face07PhoebeCC /RoboSteer-Preprocessing RoboSteer Preprocessing Reusable intermediate preprocessing assets for RoboSteer. This initial release contains Level 1 processed text instructions and image-conditioned static videos. Other levels and preprocessing stages can be added under their own directories. Release v1.0.0 213,044 text records in nine lossless UTF-8 Parquet tables. 28,666 original MP4 files: 14,333 IMG_TXT_HUMAN and 14,333 IMG_TXT_SKEL. Independent, uncompressed tar shards contain MP4 files… See the full description on the dataset page: https://huggingface.co/datasets/PhoebeCC/RoboSteer-Preprocessing.tabular100K<n<1M0 likes89 downloads5d agoHugging Face08cleanrl /summarize_from_feedback_oai_preprocessing_1705009345 Dataset Card for "summarize_from_feedback_oai_preprocessing_1705009345" More Information needed tabular100K<n<1M0 likes83 downloads3y agoHugging Face09yguooo /summarize_from_feedback_oai_preprocessing_llama3_scene1tabular100K<n<1M0 likes83 downloads2y agoHugging Face10vwxyzjn /summarize_from_feedback_oai_preprocessing_gpt2_153 Dataset Card for "summarize_from_feedback_oai_preprocessing_gpt2_153" More Information needed tabular100K<n<1M0 likes76 downloads3y agoHugging Face11vwxyzjn /summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_48 TL;DR SFT Dataset for OpenAI's Summarize from Feedback task The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset These columns are taken directly from the aforementioned dataset: id: unique identifier for the post subreddit: subreddit the post was taken from title: title of the post post: body of the post summary: summary of the post reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_48.tabular100K<n<1M0 likes73 downloads3y agoHugging Face12yguooo /summarize_from_feedback_oai_preprocessing_llama3_scene2tabular100K<n<1M0 likes73 downloads2y agoHugging Face13yguooo /summarize_from_feedback_oai_preprocessing_pythia_scene4tabular100K<n<1M0 likes71 downloads2y agoHugging Face14vwxyzjn /summarize_from_feedback_oai_preprocessing_1711138793 Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138793" More Information needed image100K<n<1M0 likes69 downloads3y agoHugging Face15yguooo /summarize_from_feedback_oai_preprocessing_llama3_scene0tabular100K<n<1M0 likes69 downloads2y agoHugging Face16yguooo /summarize_from_feedback_oai_preprocessing_llama3_scene3tabular100K<n<1M0 likes68 downloads2y agoHugging Face17yguooo /summarize_from_feedback_oai_preprocessing_pythia_scene0tabular100K<n<1M0 likes68 downloads2y agoHugging Face18vwxyzjn /summarize_from_feedback_oai_preprocessing Dataset Card for "summarize_from_feedback_oai_preprocessing" More Information needed text100K<n<1M0 likes66 downloads3y agoHugging Face19vwxyzjn /summarize_from_feedback_oai_preprocessing_1711138084 Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138084" More Information needed image100K<n<1M0 likes65 downloads3y agoHugging Face20vwxyzjn /summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_53 TL;DR SFT Dataset for OpenAI's Summarize from Feedback task The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset These columns are taken directly from the aforementioned dataset: id: unique identifier for the post subreddit: subreddit the post was taken from title: title of the post post: body of the post summary: summary of the post reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia-160m_53.tabular100K<n<1M0 likes64 downloads3y agoHugging Face21yguooo /summarize_from_feedback_oai_preprocessing_pythia_scene2tabular100K<n<1M0 likes63 downloads2y agoHugging Face22vwxyzjn /summarize_from_feedback_oai_preprocessing_1711138537 Dataset Card for "summarize_from_feedback_oai_preprocessing_1711138537" More Information needed tabular100K<n<1M0 likes61 downloads3y agoHugging Face23yguooo /summarize_from_feedback_oai_preprocessing_pythia_scene0_thetabular100K<n<1M0 likes61 downloads2y agoHugging Face24vwxyzjn /summarize_from_feedback_tldr_3_filtered_oai_preprocessing_gpt2_48 TL;DR SFT Dataset for OpenAI's Summarize from Feedback task The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset These columns are taken directly from the aforementioned dataset: id: unique identifier for the post subreddit: subreddit the post was taken from title: title of the post post: body of the post summary: summary of the post reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_gpt2_48.tabular100K<n<1M0 likes60 downloads3y agoHugging Face25vwxyzjn /summarize_from_feedback_oai_preprocessing_pythia-160m_169 Dataset Card for "summarize_from_feedback_oai_preprocessing_pythia-160m_169" More Information needed tabular100K<n<1M0 likes57 downloads3y agoHugging Face26yguooo /summarize_from_feedback_oai_preprocessing_llama3_scene4tabular100K<n<1M0 likes57 downloads2y agoHugging Face27yguooo /summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia_scene0_1incontext TL;DR SFT Dataset for OpenAI's Summarize from Feedback task The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset These columns are taken directly from the aforementioned dataset: id: unique identifier for the post subreddit: subreddit the post was taken from title: title of the post post: body of the post summary: summary of the post reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_pythia_scene0_1incontext.tabular100K<n<1M0 likes55 downloads2y agoHugging Face28yguooo /summarize_from_feedback_oai_preprocessing_llama3tabular100K<n<1M0 likes53 downloads2y agoHugging Face29cleanrl /summarize_from_feedback_oai_preprocessing_1704427060 Dataset Card for "summarize_from_feedback_oai_preprocessing_1704427060" More Information needed tabular100K<n<1M0 likes52 downloads3y agoHugging Face30vwxyzjn /summarize_from_feedback_oai_preprocessing_1708444324 Dataset Card for "summarize_from_feedback_oai_preprocessing_1708444324" More Information needed tabular100K<n<1M0 likes52 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.