CoolFace
20 results

preprocessing

TARARARAK /AI_HUB_DATASET_after_preprocessing10K<n<100K0 likes275 downloads2y agoHugging FaceYongJaeLee /Whisper_FineTuning_Su_preprocessing10K<n<100K0 likes222 downloads1y agoHugging Facevwxyzjn /summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144 TL;DR SFT Dataset for OpenAI's Summarize from Feedback task The dataset is directly taken from https://github.com/openai/summarize-from-feedback/tree/700967448d10004279f138666442bf1497d0e705#reddit-tldr-dataset These columns are taken directly from the aforementioned dataset: id: unique identifier for the post subreddit: subreddit the post was taken from title: title of the post post: body of the post summary: summary of the post reference_response: reference response for the post… See the full description on the dataset page: https://huggingface.co/datasets/vwxyzjn/summarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144.tabular100K<n<1M0 likes216 downloads3y agoHugging Facewlsaidhi /ode-preprocessing-hy15-testtextn<1K0 likes123 downloads9mo agoHugging FaceMothMalone /data-preprocessing-automl-benchmarks Data Preprocessing AutoML Benchmarks This repository contains text classification datasets with known data quality issues for preprocessing research in AutoML. Usage Load a specific dataset configuration like this: from datasets import load_dataset # Example for loading the TREC dataset dataset = load_dataset("MothMalone/data-preprocessing-automl-benchmarks", "trec") Available Datasets Below are the details for each dataset configuration available in this… See the full description on the dataset page: https://huggingface.co/datasets/MothMalone/data-preprocessing-automl-benchmarks.texttext-classification100K<n<1M0 likes112 downloads1y agoHugging Faceyuuricho /brats_preprocessing0 likes106 downloads6mo agoHugging Face