CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01manandey /id_clickbait This is the annotated full version of the dataset. Dataset Summary The CLICK-ID dataset is a collection of Indonesian news headlines that was collected from 12 local online news publishers; detikNews, Fimela, Kapanlagi, Kompas, Liputan6, Okezone, Posmetro-Medan, Republika, Sindonews, Tempo, Tribunnews, and Wowkeren. This dataset is comprised of mainly two parts; (i) 46,119 raw article data, and (ii) 15,000 clickbait annotated sample headlines. Annotation was conducted… See the full description on the dataset page: https://huggingface.co/datasets/manandey/id_clickbait.texttext-classification10K<n<100K1 likes109 downloads2y agoHugging Face02christinacdl /clickbait_detection_dataset 37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 14.280, 1 ==> 16.016 Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002 Test set label distribution: 0 ==> 1.785, 1 ==> 2.002 The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/clickbait_detection_dataset.texttext-classification10K<n<100K4 likes103 downloads3y agoHugging Face03christinacdl /clickbait_notclickbait_dataset0 : not clickbait 1 : clickbait Dataset cleaned from duplicates and kept only the first appearing text. Dataset split into train and test sets using 0.2 split ratio. Dataset split into test and validation sets using 0.2 split ratio. Size of training set: 43.802 Size of test set: 8.760 Size of validation set: 2.191 texttext-classification10K<n<100K0 likes60 downloads3y agoHugging Face04mayoooookha /clickbait_detection_dataset 37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 14.280, 1 ==> 16.016 Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002 Test set label distribution: 0 ==> 1.785, 1 ==> 2.002 The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/mayoooookha/clickbait_detection_dataset.texttext-classification10K<n<100K0 likes50 downloads4d agoHugging Face05SotirisLegkas /clickbaittext10K<n<100K0 likes36 downloads3y agoHugging Face06christinacdl /Multilingual_Clickbait_Datasettext100K<n<1M1 likes34 downloads3y agoHugging Face07christinacdl /Clickbait_Newtexttext-classification10K<n<100K0 likes32 downloads3y agoHugging Face08Tugay /clickbait-spoilingData for Semeval 2023 task, clickbait spoiling text1K<n<10K0 likes31 downloads4y agoHugging Face09pramitsahoo /clickbait-spoiling-data-question Webis Clickbait Spoiling Corpus The Webis Clickbait Spoiling Corpus 2022 (Webis-Clickbait-22) contains 5,000 spoiled clickbait posts crawled from Facebook, Reddit, and Twitter. This corpus supports the task of clickbait spoiling, which deals with generating a short text that satisfies the curiosity induced by a clickbait post. This dataset contains the clickbait posts and manually cleaned versions of the linked documents, and extracted spoilers for each clickbait post. Additionally… See the full description on the dataset page: https://huggingface.co/datasets/pramitsahoo/clickbait-spoiling-data-question.text1K<n<10K0 likes28 downloads2y agoHugging Face10intanm /indonesian-clickbait-spoilingtextn<1K0 likes15 downloads3y agoHugging Face11MateuszW /clickbait_spoilingtext1K<n<10K0 likes15 downloads3y agoHugging Face12intanm /webis-clickbait-spoiling-seq-tagtext1K<n<10K0 likes14 downloads3y agoHugging Face13melodywang2906 /clickbaitTesttextn<1K0 likes3 downloads2y agoHugging Face14sahoo1803 /clickbait_notclickbait_dataset0 : not clickbait 1 : clickbait Dataset cleaned from duplicates and kept only the first appearing text. Dataset split into train and test sets using 0.2 split ratio. Dataset split into test and validation sets using 0.2 split ratio. Size of training set: 43.802 Size of test set: 8.760 Size of validation set: 2.191 texttext-classification10K<n<100K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.