datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
id_clickbait
This is the annotated full version of the dataset.
Dataset Summary
The CLICK-ID dataset is a collection of Indonesian news headlines that was collected from 12 local online news
publishers; detikNews, Fimela, Kapanlagi, Kompas, Liputan6, Okezone, Posmetro-Medan, Republika, Sindonews, Tempo,
Tribunnews, and Wowkeren. This dataset is comprised of mainly two parts; (i) 46,119 raw article data, and (ii)
15,000 clickbait annotated sample headlines. Annotation was conducted… See the full description on the dataset page: https://huggingface.co/datasets/manandey/id_clickbait.clickbait_detection_dataset
37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 14.280, 1 ==> 16.016
Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002
Test set label distribution: 0 ==> 1.785, 1 ==> 2.002
The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/clickbait_detection_dataset.clickbait_notclickbait_dataset0 : not clickbait
1 : clickbait
Dataset cleaned from duplicates and kept only the first appearing text.
Dataset split into train and test sets using 0.2 split ratio.
Dataset split into test and validation sets using 0.2 split ratio.
Size of training set: 43.802
Size of test set: 8.760
Size of validation set: 2.191
clickbait_detection_dataset
37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 14.280, 1 ==> 16.016
Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002
Test set label distribution: 0 ==> 1.785, 1 ==> 2.002
The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/mayoooookha/clickbait_detection_dataset.clickbaitMultilingual_Clickbait_DatasetClickbait_Newclickbait-spoilingData for Semeval 2023 task, clickbait spoiling
clickbait-spoiling-data-question
Webis Clickbait Spoiling Corpus
The Webis Clickbait Spoiling Corpus 2022 (Webis-Clickbait-22) contains 5,000 spoiled clickbait posts crawled from Facebook, Reddit, and Twitter.
This corpus supports the task of clickbait spoiling, which deals with generating a short text that satisfies the curiosity induced by a clickbait post.
This dataset contains the clickbait posts and manually cleaned versions of the linked documents, and extracted spoilers for each clickbait post.
Additionally… See the full description on the dataset page: https://huggingface.co/datasets/pramitsahoo/clickbait-spoiling-data-question.indonesian-clickbait-spoilingclickbait_spoilingwebis-clickbait-spoiling-seq-tagclickbaitTestclickbait_notclickbait_dataset0 : not clickbait
1 : clickbait
Dataset cleaned from duplicates and kept only the first appearing text.
Dataset split into train and test sets using 0.2 split ratio.
Dataset split into test and validation sets using 0.2 split ratio.
Size of training set: 43.802
Size of test set: 8.760
Size of validation set: 2.191
