CoolFace
20 results

Clickbait

community-datasets /clickbait_news_bg Dataset Card for Clickbait/Fake News in Bulgarian Dataset Summary This is a corpus of Bulgarian news over a fixed period of time, whose factuality had been questioned. The news come from 377 different sources from various domains, including politics, interesting facts and tips&tricks. The dataset was prepared for the Hack the Fake News hackathon. It was provided by the Bulgarian Association of PR Agencies and is available in Gitlab. The corpus was automatically… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/clickbait_news_bg.texttext-classification1K<n<10K2 likes177 downloads3y agoHugging Facecommunity-datasets /id_clickbaitThe CLICK-ID dataset is a collection of Indonesian news headlines that was collected from 12 local online news publishers; detikNews, Fimela, Kapanlagi, Kompas, Liputan6, Okezone, Posmetro-Medan, Republika, Sindonews, Tempo, Tribunnews, and Wowkeren. This dataset is comprised of mainly two parts; (i) 46,119 raw article data, and (ii) 15,000 clickbait annotated sample headlines. Annotation was conducted with 3 annotator examining each headline. Judgment were based only on the headline. The majority then is considered as the ground truth. In the annotated sample, our annotation shows 6,290 clickbait and 8,710 non-clickbait.text-classification10K<n<100K0 likes168 downloads3y agoHugging Facerodrigoaraujorosa /detector-clickbait-br-datasets Detector Clickbait BR - Datasets Este repositório contém os datasets utilizados para o treinamento do modelo detector-clickbait-br-model, um classificador de textos em português brasileiro capaz de identificar títulos clickbait. 📚 Descrição dos Datasets 1. detector-clickbait-br-raw.csv Dataset original contendo os dados iniciais sem processamento. Características: Dados brutos coletados originalmente Pode conter duplicatas Pode conter valores nulos Formato:… See the full description on the dataset page: https://huggingface.co/datasets/rodrigoaraujorosa/detector-clickbait-br-datasets.tabulartext-classification10K<n<100K2 likes114 downloads10mo agoHugging Faceml-projects /clickbait-ml_datasettextn<1K0 likes109 downloads3y agoHugging Facemanandey /id_clickbait This is the annotated full version of the dataset. Dataset Summary The CLICK-ID dataset is a collection of Indonesian news headlines that was collected from 12 local online news publishers; detikNews, Fimela, Kapanlagi, Kompas, Liputan6, Okezone, Posmetro-Medan, Republika, Sindonews, Tempo, Tribunnews, and Wowkeren. This dataset is comprised of mainly two parts; (i) 46,119 raw article data, and (ii) 15,000 clickbait annotated sample headlines. Annotation was conducted… See the full description on the dataset page: https://huggingface.co/datasets/manandey/id_clickbait.texttext-classification10K<n<100K1 likes109 downloads2y agoHugging Facechristinacdl /clickbait_detection_dataset 37.870 texts in total, 17.850 NOT clickbait texts and 20.020 CLICKBAIT texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 14.280, 1 ==> 16.016 Validation set label distribution: 0 ==> 1.785, 1 ==> 2.002 Test set label distribution: 0 ==> 1.785, 1 ==> 2.002 The dataset was created from the… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/clickbait_detection_dataset.texttext-classification10K<n<100K4 likes103 downloads3y agoHugging Face