CoolFace
Datasetpublic

franciellevargas/FactNews

Evaluation Benchmark for Sentence-Level Factuality Prediciton in Portuguese The FactNews consits of the first large sentence-level annotated corpus for factuality prediciton in Portuguese. It is composed of 6,191 sentences annotated according to factuality and media bias definitions proposed by AllSides. We use FactNews to assess the overall reliability of news sources by formulating two text classification problems for predicting sentence-level factuality of news reporting… See the full description on the dataset page: https://huggingface.co/datasets/franciellevargas/FactNews.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
0likes25downloads
Dataset Card

Evaluation Benchmark for Sentence-Level Factuality Prediciton in Portuguese

The FactNews consits of the first large sentence-level annotated corpus for factuality prediciton in Portuguese. It is composed of 6,191 sentences annotated according to factuality and media bias definitions proposed by AllSides. We use FactNews to assess the overall reliability of news sources by formulating two text classification problems for predicting sentence-level factuality of news reporting and bias of media outlets. Our experiments demonstrate that biased sentences present a higher number of words compared to factual sentences, besides having a predominance of emotions. Hence, the fine-grained analysis of subjectivity and impartiality of news articles showed promising results for predicting the reliability of entire media outlets.

Dataset Description

<b>Proposed by </b>: Francielle Vargas (<https://franciellevargas.github.io/>) <br> <b>Funded by </b>: Google <br> <b>Language(s) (NLP)</b>: Portuguese <br>

Dataset Sources

<b>Repository</b>: https://github.com/franciellevargas/FactNews <br> <b>Demo</b>: FACTual Fact-Checking: http://143.107.183.175:14582/<br>

Paper

<b>Predicting Sentence-Level Factuality of News and Bias of Media Outlets</b> <br> Francielle Vargas, Kokil Jaidka, Thiago A.S. Pardo, Fabrício Benevenuto<br> <i>Recent Advances in Natural Language Processing (RANLP 2023) <i> <br> Varna, Bulgaria. https://aclanthology.org/2023.ranlp-1.127/ <br>

Dataset Contact

francielealvargas@gmail.com