datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs).
Bias-Debias-Alpaca
Responsible Media Content Matrix (RMCM): Overview
The RMCM is a strategic tool developed to address various forms of bias and unethical practices in media reporting. It encompasses several key categories, each focusing on a specific type of bias or ethical concern. The primary objective of the RMCM is to foster responsible journalism and content creation by providing clear guidelines on identifying and rectifying biased or harmful content.
Key Categories of RMCM:… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/Bias-Debias-Alpaca.debiased_dataset
Dataset Description
About the Dataset:
This dataset contains text data that has been processed to identify biased statements based on dimensions and aspects. Each entry has been processed using the GPT-4 language model and manually verified by 5 human annotators for quality assurance.
Purpose:
The dataset aims to help train and evaluate machine learning models in detecting, classifying, and correcting biases in text content, making it essential for NLP research related to fairness… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/debiased_dataset.debiased_test_setBias-DeBiasedidentifying-debiasing-online-media-with-chatgptdebias_benchmark
