datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vsi-bench-debiased-evalgen_debiased_nli
Overview
Original dataset available here.
@inproceedings{gen-debiased-nli-2022,
title = "Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets",
author = "Wu, Yuxiang and
Gardner, Matt and
Stenetorp, Pontus and
Dasigi, Pradeep",
booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics",
month = may,
year = "2022",
publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/gen_debiased_nli.epidemic_sound_effects_t5_debiasedlaion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs).
verifier-debias-v2
verifier-debias-v2 — de-biased SFT data for a generative verifier
This dataset was presented in the paper One Token to Fool LLM-as-a-Judge.
GitHub repository: yulaizhao/Master-RM
SFT data to train a generative verifier (GenRM, arXiv:2408.15240)
from Qwen/Qwen2.5-7B-Instruct. Each row is a conversational example
(messages = system + user + gold assistant) plus a verdict (PASS/FAIL) and
the underlying 1–5 score. The assistant target is a critique ending in
Verdict: PASS / Verdict:… See the full description on the dataset page: https://huggingface.co/datasets/narcolepticchicken/verifier-debias-v2.laion_text_debiased_100M
100M Text Debiased Subset from LAION 2B
Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images.
Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data.
CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.Winogrande_debiasedwinogrande_debiased-telugu-romanized-nodictwinogrande_debiased-teluguUnsplash-Debiased_by_AI-18K
Unsplash-Debiased-by-AI-18K
17,812 high-quality image captions selected by a label-free, dual-modality, 6-judge
debiasing pipeline from 32,135 Unsplash-40K captions. Each caption passes three gates:
(a) high/mid tier in a debiased judge consensus (family-balanced × reliability-weighted,
rank-calibrated, self-preference-dropped), (b) image quality (NIQE/MUSIQ/LIQE), and
(c) image–caption alignment (CLIPScore/SigLIP).
Images are NOT included. This dataset ships captions +… See the full description on the dataset page: https://huggingface.co/datasets/henrywch2huggingface/Unsplash-Debiased_by_AI-18K.winogrande_debiased-telugu_filteredwinogrande_debiased-telugu-romanizedBias-Debias-Alpaca
Responsible Media Content Matrix (RMCM): Overview
The RMCM is a strategic tool developed to address various forms of bias and unethical practices in media reporting. It encompasses several key categories, each focusing on a specific type of bias or ethical concern. The primary objective of the RMCM is to foster responsible journalism and content creation by providing clear guidelines on identifying and rectifying biased or harmful content.
Key Categories of RMCM:… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/Bias-Debias-Alpaca.gender_debias_disambiguate_oldgender_debias_disambiguatedebiased_dataset
Dataset Description
About the Dataset:
This dataset contains text data that has been processed to identify biased statements based on dimensions and aspects. Each entry has been processed using the GPT-4 language model and manually verified by 5 human annotators for quality assurance.
Purpose:
The dataset aims to help train and evaluate machine learning models in detecting, classifying, and correcting biases in text content, making it essential for NLP research related to fairness… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/debiased_dataset.audioset_t5_debiasedLLaMa_3B_Debiasing_Instruction_CoTtree-debiasing-stage3LLaMa_1B_Debiasing_Instruction_CoTLLaMa_3B_Debiasing_Instruction_NoCoTLLaMa_1B_Debiasing_Instruction_CoT_inferencedgender_debiasLLaMa_1B_IsCoT_DebiasingInstructionLLaMa_1B_Debiasing_Instruction_NoCoT_inferencedRM-R1-Distill-SFT-DebiasedLLaMa_1B_NoCoT_DebiasingInstructiondebiased-ultrachat-200kLLaMa_1B_Debiasing_Instruction_NoCoTRM-R1-after-Distill-RLVR-Debiased
