CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hunarbatra /vsi-bench-debiased-evalimage1K<n<10K0 likes519 downloads7mo agoHugging Face02pietrolesci /gen_debiased_nli Overview Original dataset available here. @inproceedings{gen-debiased-nli-2022, title = "Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets", author = "Wu, Yuxiang and Gardner, Matt and Stenetorp, Pontus and Dasigi, Pradeep", booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics", month = may, year = "2022", publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/gen_debiased_nli.text1M<n<10M0 likes309 downloads4y agoHugging Face03CLAPv2 /epidemic_sound_effects_t5_debiasedaudio10K<n<100K0 likes273 downloads2y agoHugging Face04zxbsmk /laion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs). image10M<n<100M1 likes219 downloads3y agoHugging Face05narcolepticchicken /verifier-debias-v2 verifier-debias-v2 — de-biased SFT data for a generative verifier This dataset was presented in the paper One Token to Fool LLM-as-a-Judge. GitHub repository: yulaizhao/Master-RM SFT data to train a generative verifier (GenRM, arXiv:2408.15240) from Qwen/Qwen2.5-7B-Instruct. Each row is a conversational example (messages = system + user + gold assistant) plus a verdict (PASS/FAIL) and the underlying 1–5 score. The assistant target is a critique ending in Verdict: PASS / Verdict:… See the full description on the dataset page: https://huggingface.co/datasets/narcolepticchicken/verifier-debias-v2.texttext-generation10K<n<100K0 likes160 downloads4mo agoHugging Face06linyq /laion_text_debiased_100M 100M Text Debiased Subset from LAION 2B Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images. Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data. CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.image100M<n<1B0 likes112 downloads2y agoHugging Face07jaypyon /Winogrande_debiasedtext10K<n<100K0 likes34 downloads2y agoHugging Face08indiehackers /winogrande_debiased-telugu-romanized-nodicttext10K<n<100K0 likes26 downloads2y agoHugging Face09indiehackers /winogrande_debiased-telugutext10K<n<100K0 likes22 downloads2y agoHugging Face10henrywch2huggingface /Unsplash-Debiased_by_AI-18K Unsplash-Debiased-by-AI-18K 17,812 high-quality image captions selected by a label-free, dual-modality, 6-judge debiasing pipeline from 32,135 Unsplash-40K captions. Each caption passes three gates: (a) high/mid tier in a debiased judge consensus (family-balanced × reliability-weighted, rank-calibrated, self-preference-dropped), (b) image quality (NIQE/MUSIQ/LIQE), and (c) image–caption alignment (CLIPScore/SigLIP). Images are NOT included. This dataset ships captions +… See the full description on the dataset page: https://huggingface.co/datasets/henrywch2huggingface/Unsplash-Debiased_by_AI-18K.tabularimage-to-text10K<n<100K0 likes21 downloads3mo agoHugging Face11indiehackers /winogrande_debiased-telugu_filteredtext10K<n<100K0 likes13 downloads2y agoHugging Face12indiehackers /winogrande_debiased-telugu-romanizedtext10K<n<100K0 likes12 downloads2y agoHugging Face13newsmediabias /Bias-Debias-Alpacagated Responsible Media Content Matrix (RMCM): Overview The RMCM is a strategic tool developed to address various forms of bias and unethical practices in media reporting. It encompasses several key categories, each focusing on a specific type of bias or ethical concern. The primary objective of the RMCM is to foster responsible journalism and content creation by providing clear guidelines on identifying and rectifying biased or harmful content. Key Categories of RMCM:… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/Bias-Debias-Alpaca.text10K<n<100K0 likes10 downloads3y agoHugging Face14samzirbo /gender_debias_disambiguate_oldtext10K<n<100K0 likes10 downloads2y agoHugging Face15samzirbo /gender_debias_disambiguatetext10K<n<100K0 likes9 downloads2y agoHugging Face16newsmediabias /debiased_datasetgated Dataset Description About the Dataset: This dataset contains text data that has been processed to identify biased statements based on dimensions and aspects. Each entry has been processed using the GPT-4 language model and manually verified by 5 human annotators for quality assurance. Purpose: The dataset aims to help train and evaluate machine learning models in detecting, classifying, and correcting biases in text content, making it essential for NLP research related to fairness… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/debiased_dataset.texttext-classification1K<n<10K0 likes8 downloads3y agoHugging Face17CLAPv2 /audioset_t5_debiasedaudio10K<n<100K0 likes8 downloads2y agoHugging Face18Nachiket-S /LLaMa_3B_Debiasing_Instruction_CoTtextn<1K0 likes7 downloads2y agoHugging Face19usamaahmedsh /tree-debiasing-stage3tabular100K<n<1M0 likes7 downloads6mo agoHugging Face20Nachiket-S /LLaMa_1B_Debiasing_Instruction_CoTtextn<1K0 likes6 downloads2y agoHugging Face21Nachiket-S /LLaMa_3B_Debiasing_Instruction_NoCoTtextn<1K0 likes6 downloads2y agoHugging Face22Nachiket-S /LLaMa_1B_Debiasing_Instruction_CoT_inferencedtextn<1K0 likes6 downloads2y agoHugging Face23samzirbo /gender_debiastext10K<n<100K0 likes5 downloads2y agoHugging Face24Nachiket-S /LLaMa_1B_IsCoT_DebiasingInstructiontextn<1K0 likes5 downloads2y agoHugging Face25Nachiket-S /LLaMa_1B_Debiasing_Instruction_NoCoT_inferencedtextn<1K0 likes5 downloads2y agoHugging Face26gloriafree /RM-R1-Distill-SFT-Debiasedtext10K<n<100K0 likes5 downloads7mo agoHugging Face27Nachiket-S /LLaMa_1B_NoCoT_DebiasingInstructiontextn<1K0 likes4 downloads2y agoHugging Face28Surtooner /debiased-ultrachat-200ktext10K<n<100K0 likes4 downloads10mo agoHugging Face29Nachiket-S /LLaMa_1B_Debiasing_Instruction_NoCoTtextn<1K0 likes3 downloads2y agoHugging Face30gloriafree /RM-R1-after-Distill-RLVR-Debiasedtext100K<n<1M0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.