CoolFace
20 results

debias

hunarbatra /vsi-bench-debiased-evalimage1K<n<10K0 likes519 downloads7mo agoHugging Facepietrolesci /gen_debiased_nli Overview Original dataset available here. @inproceedings{gen-debiased-nli-2022, title = "Generating Data to Mitigate Spurious Correlations in Natural Language Inference Datasets", author = "Wu, Yuxiang and Gardner, Matt and Stenetorp, Pontus and Dasigi, Pradeep", booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics", month = may, year = "2022", publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/pietrolesci/gen_debiased_nli.text1M<n<10M0 likes309 downloads4y agoHugging FaceCLAPv2 /epidemic_sound_effects_t5_debiasedaudio10K<n<100K0 likes273 downloads2y agoHugging Facezxbsmk /laion_text_debiased_60MFilter zxbsmk/laion_text_debiased_60M by image size and get 512 subset(12,009,641 pairs), 768 subset(4,915,850 pairs), 1024 subset(1,985,026 pairs). image10M<n<100M1 likes219 downloads3y agoHugging Facenarcolepticchicken /verifier-debias-v2 verifier-debias-v2 — de-biased SFT data for a generative verifier This dataset was presented in the paper One Token to Fool LLM-as-a-Judge. GitHub repository: yulaizhao/Master-RM SFT data to train a generative verifier (GenRM, arXiv:2408.15240) from Qwen/Qwen2.5-7B-Instruct. Each row is a conversational example (messages = system + user + gold assistant) plus a verdict (PASS/FAIL) and the underlying 1–5 score. The assistant target is a critique ending in Verdict: PASS / Verdict:… See the full description on the dataset page: https://huggingface.co/datasets/narcolepticchicken/verifier-debias-v2.texttext-generation10K<n<100K0 likes160 downloads4mo agoHugging Facelinyq /laion_text_debiased_100M 100M Text Debiased Subset from LAION 2B Captions in LAION-2B have a significant bias towards describing visual text content embedded in the images. Released CLIP models have strong text spotting bias in almost every style of web images, resulting in the CLIP-filtering datasets inherently biased towards visual text dominant data. CLIP models easily learn text spotting capacity from parrot captions while failing to connect the vision-language semantics, just like a text spotting… See the full description on the dataset page: https://huggingface.co/datasets/linyq/laion_text_debiased_100M.image100M<n<1B0 likes112 downloads2y agoHugging Face