CoolFace
15 results

perturbations

MLL-Lab /MultiBBQ-perturbations MultiBBQ: image perturbations Image-level perturbation sets used for the robustness experiments in Fairness Failure Modes of Multimodal LLMs. Each set is the GPT-Image-1 image collection from MLL-Lab/MultiBBQ with a single, controlled transform applied. Evaluating on a perturbed set measures how stable a model's fairness behavior is under everyday image degradations. Paper: Fairness Failure Modes of Multimodal LLMs Code:… See the full description on the dataset page: https://huggingface.co/datasets/MLL-Lab/MultiBBQ-perturbations.imagevisual-question-answering1K<n<10K0 likes751 downloads2mo agoHugging Facetasksource /boolq-natural-perturbationsBoolQ questions with semantic alteration and human verifications @article{khashabi2020naturalperturbations, title={Natural Perturbation for Robust Question Answering}, author={D. Khashabi and T. Khot and A. Sabhwaral}, journal={arXiv preprint}, year={2020} } tabulartext-classification10K<n<100K0 likes180 downloads3y agoHugging Faceabharadwaj123 /preference-model-perturbations preference-model-perturbations A Hugging Face dataset of paired model responses (original vs. counterfactually perturbed) along with human and reward-model preferences, generated by a Counterfactual Data Augmentation (CDA) pipeline to analyze and mitigate bias in preference models. Links Homepage: CDA Pipeline Code Description Each record contains: bias: type of bias expressed in the perturbation (5 possible values). query: the original user prompt or query.… See the full description on the dataset page: https://huggingface.co/datasets/abharadwaj123/preference-model-perturbations.texttext-generationn<1K2 likes35 downloads1y agoHugging FaceMawube /fatima-audio-perturbations Audio Perturbation TTS Gold Sentences A curated set of English sentences for perturbation-based blind-spot evaluation of audio-LLM judges on synthesised speech. Overview This dataset provides original (clean) sentences sampled from established TTS benchmarks. The sentences are designed to be fed through TTS models to generate clean audio (A_gold), then perturbed at the text level (S_gold → S_pert) and re-synthesised (A_pert) to test whether audio-LLM judges can… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/fatima-audio-perturbations.texttext-to-speech1K<n<10K0 likes32 downloads1mo agoHugging Facegupta-tanish /QwQ-Long-CoT-15k-subset-Llama3.1-8B-single-position-regex-perturbations-logps-15tabular100K<n<1M0 likes21 downloads1y agoHugging Facegupta-tanish /QwQ-Long-CoT-10k-subset-Llama3.1-8B-single-position-regex-perturbations-logps-10tabular10K<n<100K0 likes18 downloads1y agoHugging Face