iproskurina/bbq-sampled-500-each
BBQ Subset for Dike LeaderBoard We sample 500 examples per bias types from the BBQ dataset to enable fast LLM evaluation for the Dike LeaderBoard. To evaluate the model on this subset, use the following code: git clone --depth 1 https://github.com/upunaprosk/lm-evaluation-harness cd lm-evaluation-harness pip install -e . MODEL_NAME=... # meta-llama/Llama-3-8B lm_eval --model hf \ --model_args pretrained=$MODEL_NAME \ --tasks bbq \ --device cuda:0 \… See the full description on the dataset page: https://huggingface.co/datasets/iproskurina/bbq-sampled-500-each.
BBQ Subset for Dike LeaderBoard
We sample 500 examples per bias types from the BBQ dataset to enable fast LLM evaluation for the Dike LeaderBoard. To evaluate the model on this subset, use the following code:
git clone --depth 1 https://github.com/upunaprosk/lm-evaluation-harness
cd lm-evaluation-harness
pip install -e .
MODEL_NAME=... # meta-llama/Llama-3-8B
lm_eval --model hf \
--model_args pretrained=$MODEL_NAME \
--tasks bbq \
--device cuda:0 \
--batch_size 16Repository for the Bias Benchmark for QA dataset: https://github.com/nyu-mll/BBQ
Authors: Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman.
This repository is a fork of https://huggingface.co/datasets/heegyu/bbq, and adds the "All" configuration containing all subsets.
About BBQ (paper abstract)
It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA). We introduce the Bias Benchmark for QA (BBQ), a dataset of question sets constructed by the authors that highlight attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts. Our task evaluates model responses at two levels: (i) given an under-informative context, we test how strongly responses refect social biases, and (ii) given an adequately informative context, we test whether the model's biases override a correct answer choice. We fnd that models often rely on stereotypes when the context is under-informative, meaning the model's outputs consistently reproduce harmful biases in this setting. Though models are more accurate when the context provides an informative answer, they still rely on stereotypes and average up to 3.4 percentage points higher accuracy when the correct answer aligns with a social bias than when it conficts, with this difference widening to over 5 points on examples targeting gender for most models tested.
Paper Link
"BBQ: A Hand-Built Bias Benchmark for Question Answering" here. The paper has been published in the Findings of ACL 2022 here.
