Aeryx-ai/aae-dialect-fairness
AAE Dialect-Fairness Set A reusable set for debiasing hate/toxicity classifiers against African-American English (AAE) false positives. Off-the-shelf classifiers flag benign AAE text as toxic at 2x+ the rate of benign General-American English (Sap et al. 2019). This set provides (1) high-AAE benign text to augment training so a model can't use dialect as a toxicity cue, and (2) a held-out dialect-balanced benchmark to measure the residual gap. Built for the… See the full description on the dataset page: https://huggingface.co/datasets/Aeryx-ai/aae-dialect-fairness.
128
