difference-awareness
difference-awareness-version-flip
Difference Awareness: Version-Flip Set
A companion to the
Multidimensional Difference Awareness benchmark. Where the
main set asks whether a model applies the current version of a rule, this set
asks something sharper: when a rule has changed its mind about a demographic
axis, does the model follow the version the question actually cites, or its own
training prior?
Each flip pair is two byte-identical patient scenarios. The only difference
is which version of a guideline or law… See the full description on the dataset page: https://huggingface.co/datasets/Complementarity/difference-awareness-version-flip.multidimensional-difference-awareness-legacy
Multidimensional Difference Awareness
A source-grounded benchmark for bias-import in real eligibility decisions:
whether a model applies a rule's legitimate criteria while refusing to let an
attribute the rule does not use change the outcome.
What is new here. Wang et al. (ACL 2025,
arXiv:2502.01926) measure difference
awareness on general-knowledge facts. Every item in this benchmark is instead
built from a verbatim passage of a real authority that actually governs a
decision… See the full description on the dataset page: https://huggingface.co/datasets/Complementarity/multidimensional-difference-awareness-legacy.Law-Demographic-Bias-Difference-Awareness
Law and Demographic Bias Difference-Awareness Benchmark
A multiple-choice benchmark for testing whether a language model can tell apart two situations
that look alike and demand opposite answers:
neq — the law grants an entitlement to one specific group, so treating both groups
identically is the wrong answer.
eq — the law grants the same right to everyone, so drawing a distinction between the
groups is the wrong answer.
Every item presents two demographic or legal groups, a… See the full description on the dataset page: https://huggingface.co/datasets/Debk/Law-Demographic-Bias-Difference-Awareness.
