datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bbq
BBQ
Repository for the Bias Benchmark for QA dataset.
https://github.com/nyu-mll/BBQ
Authors: Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman.
This repository is a fork of https://huggingface.co/datasets/heegyu/bbq, and adds the "All" configuration containing all subsets.
About BBQ (paper abstract)
It is well documented that NLP models learn social biases, but little work has been done… See the full description on the dataset page: https://huggingface.co/datasets/oskarvanderwal/bbq.olmo-eval-bbqThis data comes from the BBQ benchmark.
This is one of the datasets included in the Ai2 Safety Evaluation Suite, and the Olmo evaluation suite.
The repo for Ai2's safety suite includes instructions on how to evaluate models on various safety-related evaluations, including this one.
Permitted Use
The data is provided for benchmarking and evaluation purposes only. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
Disclaimer
This benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-eval-bbq.bbq
BBQ
Repository for the Bias Benchmark for QA dataset.
https://github.com/nyu-mll/BBQ
Authors: Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman.
This repository is a fork of https://huggingface.co/datasets/heegyu/bbq, and adds the "All" configuration containing all subsets.
About BBQ (paper abstract)
It is well documented that NLP models learn social biases, but little work has been done… See the full description on the dataset page: https://huggingface.co/datasets/quantiles/bbq.BBQ_Benchmark_Reasoning_TraceGG-BBQ
Dataset Card for GG-BBQ
German Gender Bias Benchmark for Question Answering (GG-BBQ) for gender bias evaluation in LLMs that support German language.
Dataset Details
Dataset Description
Language(s) (NLP): German
License: cc-by-4.0
Dataset Sources
Repository: https://github.com/shalakasatheesh/GG-BBQ
Paper: https://arxiv.org/abs/2507.16410
Uses
This dataset is to be used to carry out the evaluation of gender bias in language… See the full description on the dataset page: https://huggingface.co/datasets/shalakasatheesh/GG-BBQ.BBQ_Target_Loc_Datasetlanguage: - en pretty_name: "BBQ: Bias Benchmark for Question Answering" tags: - bias-detection - question-answering - fairness - ethics - nlp license: "CC-BY-4.0" task_categories: - question-answering - bias-evaluation
Dataset Card for BBQ: Bias Benchmark for Question Answering
Dataset Summary
The Bias Benchmark for Question Answering (BBQ) is a hand-crafted dataset designed to evaluate implicit social biases in large language models (LLMs) through question-answering tasks. It systematically… See the full description on the dataset page: https://huggingface.co/datasets/bitlabsdb/BBQ_Target_Loc_Dataset.BBQ_target_locBBQ_dataset
language:
- en
pretty_name: "BBQ: Bias Benchmark for Question Answering"
tags:
- bias-detection
- question-answering
- fairness
- ethics
- nlp
license: "CC-BY-4.0"
task_categories:
- question-answering
- bias-evaluation
Dataset Card for BBQ: Bias Benchmark for Question Answering
Dataset Summary
The Bias Benchmark for Question Answering (BBQ) is a hand-crafted dataset designed to evaluate implicit social biases in large language models (LLMs) through… See the full description on the dataset page: https://huggingface.co/datasets/bitlabsdb/BBQ_dataset.bbq_target_loc_dedupBBQABBQ_This dataset contains the same subcategory "Disability_status" from heegyu/bbq but with a longer context that defines the Disability in question, generated by meta-llama/llama-3-3-70b-instruct.
bbqdataset
