cair-nepal/ai-bias-research-landscape
Dataset Card: AI Bias Research Landscape Dataset Summary This dataset contains 692 curated bibliographic records of peer-reviewed and preprint publications on artificial intelligence (AI) and algorithmic bias, published between 2012 and 2026. Each record includes publication metadata (paper title, DOI, authors, author regions, affiliations, publication year, and research domain), author ORCID identifiers, and OpenAlex-derived metadata, including OpenAlex IDs… See the full description on the dataset page: https://huggingface.co/datasets/cair-nepal/ai-bias-research-landscape.
Dataset Card: AI Bias Research Landscape
Dataset Summary
This dataset contains 692 curated bibliographic records of peer-reviewed and preprint publications on artificial intelligence (AI) and algorithmic bias, published between 2012 and 2026. Each record includes publication metadata (paper title, DOI, authors, author regions, affiliations, publication year, and research domain), author ORCID identifiers, and OpenAlex-derived metadata, including OpenAlex IDs, citation counts, referenced works, open-access status, and open-access URLs. Six records lack a DOI and therefore do not contain OpenAlex-sourced identifiers, citation counts, or open-access data; abstracts for these six records were curated manually rather than retrieved from OpenAlex. The dataset is provided in CSV format.
This dataset is the underlying corpus for an interactive atlas of AI bias research, available at https://biasatlas.cair-nepal.org, and for the accompanying paper "Whose fairness? Structural concentration in AI bias research" (under review).
Dataset Structure
Data Instances
Each row is one publication. The dataset has 692 rows and 21 columns.
Data Fields
Data Splits
No splits — the dataset is a single flat table.
Dataset Creation
Curation Rationale
The corpus was assembled to characterize the geographic, institutional, and thematic structure of AI bias research — who produces it, where, and in which application areas — rather than to catalogue bias-mitigation methods themselves.
Source Data
Records were identified by systematically querying IEEE Xplore, the ACM Digital Library, Scopus, ScienceDirect, and Engineering Village, using search strings centered on bias, artificial intelligence, and decision-making. This was supplemented by citation snowballing (via Connected Papers and Litmaps) and by directly searching the ACM Conference on Fairness, Accountability, and Transparency (FAccT) proceedings, given the venue's direct relevance to AI bias research.
Inclusion criteria required peer-reviewed status (or, for 69 records, arXiv preprints — 10 of which are FAccT papers confirmed peer-reviewed) and explicit focus on AI bias. The search was limited to publications from 2015 onward to ensure relevance to contemporary AI systems, with one exception: Dwork et al. (2012), a foundational fairness paper identified through citation snowballing and retained given its centrality to the field.
Each record was manually screened for topical relevance and assigned to one of five thematic domains. Abstracts were retrieved via the OpenAlex API where a DOI match existed; for the 6 records without a matching DOI, abstracts were collected manually.
Annotations
Thematic domain assignment (Domain) and geographic focus (Focus Region) were assigned manually during corpus curation, not derived from an automated classifier. A separate semantic-clustering analysis (SBERT embeddings + UMAP + HDBSCAN) was used in the accompanying paper to validate these manual domain assignments post hoc; domain labels were never used to inform the clustering itself.
Considerations for Using the Data
Discussion of Biases and Limitations
- Database and language coverage. The corpus is drawn primarily from English-language, Global-North-indexed sources (IEEE Xplore, ACM Digital Library, Scopus, ScienceDirect, Engineering Village), supplemented with FAccT proceedings and citation snowballing. This sampling strategy likely under-represents research published in other languages, regional venues, and outlets not covered by major indexing services. Geographic disparities observed using this dataset cannot be fully disentangled from database coverage.
- Domain labels are analytical groupings, not strict categories. Semantic clustering shows substantial but incomplete correspondence with the five manually assigned domains; some overlap between domains should be expected.
- Citation counts are a point-in-time snapshot from the OpenAlex API at time of curation and will not reflect citations accrued afterward.
- 6 records have no DOI and consequently no OpenAlex-sourced citation count, open-access status, or referenced-works data; their
abstractfield was populated manually rather than from OpenAlex.
Personal and Sensitive Information
The dataset contains author names, affiliations, and ORCID identifiers — all already public bibliographic information associated with the cited publications. No additional personal data is included.
Additional Information
Licensing Information
CC-BY-4.0. (Confirm this matches your intended license before publishing — this reflects that the dataset is aggregated bibliographic metadata, not full-text content, but you should verify this is what you want.)
Citation Information
@misc{shrestha2026fairnessstructuralconcentrationai,
title={Whose fairness? Structural concentration in AI bias research},
author={Abhash Shrestha and Subigya Gautam and Anu Sapkota and Sanju Tiwari and Tek Raj Chhetri},
year={2026},
eprint={2607.05574},
archivePrefix={arXiv},
primaryClass={cs.CY},
url={https://arxiv.org/abs/2607.05574},
}Links
- Interactive atlas / dashboard: https://biasatlas.cair-nepal.org
- Code repository: https://github.com/CAIRNepal/biasatlas
Dataset Curators
Abhash Shrestha, Subigya Gautam, Anu Sapkota, Sanju Tiwari, Tek Raj Chhetri — Center for Artificial Intelligence (AI) Research Nepal.
