MilaNLProc/survey-language-technologies
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions π Best Social Impact Paper Award at ACL 2025 Dataset Summary This dataset comprises responses from 1,000 individuals from diverse socioeconomic backgrounds, collected to study how socioeconomic status (SES) influences interaction with language technologies, particularly generative AI and large language models (LLMs). Participants shared demographic and socioeconomic data, as well asβ¦ See the full description on the dataset page: https://huggingface.co/datasets/MilaNLProc/survey-language-technologies.
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
π Best Social Impact Paper Award at ACL 2025
Dataset Summary
This dataset comprises responses from 1,000 individuals from diverse socioeconomic backgrounds, collected to study how socioeconomic status (SES) influences interaction with language technologies, particularly generative AI and large language models (LLMs). Participants shared demographic and socioeconomic data, as well as up to 10 real prompts they previously submitted to LLMs like ChatGPT, totaling 6,482 unique prompts.
Dataset Structure
The dataset is provided as a single CSV file:
survey_language_technologies.csvNote: All multi-select fields are semicolon (;) separated.Citation
If you use this dataset in your research, please cite the associated paper:
@inproceedings{bassignana-etal-2025-ai,
title = "The {AI} Gap: How Socioeconomic Status Affects Language Technology Interactions",
author = "Bassignana, Elisa and
Curry, Amanda Cercas and
Hovy, Dirk",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.914/",
pages = "18647--18664",
ISBN = "979-8-89176-251-0",
abstract = "Socioeconomic status (SES) fundamentally influences how people interact with each other and, more recently, with digital technologies like large language models (LLMs). While previous research has highlighted the interaction between SES and language technology, it was limited by reliance on proxy metrics and synthetic data. We survey 1,000 individuals from `diverse socioeconomic backgrounds' about their use of language technologies and generative AI, and collect 6,482 prompts from their previous interactions with LLMs. We find systematic differences across SES groups in language technology usage (i.e., frequency, performed tasks), interaction styles, and topics. Higher SES entail a higher level of abstraction, convey requests more concisely, and topics like `inclusivity' and `travel'. Lower SES correlates with higher anthropomorphization of LLMs (using ``hello'' and ``thank you'') and more concrete language. Our findings suggest that while generative language technologies are becoming more accessible to everyone, socioeconomic linguistic differences still stratify their use to create a digital divide. These differences underscore the importance of considering SES in developing language technologies to accommodate varying linguistic needs rooted in socioeconomic factors and limit the AI Gap across SES groups."
}Dataset Curators
- Elisa Bassignana (IT University of Copenhagen)
- Amanda Cercas Curry (CENTAI Institute)
- Dirk Hovy (Bocconi University)
Links
- π Dataset file:
survey_language_technologies.csv - π Survey Interface (may take some time to load): https://nlp-use-survey.streamlit.app/
- π Paper: https://aclanthology.org/2025.acl-long.914/
