elledilara/CoDIGIT-Edinburgh
This is an experimental, creative and participatory effort to implement a co-designed, situated and small-scale fine-tuning dataset for Large Language Models (LLMs). The dataset contains 1197 items: • 60 gender-oriented, co-designed prompts for sentence completion; • 180 model completions generated by LLaMA 3.1 8B (three for each prompt); • 897 gender bias scores, assigned to model responses by participants based on the experimental gender bias scale reported below; • 60… See the full description on the dataset page: https://huggingface.co/datasets/elledilara/CoDIGIT-Edinburgh.
This is an experimental, creative and participatory effort to implement a co-designed, situated and small-scale fine-tuning dataset for Large Language Models (LLMs).
The dataset contains 1197 items:
• 60 gender-oriented, co-designed prompts for sentence completion;
• 180 model completions generated by LLaMA 3.1 8B (three for each prompt);
• 897 gender bias scores, assigned to model responses by participants based on the experimental gender bias scale reported below;
• 60 participant-written completions of the initial 60 gender-oriented, co-designed prompts.
For full information please see:
CoDIGIT-Edinburgh datasheet.pdf
If you use the dataset, please cite:
Dal Molin, L. (2026) Co-Designed Gender Instruction Tuning Dataset (CoDIGIT-Edinburgh). Zenodo. DOI: 10.5281/zenodo.19064467
