datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dialectic-preferences-bias-aae-sae-parallel
Dialectic Preferences Bias Dataset
Dataset Description
Overview
This dataset is part of a research study examining dialectic preference bias in Large Language Models (LLMs). It contains paired sentences in African American English (AAE) and Standard American English (SAE), used to analyze potential biases in language models' treatment of different dialects.
The dataset contains two columns:
african_american_english: Text samples in African American English… See the full description on the dataset page: https://huggingface.co/datasets/furquan/dialectic-preferences-bias-aae-sae-parallel.AAEMime
Dataset README
Introduction
This dataset contains the Annotated Responses for our AAE Study which comprises annotated conversational data and Tweets.
Dataset Overview
Annotations: Each entry includes response type, text, annotation prefix, labels, and Likert scale values.
Files: Final MUSE CORAAL Many Shot Annotations (contains annotations for the CORAAL and NPR corpus) and Final MUSE Tweets Many Shot Annotations (contains annotations for Tweets and NPR corpus)
Formats Available:… See the full description on the dataset page: https://huggingface.co/datasets/kweCobi/AAEMime.
