datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
InterviewForge_GenDS
Synthetic Data Generation
Model & Infrastructure
The dataset was generated using the mistral:latest Large Language Model running locally via the Ollama framework. This model was explicitly selected because it balances advanced reasoning capabilities with hardware efficiency, allowing the execution of 10,944 complex generation requests entirely locally on an RTX 3080 GPU without incurring API costs. Additionally, Mistral demonstrated exceptional reliability in… See the full description on the dataset page: https://huggingface.co/datasets/Davichick/InterviewForge_GenDS.gender-bias-PE
Dataset Card for gender-bias-PE data
Dataset Description
The gender-bias-PE dataset contains the post-edits and associated behavioural data of the human-centered experiments presented in the paper:
What the Harm? Quantifying the Tangible Impact of Gender Bias in Machine Translation with a Human-centered Study accepted at EMNLP 2024.
The dataset allows to study the impact of gender bias in Machine Translation (MT) via human-centered measures like post-editing effort (i.e.… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/gender-bias-PE.gender_congress_117-118
Dataset Card for Dataset gender_congress_117-118
This dataset consists of definitions for gender and gender-related terms from congressional bills proposed between January 2021 and November 2023 (US Congress Sessions 117 and 118).
It focuses on bills that feature the term "gender" prominently in the bill title, bill summary, or in the bill text.
Dataset Description
Curated by: Filipa Calado
Language(s) (NLP): English
License: Apache 2.0
Uses
This… See the full description on the dataset page: https://huggingface.co/datasets/gofilipa/gender_congress_117-118.
