datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StereoMIS_processedStereoSet-UK-Unlearning
StereoSet-UK Unlearning
StereoSet-UK Unlearning contains 2,101 Ukrainian full-sentence triplets derived from the
intrasentence portion of the StereoSet development set. The Ukrainian sentences were translated
with the DeepL API and received technical cleanup. English source text is omitted.
Each triplet assigns the stereotype sentence to forget_uk, the anti-stereotype sentence to
retain_uk, and the unrelated sentence to control_uk. Five items with duplicate translated
candidates… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/StereoSet-UK-Unlearning.StereoSet-UK-Eval
StereoSet-UK Eval
StereoSet-UK Eval contains 949 Ukrainian masked triplets derived from the intrasentence portion
of the StereoSet development set. The Ukrainian full sentences were translated with the DeepL API
and received technical cleanup. English source text is omitted.
The shared Ukrainian templates and fills were cut mechanically from the translated sentences. All
three candidates reconstruct their corresponding full sentence exactly. Each retained template
passed the… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/StereoSet-UK-Eval.hexopyranose_stereoisomers
Hexopyranose Stereoisomers — DFT-Optimized Geometries
This dataset contains DFT-optimized 3D geometries for all 32 stereoisomers of hexopyranose (OCC1OC(O)C(O)C(O)C1O), a six-membered sugar ring with 5 stereocenters. It is designed for benchmarking molecular embedding methods that require fine stereochemical discrimination, such as the Coulomb Matrix and Bag of Bonds representations.
Background
Hexopyranose has 5 stereocenters, yielding 2⁵ = 32 possible stereoisomers.… See the full description on the dataset page: https://huggingface.co/datasets/UnidentifiedHidden/hexopyranose_stereoisomers.stereotactic-radiosurgery-k1-with-segmentation
🎯 Stereotactic Radiosurgery Dataset (SRS)
🏥 400 synthetic patient records describing the clinical, imaging, segmentation, and treatment-planning metadata of a stereotactic radiosurgery workflow, delivered as a single CSV with placeholder file paths.
⚠️ Disclaimer: This is a metadata-only synthetic dataset. It contains no real patients, no image files, and no segmentation files. Every record is generated; paths in the imaging and segmentation columns are placeholders that do… See the full description on the dataset page: https://huggingface.co/datasets/Taylor658/stereotactic-radiosurgery-k1-with-segmentation.StereoBiasThis repository accompanies the paper:
Aditya Tomar, Rudra Murthy, and Pushpak Bhattacharyya. 2025. Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach. In Findings of the Association for Computational Linguistics: ACL 2025, pages 17304–17317, Vienna, Austria. Association for Computational Linguistics.
Citation
@inproceedings{tomar-etal-2025-stereotype,
title = "Stereotype Detection as a Catalyst for Enhanced Bias Detection: A… See the full description on the dataset page: https://huggingface.co/datasets/aditya20t/StereoBias.StereoHoax-GL
StereoHoax-GL
Dataset Summary
StereoHoax-GL is a Galician translated and linguistically revised version of the test partition of the StereoHoax-ES subset included in DETESTS-Dis.
The dataset is intended as an evaluation resource for stereotype and discriminatory-content detection in Galician. It preserves the original identifiers, metadata fields, and annotation labels from the StereoHoax-ES test partition, while replacing the original Spanish text field with its… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/StereoHoax-GL.Stereotype-Elicitation-Prompt-LibraryStereoDetect
Paper
You can access the paper on here: StereoDetect: Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological Underpinnings
This work is accepted at the Findings of the Association for Computational Linguistics: EMNLP 2025.
Dataset Overview
The dataset is divided into three splits:
train.csv — used for model training
val.csv — used for validation and hyperparameter tuning
test.csv — used for final evaluation
Label Definitions… See the full description on the dataset page: https://huggingface.co/datasets/thenlpresearcher/StereoDetect.GBEM-UA
Overview
This dataset was created for the paper “GBEM-UA: Gender Bias Evaluation and Mitigation for Ukrainian Large Language Models” to study gender bias in the "hiring problem" within the Ukrainian language, focusing on how grammatical gender (e.g., feminitive vs. non-feminitive forms) may influence model predictions.
Dataset Structure
Each row includes:
sentence: the candidate description
profession: base profession name
experience: "relevant" or "irrelevant"… See the full description on the dataset page: https://huggingface.co/datasets/Stereotypes-in-LLMs/GBEM-UA.
