datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fairness-pruning-pairs-es
Fairness Pruning Prompt Pairs — Spanish
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis, with a focus on Spanish-language bias patterns.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs. It is the Spanish companion to the English dataset, enabling… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-es.fairness-pruning-pairs-en
Fairness Pruning Prompt Pairs — English
Prompt pair dataset for neuronal bias mapping in Large Language Models. Designed to identify which MLP neurons encode demographic bias through differential activation analysis.
This dataset is part of the Fairness Pruning research project, which investigates bias mitigation through activation-guided MLP width pruning in LLMs.
Dataset Summary
Each record contains a pair of prompts that are identical except for a single… See the full description on the dataset page: https://huggingface.co/datasets/oopere/fairness-pruning-pairs-en.stablebridge-pruning-eval
Stablebridge Pruning Evaluation Dataset
Evaluation dataset for the Stablebridge context pruner/highlighter model, measuring sentence-level pruning quality on US stablecoin regulatory documents.
Dataset Structure
File
Records
Description
queries.jsonl
93
Regulatory queries (JSONL with _id and text fields)
corpus.jsonl
38
US stablecoin regulatory documents (full text)
qrels/test.tsv
2,704
Query-document relevance judgments
pruning_labels/test.jsonl
10,006… See the full description on the dataset page: https://huggingface.co/datasets/sugiv/stablebridge-pruning-eval.visual_pruning_klstablebridge-regulatory-pruning-eval
