CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01harvardairobotics /FairSeg Dataset Card: FairSeg Dataset Summary FairSeg is a large-scale ophthalmology dataset for studying fairness in medical image segmentation. It contains 10,000 SLO fundus images with pixel-wise optic disc and cup segmentation masks, paired with comprehensive demographic annotations. The dataset is designed to benchmark and improve demographic equity in segmentation models, including foundation models such as SAM (Segment Anything Model). This dataset was introduced at ICLR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairSeg.textimage-segmentation10K<n<100K0 likes3.3k downloads5mo agoHugging Face02harvardairobotics /FairVision Dataset Card: Harvard-FairVision Dataset Summary Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes. This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.imageimage-classification10K<n<100K0 likes499 downloads5mo agoHugging Face03harvardairobotics /FairVLMed Dataset Card: Harvard-FairVLMed Dataset Summary Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language. This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.imageimage-classification10K<n<100K0 likes282 downloads5mo agoHugging Face04harvardairobotics /FairFedMed Dataset Card: FairFedMed Dataset Summary FairFedMed is the first federated learning (FL) benchmark dataset for medical imaging with demographic annotations, designed to study group fairness across institutions in a federated setting. It comprises two subsets spanning ophthalmology and chest radiology, enabling research on fairness-aware federated learning under realistic cross-institutional data heterogeneity. This dataset was introduced in the IEEE Transactions on… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairFedMed.tabularimage-classification10K<n<100K1 likes189 downloads5mo agoHugging Face05FairForget /BBQ-UK BBQ-UK: Ukrainian Translation BBQ-UK is a Ukrainian translation of the Bias Benchmark for Question Answering (BBQ). It preserves the original paired ambiguous and disambiguated contexts, answer positions, labels, bias-target metadata, categories, and question polarity. The public release contains Ukrainian task text only. English source text is not included. Dataset status 28,503 context pairs 57,006 task rows 28,503 ambiguous and 28,503 disambiguated rows 11… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/BBQ-UK.tabularquestion-answering10K<n<100K0 likes140 downloads1mo agoHugging Face06fairnlp /holistic-bias Usage When downloading, specify which files you want to download and set the split to train (required by datasets). from datasets import load_dataset nouns = load_dataset("fairnlp/holistic-bias", data_files=["nouns.csv"], split="train") sentences = load_dataset("fairnlp/holistic-bias", data_files=["sentences.csv"], split="train") Dataset Card for Holistic Bias This dataset contains the source data of the Holistic Bias dataset as described by Smith et. al. (2022).… See the full description on the dataset page: https://huggingface.co/datasets/fairnlp/holistic-bias.text100K<n<1M8 likes132 downloads3y agoHugging Face07harvardairobotics /FairDomain Dataset Card: Harvard-FairDomain Dataset Summary Harvard-FairDomain is a large-scale ophthalmology dataset designed for studying fairness under domain shift in medical image analysis. It supports both image segmentation and classification tasks, with 10,000 samples per task drawn from 10,000 unique patients. The dataset introduces an additional imaging modality — en-face fundus images — alongside the original scanning laser ophthalmoscopy (SLO) fundus images, enabling… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairDomain.tabularimage-segmentation10K<n<100K0 likes107 downloads5mo agoHugging Face08FAIRC /SLMTrainBench SLMTrainBench SLMTrainBench is the measurement dataset for When Peak Floating-Point Throughput Misleads: Utilization and Cost Frontiers for Small Language Model Pretraining. It maps batch-saturated, single-GPU training performance for nine dense decoder-only models from 150 million to 8 billion parameters across ten NVIDIA GPUs and context lengths from 512 to 32,768 tokens. The dataset contains 2,963 tested batch configurations, including successful measurements and… See the full description on the dataset page: https://huggingface.co/datasets/FAIRC/SLMTrainBench.tabular1K<n<10K0 likes90 downloads1mo agoHugging Face09criteo /FairJob FairJob: A Real-World Dataset for Fairness in Online Systems Summary This dataset is released by Criteo to foster research and innovation on Fairness in Advertising and AI systems in general. See also Criteo pledge for Fairness in Advertising. The dataset is intended to learn click predictions models and evaluate by how much their predictions are biased between different gender groups. The associated paper is available at Vladimirova et al. 2024. License… See the full description on the dataset page: https://huggingface.co/datasets/criteo/FairJob.tabulartabular-classification1M<n<10M9 likes84 downloads2y agoHugging Face10fairleap-ai /fairleap-driver-earnings-regression-500 Fairleap Driver Earnings Regression 500 📘 Dataset Overview A small synthetic tabular dataset of ride-hailing driver work sessions, built for the Fairleap AI project — a platform addressing income uncertainty and wellbeing for Gojek/GOTO drivers in Indonesia. Each row is one work session: when it happened, where, how long it ran, how many rides were completed, and what it paid. The dataset backs two regression targets: earnings (Indonesian Rupiah per session) and… See the full description on the dataset page: https://huggingface.co/datasets/fairleap-ai/fairleap-driver-earnings-regression-500.tabulartabular-regressionn<1K1 likes74 downloads1mo agoHugging Face11harvardairobotics /FairGenMed Dataset Card: FairGenMed Dataset Summary FairGenMed is the first dataset for studying fairness in medical generative models. It provides detailed quantitative clinical measurements alongside demographic annotations to investigate the semantic correlation between text prompts and anatomical regions across demographic subgroups. The dataset supports both generative model evaluation and downstream classification tasks for glaucoma detection. This dataset accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairGenMed.imageimage-classification10K<n<100K0 likes71 downloads5mo agoHugging Face12FairForget /WinoPron-UK WinoPron-UK WinoPron-UK is a Ukrainian translation and grammatical adaptation of WinoPron, a coreference benchmark for evaluating performance and pronoun-related bias. The release covers all 180 complementary WinoPron pairs. Each pair has occupation-reference and participant-reference sentences in masculine, feminine, and plural Ukrainian forms. Configurations double: 1,080 sentences containing an occupation and another participant. single: 1,080 matched control… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoPron-UK.textquestion-answering1K<n<10K0 likes68 downloads1mo agoHugging Face13FairForget /WinoBias-UK-Controlled WinoBias-UK Controlled WinoBias-UK Controlled is a Ukrainian coreference-bias evaluation set derived from WinoBias. Occupation candidates are expressed through gender-neutral Ukrainian descriptions, while the evaluated pronoun remains gendered. Current release This release contains 1,578 source items and 3,156 evaluation rows: validation: 788 source items and 1,576 rows; test: 790 source items and 1,580 rows; each source item has one masculine-pronoun and one… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoBias-UK-Controlled.textquestion-answering1K<n<10K0 likes65 downloads1mo agoHugging Face14FairForget /WinoBias-UK-Natural WinoBias-UK Natural WinoBias-UK Natural is a Ukrainian gender-counterfactual coreference evaluation set derived from WinoBias. It provides natural masculine, feminine, mixed, and cross-reference variants while preserving the source event and participant roles. Current release This preview contains 279 validated WinoBias source pairs and 1,674 Ukrainian variants. The current release covers the validation Type 1 stratum. Full validation and test coverage is in… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WinoBias-UK-Natural.textquestion-answering1K<n<10K0 likes61 downloads1mo agoHugging Face15holistic-ai /job-fair-resumetextn<1K3 likes46 downloads2y agoHugging Face16FairForget /StereoSet-UK-Unlearning StereoSet-UK Unlearning StereoSet-UK Unlearning contains 2,101 Ukrainian full-sentence triplets derived from the intrasentence portion of the StereoSet development set. The Ukrainian sentences were translated with the DeepL API and received technical cleanup. English source text is omitted. Each triplet assigns the stereotype sentence to forget_uk, the anti-stereotype sentence to retain_uk, and the unrelated sentence to control_uk. Five items with duplicate translated candidates… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/StereoSet-UK-Unlearning.text1K<n<10K0 likes46 downloads1mo agoHugging Face17FairForget /StereoSet-UK-Eval StereoSet-UK Eval StereoSet-UK Eval contains 949 Ukrainian masked triplets derived from the intrasentence portion of the StereoSet development set. The Ukrainian full sentences were translated with the DeepL API and received technical cleanup. English source text is omitted. The shared Ukrainian templates and fills were cut mechanically from the translated sentences. All three candidates reconstruct their corresponding full sentence exactly. Each retained template passed the… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/StereoSet-UK-Eval.textn<1K0 likes42 downloads1mo agoHugging Face18frostMap /fair-refrigerator-13ba64 fair-refrigerator-13ba64 Synthetic weather test data: 46 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/frostMap/fair-refrigerator-13ba64.tabularn<1K0 likes31 downloads12d agoHugging Face19renhehuang /resume-job-fairness-eval Resume-Job Fairness Evaluation Dataset (pairs_longtext) English | 中文 English Dataset Summary This dataset contains 960 resume-job pairs designed for fairness evaluation in AI-powered hiring systems. Each pair includes full-text resumes and job descriptions, along with sensitive attribute labels (educational background category) to enable demographic parity and counterfactual fairness testing. Primary Use Case: Evaluate bias and fairness in resume-job matching… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/resume-job-fairness-eval.tabulartext-classificationn<1K0 likes30 downloads10mo agoHugging Face20kuldeepbishnoi29 /adult-fairness-experimentstabular1M<n<10M0 likes30 downloads9mo agoHugging Face21fatmaElsafoury2022 /SST_sentiment_fairness_data Sentiment fairness dataset ================================ This dataset is to measure gender fairness in the downstream task of sentiment analysis. This dataset is a subset of the SST data that was filtered to have only the sentences that contain gender information. The python code used to create this dataset can be found in the prepare_sst.ipyth file. Then the filtered datset was labeled by 4 human annotators who are the authors of this dataset. The annotations… See the full description on the dataset page: https://huggingface.co/datasets/fatmaElsafoury2022/SST_sentiment_fairness_data.tabulartext-classificationn<1K2 likes27 downloads3y agoHugging Face22fair-forward /hcsa_indonesia Dataset Card for HCSA Forest Plot Data 2023 This dataset contains information on forest field plot inventory data collected using the High Carbon Stock Approach (HCSA) methodology. The data serves as validation and training data for large-scale indicative HCS forest maps produced with the HCSA Largescale Mapping Framework, as part of a project funded by the GIZ Fair Forward Initiative. It encompasses various parameters pertaining to land cover, carbon content, tree characteristics… See the full description on the dataset page: https://huggingface.co/datasets/fair-forward/hcsa_indonesia.tabular1K<n<10K0 likes13 downloads2y agoHugging Face23kuldeepbishnoi29 /bank-fairness-experimentstabular10K<n<100K0 likes12 downloads9mo agoHugging Face24fairxllm /allsides-8valuesgated Dataset Card for allsides-8values The allside-8values dataset is collected and used in the work of 'Fine-Grained Interpretation of Political Opinions in Large Language Models'. Source Data: The dataset is extracted from Allsides across different domain dimensions (based on a fine-grained 8-values scheme). Exploratory Data Analysis The synthetic dataset can be constructed using this dataset and a given LLM (e.g., GPT-4o). Below is the EDA of constructed dataset.… See the full description on the dataset page: https://huggingface.co/datasets/fairxllm/allsides-8values.textn<1K0 likes11 downloads1y agoHugging Face25xuzihao112 /FairJob FairJob: A Real-World Dataset for Fairness in Online Systems Summary This dataset is released by Criteo to foster research and innovation on Fairness in Advertising and AI systems in general. See also Criteo pledge for Fairness in Advertising. The dataset is intended to learn click predictions models and evaluate by how much their predictions are biased between different gender groups. The associated paper is available at Vladimirova et al. 2024. License… See the full description on the dataset page: https://huggingface.co/datasets/xuzihao112/FairJob.tabulartabular-classification1M<n<10M0 likes11 downloads6mo agoHugging Face26fairnightzz /os-ankitextn<1K0 likes9 downloads3y agoHugging Face27academic-datasets /building-bridges-gender-fair-german-mttabularn<1K0 likes8 downloads2y agoHugging Face28fairnightzz /anki-generatedtextn<1K0 likes6 downloads3y agoHugging Face29tools-o /Fair-PP-CNtabular10K<n<100K0 likes4 downloads1y agoHugging Face30tools-o /Fair-PP-CN-subtabularn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.