datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
music
music
A large-scale music dataset containing artist, release and track names + URLs.
Sources
Source
Rows
soundcloud
200M
discogs
178M
applemusic
104M
lastfm
102M
deezer
90M
bandcamp
49M
musicbrainz
24M
youtube-videos
10M
youtube
7M
metal-archives
5M
vgmdb
2M
film-tv
3M
newgrounds
1M
onlineradiobox
466K
beatport
454K
Total
776M
Please note that this dataset has not yet been deduplicated across sources. A cross-source… See the full description on the dataset page: https://huggingface.co/datasets/fairygaze/music.FAIR1M
FAIR1M
The FAIR1M dataset is a fine-grained object recognition and detection dataset that focuses on high-resolution (0.3-0.8m) RGB images taken by the Gaogen (GF) satellites and extracted from Google Earth. It consists of a collection of 15,000 high-resolution images that cover various objects and scenes. The dataset provides annotations in the form of rotated bounding boxes for objects belonging to 5 main categories (ships, vehicles, airplanes, courts, and roads), further… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/FAIR1M.fairytail
Bangumi Image Base of Fairy Tail
This is the image base of bangumi Fairy Tail, we detected 270 characters, 33650 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the characters'… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/fairytail.FairSeg
Dataset Card: FairSeg
Dataset Summary
FairSeg is a large-scale ophthalmology dataset for studying fairness in medical image segmentation. It contains 10,000 SLO fundus images with pixel-wise optic disc and cup segmentation masks, paired with comprehensive demographic annotations. The dataset is designed to benchmark and improve demographic equity in segmentation models, including foundation models such as SAM (Segment Anything Model).
This dataset was introduced at ICLR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairSeg.fairytail100nenquest
Bangumi Image Base of Fairy Tail: 100-nen Quest
This is the image base of bangumi Fairy Tail: 100-nen Quest, we detected 134 characters, 10767 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/fairytail100nenquest.FairFace
Dataset Card for FairFace
Dataset Summary
FairFace is a face image dataset which is race balanced. It contains 108,501 images from 7 different race groups: White, Black, Indian, East Asian, Southeast Asian, Middle Eastern, and Latino.
Images were collected from the YFCC-100M Flickr dataset and labeled with race, gender, and age groups.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceM4/FairFace.mmlu-autotranslatedtruthfulqa-autotranslatedFAIR1M1This is the train set for FAIR1M1.
The validation set is in LittleCollections/FAIR1M2
**Data Structure:
--train
--part1
--images.zip
--labelXml.zip
--part2
--images-1.zip
--images-2.zip
--labelXmls.zip
--readme.txt
nanopath-fairness-tiles
nanopath-fairness-tiles
Pre-tiled histopathology patches from CPTAC whole-slide images, used as the
external out-of-distribution validation set for a study on pretraining-time
vs. post-hoc fairness in histopathology foundation models.
Contents
Per-cohort folders, each slides_full/<slide_id>.parquet (one row per tile:
case_id, slide_id, tile_idx, image) + labels.tsv:
cohort
organ / task
slides
cptac_lung
NSCLC — LUAD vs LSCC subtype
604
cptac_gbm
GBM —… See the full description on the dataset page: https://huggingface.co/datasets/ryankim17920/nanopath-fairness-tiles.arc-easy-autotranslatedFairVision
Dataset Card: Harvard-FairVision
Dataset Summary
Harvard-FairVision is the first large-scale medical fairness dataset with both 2D and 3D imaging data, covering three major eye diseases affecting approximately 380 million people worldwide. It contains 30,000 subjects (10,000 per disease) across Age-Related Macular Degeneration (AMD), Diabetic Retinopathy (DR), and glaucoma, each with paired SLO fundus photos and 3D OCT B-scans and six demographic identity attributes.
This… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVision.FAIR_SAR_2M
FAIR-SAR-2M
FAIR-SAR-2M is a large-scale synthetic aperture radar (SAR) target-image dataset containing measured and generated samples with dense full-angle coverage. It is organized into three resolution domains and supports SAR target recognition, image generation, data augmentation, and downstream object-detection research.
Dataset Domains
Domain
Source
Angle semantics
Paste-back support
HR
High-resolution SAR target chips
Azimuth angle
No; HR is… See the full description on the dataset page: https://huggingface.co/datasets/SolrenLenira/FAIR_SAR_2M.FairytaleQAFairytaleQA dataset, an open-source dataset focusing on comprehension of narratives, targeting students from kindergarten to eighth grade. The FairytaleQA dataset is annotated by education experts based on an evidence-based theoretical framework. It consists of 10,580 explicit and implicit questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations.WarBias
WarBias
WarBias is a synthetic bilingual dataset of 1,059 stereotype / counter-stereotype / unrelated triplets per language, covering 127 case families about veterans and internally displaced people (IDPs). It also includes 240 base evaluation questions and 1,320 matched intersectional QA variants per language.
Languages: English (en) and Ukrainian (uk).Status: Model-reviewed; human validation pending.
Load
from datasets import load_dataset
uk =… See the full description on the dataset page: https://huggingface.co/datasets/FairForget/WarBias.FairytaleQA\
The FairytaleQA dataset focusing on narrative comprehension of kindergarten to eighth-grade students. Generated by educational experts based on an evidence-based theoretical framework, FairytaleQA consists of 10,580 explicit and implicit questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations. This is for the Question Generation Task of FairytaleQA.fairness-prm-training-dataFairyTaleFusionLibra_bimanual_fairinoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "FAIRINO_BIMANUAL_SIM",
"total_episodes": 100,
"total_frames": 26368,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Libra-TeV/Libra_bimanual_fairino.arch-fairness-glbsLibra_bimanual_fairino_plug_in_container_r_a_angle_variationsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "FAIRINO_BIMANUAL_SIM",
"total_episodes": 95,
"total_frames": 20345,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:95"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Libra-TeV/Libra_bimanual_fairino_plug_in_container_r_a_angle_variations.fairy_talesConcatenated and edited collection of fairy tales taken from Project Gutenberg.
Texts:
https://www.gutenberg.org/files/2591/2591-0.txt
https://www.gutenberg.org/files/503/503-0.txt
https://www.gutenberg.org/files/7277/7277-0.txt
https://www.gutenberg.org/cache/epub/35862/pg35862.txt
https://www.gutenberg.org/cache/epub/69739/pg69739.txt
https://www.gutenberg.org/files/2435/2435-0.txt
https://www.gutenberg.org/cache/epub/7871/pg7871.txt
https://www.gutenberg.org/files/8933/8933-0.txt… See the full description on the dataset page: https://huggingface.co/datasets/vicclab/fairy_tales.FairVLMed
Dataset Card: Harvard-FairVLMed
Dataset Summary
Harvard-FairVLMed is the first fair vision-language medical dataset designed for studying fairness in medical vision-language (VL) foundation models. It contains 10,000 SLO fundus images paired with de-identified clinical notes and comprehensive demographic annotations, enabling in-depth fairness analysis across four protected attributes: race, gender, ethnicity, and preferred language.
This dataset was introduced at CVPR… See the full description on the dataset page: https://huggingface.co/datasets/harvardairobotics/FairVLMed.FairDialogue
Dataset Card for FairDialogue
Dataset Description
FairDialogue is a benchmark resource for evaluating bias in end-to-end spoken dialogue models (SDMs).
While biases in large language models (LLMs) have been widely studied, spoken dialogue systems with audio input/output remain underexplored. FairDialogue provides stimulus data (audio, transcripts, and prompts) that can be used together with the official evaluation scripts to measure fairness in decision-making and… See the full description on the dataset page: https://huggingface.co/datasets/yihao005/FairDialogue.Libra_bimanual_fairino_plug_in_container_mergedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "FAIRINO_BIMANUAL_SIM",
"total_episodes": 165,
"total_frames": 31881,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:165"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Libra-TeV/Libra_bimanual_fairino_plug_in_container_merged.Libra_bimanual_fairino_plug_in_container_right_armThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "FAIRINO_BIMANUAL_SIM",
"total_episodes": 70,
"total_frames": 11468,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 15,
"splits": {
"train": "0:70"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Libra-TeV/Libra_bimanual_fairino_plug_in_container_right_arm.fairtalking-second-work
Second Work on FairTalking-Bench: CTA / PSM / RADR
Three novel, identity-agnostic, generator-agnostic deepfake detection methods
built on top of the FairTalking-Bench dataset.
Methods
Name
Core idea
Trainable?
Backbone
CTA
Causal Time-Arrow asymmetry (A->V vs V->A predictive loss gap)
Yes
TimeSformer / ViT + Wav2Vec2
PSM
Phoneme-Synchrony Manifold (phoneme-conditioned lip prior)
Yes
Wav2Vec2-phoneme + MediaPipe lip
RADR
Reference-Anchored Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/huahua123313/fairtalking-second-work.speech_fairness_synthfair1m
FAIR1M
The FAIR1M dataset is a fine-grained object recognition and detection dataset that focuses on high-resolution (0.3-0.8m) RGB images taken by the Gaogen (GF) satellites and extracted from Google Earth. It consists of a collection of 15,000 high-resolution images that cover various objects and scenes. The dataset provides annotations in the form of rotated bounding boxes for objects belonging to 5 main categories (ships, vehicles, airplanes, courts, and roads), further… See the full description on the dataset page: https://huggingface.co/datasets/ShantyCam/fair1m.fairy-stockfish_data
