datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dcvlm-balanced-200b
DCVLM-Balanced (200B tokens)
DCVLM-Balanced is the balanced-mixture training set from our DataComp-VLM paper.
It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat
WebDataset tar shards so it can be consumed by any training
stack.
This is a 200B-token release consisting of 112,358,849 samples, curated from our DCVLM-large data pool.
The instruction-heavy counterpart (DCVLM-baseline) is available as
dcvlm-baseline-200b, along with… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-balanced-200b.crop-disease-balanced-5022ai2thor-perspective-qa-20k-balanced-splits-with-objai2thor-perspective-qa-100k-balanced-training-v1-splitsprocthor-100-counting-balanceddeeplesion-balanced-2k
DeepLesion Benchmark Subset (Balanced 2K)
This dataset is a curated subset of the DeepLesion dataset, prepared for demonstration and benchmarking purposes. It consists of 2,000 CT lesion samples, balanced across 8 coarse lesion types, and filtered to include lesions with a short diameter > 10mm.
Dataset Details
Source: DeepLesion
Institution: National Institutes of Health (NIH) Clinical Center
Subset size: 2,000 images
Lesion types: lung, abdomen, mediastinum, liver… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/deeplesion-balanced-2k.balanced_wildfire_datasetai2thor-perspective-qa-balanced-400-v2ai2thor-perspective-qa-800-balanced-val-v1ai2thor-perspective-qa-400-balanced-diverseai2thor-perspective-qa-400-balanced-v2refchartqa-balanced-10kai2thor-perspective-qa-800-balanced-diverse-v1ai2thor-perspective-qa-400-balanced-diverse-v2airbus-balanced-subsetBALANCED_MSPP_MSPI_IEMOxray-balanced-datasetSEED_balanced
SEED_balanced
SEED_balanced is the public balanced release of SEED, a benchmark for provenance tracing in sequential deepfake facial edits. Unlike conventional deepfake datasets that focus on single-step manipulations or binary real/fake detection, SEED models multi-step diffusion-based facial editing trajectories and supports three complementary tasks: Authenticity Analysis, Editing Trace Analysis, and Spatial Evidence Analysis. The full SEED benchmark contains 91,526 images… See the full description on the dataset page: https://huggingface.co/datasets/Mengieong/SEED_balanced.simworld-20k-balancedNIH-ChestXray14-Balancedffhq-balanced-dataset
Flickr-Faces-HQ Balanced Dataset
Overview
Flickr-Faces-HQ Balanced Dataset is a race-balanced subset of the original Flickr-Faces-HQ Dataset (FFHQ), aiming to provide a high-resolution facial dataset that reduces racial bias.
Race is automatic labeled using Anzhc/Race-Classification-FairFace-YOLOv8 model, since the official weights of FairFace are unreachable.
This subset includes 7,252 samples at 1024×1024 resolution, with 1,036 images per race class, across 7 race… See the full description on the dataset page: https://huggingface.co/datasets/Yana-Hangabina/ffhq-balanced-dataset.BalanceCC
Dataset Card for Dataset Name
This is the BalanceCC benchmark published in CCEdit, containing 100 videos with varied attributes, designed to offer a comprehensive platform
for evaluating generative video editing, focusing on both controllability and creativity.
Paper Link
Project Page
Dataset Details
Dataset Description
Our objective is to develop a benchmark dataset specifically designed for tasks involving controllable and creative video editing.… See the full description on the dataset page: https://huggingface.co/datasets/RuoyuFeng/BalanceCC.refchartqa-balanced-10k-gt-annotatedHAM_db_enhanced_balancedAI-vs-Real-balancedGTSRB_224x224_balanced
Balanced GTSRB (224x224)
This is a balanced GTSRB dataset containing 43 classes, with 1,000 training samples per class and the same number of test samples as in the original dataset.All images have been resized to 224×224 using interpolation and padding, maintaining aspect ratio.
For classes with fewer than 1,000 training samples, data augmentation was used to supplement the dataset (Note: without any flip transforms, thanks to this post).
For details on how the dataset was… See the full description on the dataset page: https://huggingface.co/datasets/SomeBottle/GTSRB_224x224_balanced.balanced_textileBigEarthS2-all-train-10k-balancedai2thor-path-tracing-qa-train-2point-balanced8-16krefchartqa-balanced-1k-gt-annotated-mini
