datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FUSU-Fine_grained_Urban_Semantic_Understanding
About:
FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection and domain adaptation tasks. This data comprises 2 parts:
Bi-temporal high-resolution satellite RGB images with fine-grained annotations.
Monthly revisited Sentinel-2 and Sentinel-1 images.
Details:
1.… See the full description on the dataset page: https://huggingface.co/datasets/sp-juni/FUSU-Fine_grained_Urban_Semantic_Understanding.FiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/LIMinghan/FiVE-Fine-Grained-Video-Editing-Benchmark.test-fine-grained-challenges
Fine-Grained Challenges
Groups of morphologically similar species for evaluating and tuning fine-grained classifiers on top of BioCLIP ecosystem model embeddings. Each group gathers species that are easily confused with one another, and the groups span several clades so a classifier can be probed on the distinctions that actually matter rather than on coarse taxonomy.
The corpus lives in a Lance dataset and can be acted on as a whole, on any single group independently, or on any… See the full description on the dataset page: https://huggingface.co/datasets/thompsonmj/test-fine-grained-challenges.fine-grained-challenges
Dataset Card for Fine-Grained Challenges
Fine-Grained Challenges collects focused groups of visually similar animals for testing biological image classifiers. It combines images, taxonomy, provenance, and frozen embeddings from three BioCLIP-family models in one Lance dataset.
Dataset Details
Release v0.1.0 contains three challenge groups:
Challenge group
Focus
Rows
Species-labeled rows
Genus or higher rows
Labeled species
Peromyscus
Deermice and close… See the full description on the dataset page: https://huggingface.co/datasets/imageomics/fine-grained-challenges.OpenCaption-FineGrained
OpenCaption-FineGrained
OpenCaption-FineGrained is a high-quality dense image captioning dataset containing fine-grained, long-form image descriptions synthesized using the Qwen3.6 Multimodal model. Rather than generating a single caption per image, the dataset follows a multi-sample caption selection pipeline, where multiple candidate captions are generated for every image and the highest-quality caption is selected using an automated quality filtering strategy.
The resulting… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenCaption-FineGrained.formosa-vision-finegrained
Formosa Vision Fine-grained (Expanded)
Dataset Summary
此資料集以台灣在地文化與地景為核心,提供具細節的中文描述,並保留原始圖像。
擴充版本針對每張圖像生成更長、更密集的語義描述,以強化模型在細節理解上的表現。
Motivation
『資料合成』FLAIR 的核心在於訓練模型「聽得懂細節」。這意味著「長文本」越具體、包含越多方位詞 (左上角、紅色物體旁...),模型學到的局部特徵就越好。因為在此階段會透過大型多模態模型生成豐富且長的中文描述夠「碎唸」(包含大量方位、顏色、材質等細節)。相較於網路爬蟲數據,此資料庫具備高品質的本土文化實體 (Entity) 標註,是訓練台灣在地化 AI 的最佳基石。
Source Data
原始資料集:twinkle-ai/Formosa-Vision(Hugging Face Datasets)
擴充流程:以本地 VLM 產生更細緻的中文長描述
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/renhehuang/formosa-vision-finegrained.FiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/CiaranCw/FiVE-Fine-Grained-Video-Editing-Benchmark.fine_grained_cupreplanner_fine_grainedreplanner_fine_grained2replanner_fine_grained_full_3_allfinegrained_vehicle_labelscar-bdd-fine-grained
Fine-Grained Vehicle Detection Dataset (Corolla × BMW 3-Series, BDD100K-derived)
Real-world dashcam frames with fine-grained make annotations on top of
BDD100K's existing car bounding boxes. Identifies which BDD-labeled cars
are specifically Toyota Corolla sedans or BMW 3-Series sedans.
Images: 1410 unique frames · Annotations: 1524
(558 Corolla + 966 BMW 3-Series)
Source imagery: BDD100K dashcam corpus (Berkeley DeepDrive)
Why this dataset
BDD100K labels every car as… See the full description on the dataset page: https://huggingface.co/datasets/arrmlet/car-bdd-fine-grained.fine_grained_hammer_maskonly_gazeCoT-SFT-Fine-grained-part1Classification_4shot_Fine_Grainedreplanner_fine_grained3replanner_fine_grained_full_allreplanner_fine_grained_full3replanner_fine_grained_full
