fine-grained
FUSU-Fine_grained_Urban_Semantic_Understanding
About:
FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection and domain adaptation tasks. This data comprises 2 parts:
Bi-temporal high-resolution satellite RGB images with fine-grained annotations.
Monthly revisited Sentinel-2 and Sentinel-1 images.
Details:
1.… See the full description on the dataset page: https://huggingface.co/datasets/sp-juni/FUSU-Fine_grained_Urban_Semantic_Understanding.FiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/LIMinghan/FiVE-Fine-Grained-Video-Editing-Benchmark.cantonese-drama-voice-fine-grained语料集名称:面向影视剧AI配音的粤语语料库
语料来源:AI DimSum Lab
简介:
本语料库是专为粤语影视剧 AI 配音模型训练构建的专用语料资源,核心涵盖《神雕侠侣(1995 古天乐版)》《乘龙怪婿》《寻秦记》等经典粤语影视内容。语料库匹配 AI 配音模型的人物识别、语音情绪识别、语音生成三大模块需求,并根据下游任务对《神雕侠侣》等语音数据进行了多情感、多人物、文本标注。其数据规模总计约800MB,影视时长超 7 小时,可提供丰富的粤语语音、情感、人物关联样本,能有效支撑模型训练中人物区分、情感还原、语音生成的精度提升,是粤语影视剧 AI 配音落地的核心数据基础。
适用场景:
粤语影视剧 AI 配音模型训练:直接用于模型的人物区分、情感还原、语音生成模块优化;
粤语语音研究:可作为粤语语音特征、情感语音分析的基础数据集;
影视 AI 技术开发:为影视领域的语音合成、角色语音克隆等技术提供数据支持。
使用说明:
本语料库仅用于非商业研究与技术开发(商业使用需联系维护者确认授权);
使用前建议对语音数据进行预处理(如降噪、采样率统一),以提升模型训练效果。
fine-grained-medical-reasoning
Dataset Card for Fine-Grained Medical Reasoning
Fine-grained medical reasoning QA dataset introduced in "Can LLMs Reason Like Doctors? Exploring the Limits of Large Language Models in Complex Medical Reasoning"
(Findings of EACL 2026). Manually annotated from the MedAgentsBench test_hard set,
it evaluates LLMs’ abduction, deduction, and induction capabilities, offering detailed insights into physician-like reasoning.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/expertailab/fine-grained-medical-reasoning.Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes
Dataset Card for Dataset Name
The Amazon reviews full score dataset is constructed by randomly taking 600,000 training samples and 130,000 testing samples for each review score from 1 to 5. In total there are 3,000,000 trainig samples and 650,000 testing samples.
Dataset Details
Dataset Description
The files train.csv and test.csv contain all the training samples as comma-sparated values. There are 3 columns in them, corresponding to class index (1 to 5)… See the full description on the dataset page: https://huggingface.co/datasets/yassiracharki/Amazon_Reviews_for_Sentiment_Analysis_fine_grained_5_classes.test-fine-grained-challenges
Fine-Grained Challenges
Groups of morphologically similar species for evaluating and tuning fine-grained classifiers on top of BioCLIP ecosystem model embeddings. Each group gathers species that are easily confused with one another, and the groups span several clades so a classifier can be probed on the distinctions that actually matter rather than on coarse taxonomy.
The corpus lives in a Lance dataset and can be acted on as a whole, on any single group independently, or on any… See the full description on the dataset page: https://huggingface.co/datasets/thompsonmj/test-fine-grained-challenges.
