datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RynnBrain-Bench
RynnBrain-Bench
Introduction
We introduce RynnBrain-Bench, a high-dimensional evaluation suite designed to holistically benchmark the cognition and localization capabilities of embodied understanding models in complex household environments.
Advancing beyond existing benchmarks, RynnBrain-Bench features a unique emphasis on fine-grained understanding and precise spatiotemporal localization within episodic video sequences.
RynnBrain-Bench systematically… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnBrain-Bench.multialpacaMulti-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.ClinHallu
CLINHALLU Benchmark
CLINHALLU is a benchmark for diagnosing stage-wise hallucinations in medical MLLM reasoning.
Paper: CLINHALLU: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM ReasoningGitHub: alibaba-damo-academy/ClinHallu
Benchmark Results
Accuracy and stage-wise hallucination rates on CLINHALLU. We report answer accuracy (Acc) and hallucination rates for visual recognition (H^V), knowledge recall (H^K), and reasoning integration (H^R).… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinHallu.VL3-Syn7M
The re-caption dataset used in VideoLLaMA 3: Frontier Multimodal Foundation Models for Video Understanding
If you like our project, please give us a star ⭐ on Github for the latest update.
🌟 Introduction
This dataset is the re-captioned data we used during the training of VideoLLaMA3. It consists of 7 million diverse, high-quality images, each accompanied by a short caption and a detailed caption.
The images in this dataset originate from COYO-700M, MS-COCO 2017… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/VL3-Syn7M.legal-retrieval-decisionDeidentification-of-EHRde_shop_api_v3softprompt0ingDamorkDataSet
