damo
Datasets
All datasets matching “damo”multimodal_textbook
Multimodal-Textbook-6.5M
Overview
This dataset is for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining", containing 6.5M images interleaving with 0.8B text from instructional videos.
It contains pre-training corpus using interleaved image-text format. Specifically, our multimodal-textbook includes 6.5M keyframesextracted from instructional videos, interleaving with 0.8B ASR texts.
All the images and text are extracted from online… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/multimodal_textbook.C_damoxingVideoRefer-700K
VideoRefer-700K
Paper | Project Page | Code
VideoRefer-700K is a large-scale, high-quality object-level video instruction dataset. Curated using a sophisticated multi-agent data engine to fill the gap for high-quality object-level video instruction data.
VideoRefer consists of three types of data:
Object-level Detailed Caption
Object-level Short Caption
Object-level QA
Video sources:
Detailed&Short Caption
Panda-70M.
QA
MeViS
A2D
Youtube-VOS
Data format:
[
{… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-700K.LIb_damoxingClinFusion-Eval-Data
🏥 ClinFusion-Eval-Data
The Holistic Evaluation Suite for Vision-Centric Medical Multimodal LLMs
ClinFusion-Eval-Data is the unified evaluation corpus used to benchmark the ClinFusion model series (ClinFusion-8B, ClinFusion-32B). It packages 211,810 evaluation records spanning 22 public medical benchmarks into a single, consistently-formatted suite, together with 509 GiB of the underlying 2D images and native 3D CT volumes they refer to.
The goal is reproducibility:… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinFusion-Eval-Data.LIb_damoxing2
