datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal_textbook
Multimodal-Textbook-6.5M
Overview
This dataset is for "2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining", containing 6.5M images interleaving with 0.8B text from instructional videos.
It contains pre-training corpus using interleaved image-text format. Specifically, our multimodal-textbook includes 6.5M keyframesextracted from instructional videos, interleaving with 0.8B ASR texts.
All the images and text are extracted from online… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/multimodal_textbook.C_damoxingVideoRefer-700K
VideoRefer-700K
Paper | Project Page | Code
VideoRefer-700K is a large-scale, high-quality object-level video instruction dataset. Curated using a sophisticated multi-agent data engine to fill the gap for high-quality object-level video instruction data.
VideoRefer consists of three types of data:
Object-level Detailed Caption
Object-level Short Caption
Object-level QA
Video sources:
Detailed&Short Caption
Panda-70M.
QA
MeViS
A2D
Youtube-VOS
Data format:
[
{… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/VideoRefer-700K.LIb_damoxingClinFusion-Eval-Data
🏥 ClinFusion-Eval-Data
The Holistic Evaluation Suite for Vision-Centric Medical Multimodal LLMs
ClinFusion-Eval-Data is the unified evaluation corpus used to benchmark the ClinFusion model series (ClinFusion-8B, ClinFusion-32B). It packages 211,810 evaluation records spanning 22 public medical benchmarks into a single, consistently-formatted suite, together with 509 GiB of the underlying 2D images and native 3D CT volumes they refer to.
The goal is reproducibility:… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinFusion-Eval-Data.LIb_damoxing2MultiJail
Multilingual Jailbreak Challenges in Large Language Models
This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models".
[Github repo]
Annotation Statistics
We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below:
High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi)
Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.RynnBrain-Bench
RynnBrain-Bench
Introduction
We introduce RynnBrain-Bench, a high-dimensional evaluation suite designed to holistically benchmark the cognition and localization capabilities of embodied understanding models in complex household environments.
Advancing beyond existing benchmarks, RynnBrain-Bench features a unique emphasis on fine-grained understanding and precise spatiotemporal localization within episodic video sequences.
RynnBrain-Bench systematically… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnBrain-Bench.multialpacaoag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.Qwen2.5-7B-LongPO-128K-tokenizedMulti-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.LMA-Individual-ProjectInterVBench
Video Drift Evaluation (vde.py)
This repository contains a single entry point, vde.py, that computes Video Drift Error (VDE) scores for every .mp4 file inside a target directory. VDE provides a simple way to monitor how quality-related metrics drift across chunks of the same video. The script already supports several metric backends (clarity, motion, aesthetic, dynamic, subject, background) via the vbench tooling.
Environment Setup
Install the project… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/InterVBench.Mistral-7B-LongPO-256K-tokenizedDamogranLeftCup_20260911_120558This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranLeftCup_20260911_120558.RynnEC-Bench
RynnEC-Bench
RynnEC-Bench evaluates fine-grained embodied understanding models from the perspectives of object cognition and spatial cognition in open-world scenario. The benchmark includes 507 video clips captured in real household scenarios.
Model
Overall Mean
Object Properties
Seg. DR
Seg. SR
Object Mean
Ego. His.
Ego. Pres.
Ego. Fut.
World Size
World Dis.
World PR
Spatial Mean
GPT-4o
28.3
41.1
---
---
33.9
13.4
22.8
6.0
24.3
16.7
36.1
22.2
GPT-4.1
33.5… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnEC-Bench.DamogranMiddleCup_20260911_124118This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranMiddleCup_20260911_124118.DamogranRightCup_20260911_140547This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranRightCup_20260911_140547.SpanUQ-Benchmark
SpanUQ Benchmark
A span-level uncertainty estimation benchmark for large language model generation. Each example contains an LLM-generated response decomposed into spans (contiguous text segments expressing single verifiable assertions), with uncertainty labels derived from sampling-based consistency verification.
Quick Start
from datasets import load_dataset
# Load a specific model configuration
ds = load_dataset("DamonDemon/SpanUQ-Benchmark", "Qwen3-14B")… See the full description on the dataset page: https://huggingface.co/datasets/DamonDemon/SpanUQ-Benchmark.DamogranRightCup2_20260911_142724This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranRightCup2_20260911_142724.Mistral-7B-LongPO-128K-tokenizedDamogranLeftCup_20260911_122012This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranLeftCup_20260911_122012.DamogranRightCup_20260911_130855This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranRightCup_20260911_130855.DamogranRightCup_20260911_140000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranRightCup_20260911_140000.DamogranRightCup_20260911_130618This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranRightCup_20260911_130618.DamogranRightCup_20260911_140335This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranRightCup_20260911_140335.DamogranMiddleCup_20260911_130256This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranMiddleCup_20260911_130256.DamogranRightCup_20260911_142409This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/rbalazs/DamogranRightCup_20260911_142409.ClinHallu
CLINHALLU Benchmark
CLINHALLU is a benchmark for diagnosing stage-wise hallucinations in medical MLLM reasoning.
Paper: CLINHALLU: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM ReasoningGitHub: alibaba-damo-academy/ClinHallu
Benchmark Results
Accuracy and stage-wise hallucination rates on CLINHALLU. We report answer accuracy (Acc) and hallucination rates for visual recognition (H^V), knowledge recall (H^K), and reasoning integration (H^R).… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinHallu.
