datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MultiJail
Multilingual Jailbreak Challenges in Large Language Models
This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models".
[Github repo]
Annotation Statistics
We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below:
High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi)
Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/MultiJail.RynnBrain-Bench
RynnBrain-Bench
Introduction
We introduce RynnBrain-Bench, a high-dimensional evaluation suite designed to holistically benchmark the cognition and localization capabilities of embodied understanding models in complex household environments.
Advancing beyond existing benchmarks, RynnBrain-Bench features a unique emphasis on fine-grained understanding and precise spatiotemporal localization within episodic video sequences.
RynnBrain-Bench systematically… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/RynnBrain-Bench.multialpacaoag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.LMA-Individual-ProjectSpanUQ-Benchmark
SpanUQ Benchmark
A span-level uncertainty estimation benchmark for large language model generation. Each example contains an LLM-generated response decomposed into spans (contiguous text segments expressing single verifiable assertions), with uncertainty labels derived from sampling-based consistency verification.
Quick Start
from datasets import load_dataset
# Load a specific model configuration
ds = load_dataset("DamonDemon/SpanUQ-Benchmark", "Qwen3-14B")… See the full description on the dataset page: https://huggingface.co/datasets/DamonDemon/SpanUQ-Benchmark.ClinHallu
CLINHALLU Benchmark
CLINHALLU is a benchmark for diagnosing stage-wise hallucinations in medical MLLM reasoning.
Paper: CLINHALLU: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM ReasoningGitHub: alibaba-damo-academy/ClinHallu
Benchmark Results
Accuracy and stage-wise hallucination rates on CLINHALLU. We report answer accuracy (Acc) and hallucination rates for visual recognition (H^V), knowledge recall (H^K), and reasoning integration (H^R).… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-DAMO-Academy/ClinHallu.DAMO-MultiJailThe dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license.
ciaa-annual-reports
CIAA Annual Reports — Nepali transcripts, ruled tables and chart data
Machine-readable transcripts of the annual reports of Nepal's Commission for the
Investigation of Abuse of Authority (अख्तियार दुरुपयोग अनुसन्धान आयोग, CIAA) —
all 35 it has published to date. The 1st to 35th reports, fiscal years
BS 2047/48 – 2081/82 (AD 1990–2025).
The CIAA publishes these as PDFs whose text layer is, for several years, legacy
pre-Unicode Devanagari that ordinary extractors turn into… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/ciaa-annual-reports.NLP-A2VL3-Syn7M
The re-caption dataset used in VideoLLaMA 3: Frontier Multimodal Foundation Models for Video Understanding
If you like our project, please give us a star ⭐ on Github for the latest update.
🌟 Introduction
This dataset is the re-captioned data we used during the training of VideoLLaMA3. It consists of 7 million diverse, high-quality images, each accompanied by a short caption and a detailed caption.
The images in this dataset originate from COYO-700M, MS-COCO 2017… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/VL3-Syn7M.SOULThis repo contains the data for our paper "SOUL: Towards Sentiment and Opinion Understanding of Language" in EMNLP 2023.
Github repo
Statistics
The SOUL dataset comprises 15,028 statements related to 3,638 reviews, resulting in an average of 4.13 statements per review. To create training, development, and test sets, we split the reviews in a ratio of 6:1:3, respectively.
Split
# reviews
# statements
True
False
Not-given
Total
Train
2,182
3,675
2,159
8,834
3,000
8,834… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/SOUL.Agent1oasst1-guanaco-damo-convai-pro
Dataset Card for "oasst1-guanaco-damo-convai-pro"
More Information needed
matterjslegal-retrieval-decisionDeidentification-of-EHRde_shop_api_v3softprompt0ingautotrain-datasetv4_100k_processedparler_spark_trainDamorkDataSetdata.csv
Nanbeige4-3B Base Model Blind Spot Dataset
Model Tested
Nanbeige/Nanbeige4-3B-Base
https://huggingface.co/Nanbeige/Nanbeige4-3B-Base
This dataset documents examples where the model produces incorrect predictions.
Dataset Structure
column
description
input
prompt given to the model
expected_output
correct output
model_output
output generated by the model
error_type
category of error
How the Model Was Loaded
from transformers… See the full description on the dataset page: https://huggingface.co/datasets/DamolaRachael/data.csv.
