datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visually-dependent-ambiguity
VIDA: Visually-Dependent Ambiguity for Multimodal MT
VIDA is an English-Chinese multimodal machine translation dataset for visual ambiguity resolution.Each instance contains an English source sentence, its paired image, and Chinese references that resolve annotated ambiguity spans using visual evidence.
Paper: A Multimodal Dataset for Visually Grounded Ambiguity in Machine Translation
Dataset composition
This release contains four splits:
Split
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/p1k0/visually-dependent-ambiguity.ambiguity-filtered-mm-safetybench
