datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WebLINX-full
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WARNING: This is not the main WebLINX data card! You might want to use the main WebLINX data card instead:
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue
Xing Han Lù*, Zdeněk Kasner*, Siva Reddy
💾Code
📄Paper
🌐Website
📓Colab
🤖Models
💻Explorer
🐦Tweets
🏆Leaderboard
Your browser does not support the… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/WebLINX-full.agent-reward-bench
AgentRewardBench
💾Code
📄Paper
🌐Website
🤗Dataset
💻Demo
🏆Leaderboard
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor
Loading dataset
You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.CharXiv
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
NeurIPS 2024
🏠Home (🚧Still in construction) | 🤗Data | 🥇Leaderboard | 🖥️Code | 📄Paper
This repo contains the full dataset for our paper CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs, which is a diverse and challenging chart understanding benchmark fully curated by human experts. It includes 2,323 high-resolution charts manually sourced from arXiv preprints. Each chart is… See the full description on the dataset page: https://huggingface.co/datasets/princeton-nlp/CharXiv.flickr30k
Flickr30k
Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Homepage: https://shannon.cs.illinois.edu/DenotationGraph/
Bibtex:
@article{young2014image,
title={From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions},
author={Young, Peter and Lai, Alice and Hodosh, Micah and Hockenmaier, Julia},
journal={Transactions of the Association… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/flickr30k.Omnimodal-Agent-SFT-2K
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
This dataset contains omni-modal agent supervised fine-tuning (SFT) trajectories in the LlamaFactory SFT data format. You can directly follow LlamaFactory's instructions to fine-tune your omni-modal LLMs.OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/Omnimodal-Agent-SFT-2K.MedThinkVQA
MedThinkVQA
MedThinkVQA is an expert-annotated benchmark for multi-image diagnostic reasoning in radiology. Unlike prior medical VQA benchmarks that typically contain at most one image per case, MedThinkVQA requires models to extract evidence from each image, integrate cross-view information, and perform differential-diagnosis reasoning.
Links
GitHub: https://github.com/benluwang/MedThinkVQA
Leaderboard: https://benluwang.github.io/MedThinkVQA/
Submission Guide:… See the full description on the dataset page: https://huggingface.co/datasets/bio-nlp-umass/MedThinkVQA.mscoco_2014_5k_test_image_text_retrieval
MSCOCO (5K test set)
Original paper: Microsoft COCO: Common Objects in Context
Homepage: https://cocodataset.org/#home
5K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@inproceedings{lin2014microsoft,
title={Microsoft coco: Common objects in context},
author={Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Doll{\'a}r, Piotr and Zitnick, C Lawrence}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/mscoco_2014_5k_test_image_text_retrieval.flickr_1k_test_image_text_retrieval
Flickr30k (1K test set)
Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Homepage: https://shannon.cs.illinois.edu/DenotationGraph/
1K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@article{young2014image,
title={From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/flickr_1k_test_image_text_retrieval.Design2CodeThis dataset consists of 484 webpages from the C4 validation set, serving the purpose of testing multimodal LLMs on converting visual designs into code implementations.
Each example is a pair of source HTML and screenshot ({id}.html and {id}.png).
See the dataset in the huggingface format here.
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
Example Usage
For example, you… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/Design2Code.OmniGAIA
OmniGAIA: Omni-Modal General AI Assistant Benchmark
📄 Paper
•
💻 Code & Demo
•
🤗 Dataset & Model
•
📈 Leaderboard
OmniGAIA is a benchmark for Omni-Modal General AI Assistants that jointly reason over vision, audio, and language with external tools. It is designed to evaluate long-horizon, multi-hop, open-form problem solving in realistic settings rather than short perception-only QA.
Benchmark Construction
The OmniGAIA construction… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/OmniGAIA.whoops
Dataset Card for WHOOPS!
Dataset Description
Contribute Images to Extend WHOOPS!
Languages
Dataset
Data Fields
Data Splits
Data Loading
Licensing Information
Annotations
Considerations for Using the Data
Citation Information
Dataset Description
WHOOPS! is a dataset and benchmark for visual commonsense. The dataset is comprised of purposefully commonsense-defying images created by designers using publicly-available image generation tools like Midjourney. It contains… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/whoops.paradetox
ParaDetox: Text Detoxification with Parallel Data (English)
This repository contains information about ParaDetox dataset -- the first parallel corpus for the detoxification task -- as well as models and evaluation methodology for the detoxification of English texts. The original paper "ParaDetox: Detoxification with Parallel Data" was presented at ACL 2022 main conference.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/paradetox.AURORA
Read the paper here: https://arxiv.org/abs/2407.03471. IMPORTANT: Please check out our GitHub repository for more instructions on how to also access the Something-Something-Edit subdataset, which we can't publish directly: https://github.com/McGill-NLP/AURORA
MMc-Instruct-Stage2Sketch2CodeThe Sketch2Code dataset consists of 731 human-drawn sketches paired with 484 real-world webpages from the Design2Code dataset, serving to benchmark Vision-Language Models (VLMs) on converting rudimentary sketches into web design prototypes.
Each example consists of a pair of source HTML and rendered webpage screenshot (stored in webpages/ directory under name {webpage_id}.html and {webpage_id}.png), as well as 1 to 3 sketches drawn by human annotators (stored in sketches/ directory under name… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/Sketch2Code.LLaVAR
LLaVAR Data: Enhanced Visual Instruction Data with Text-Rich Images
More info at LLaVAR project page, Github repo, and paper.
Training Data
Based on the LAION dataset, we collect 422K pretraining data based on OCR results. For finetuning data, we collect 16K high-quality instruction-following data by interacting with langauge-only GPT-4. Note that we also release a larger and more diverse finetuning dataset below (20K), which contains the 16K we used for the paper. The… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/LLaVAR.MultiExpArt
Dataset Card for Multilingual Explain Artworks: MultiExpArt
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Description
Dataset Summary
As the performance of Large-scale Vision Language Models (LVLMs) improves, they are increasingly capable of responding in multiple languages, and there is an expectation that the demand for explanations generated by LVLMs will grow. However, pre-training… See the full description on the dataset page: https://huggingface.co/datasets/naist-nlp/MultiExpArt.Design2Code-HARDThis dataset consists of 80 extra difficult webpages from Github Pages, which challenges SoTA multimodal LLMs on converting visual designs into code implementations.
Each example is a pair of source HTML and screenshot ({id}.html and {id}.png).
See the "easy" version of the Design2Code testset here
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
Design2Code-hfThis dataset consists of 484 webpages from the C4 validation set, serving the purpose of testing multimodal LLMs on converting visual designs into code implementations.
See the dataset in the raw files format here.
Note that all images in these webpages are replaced by a placeholder image (rick.jpg)
Please refer to our project page and our paper for more information.
TOMATO
🍅 TOMATO
📄 Paper | 💻 Code | 🎬 Videos
This repository contains the QAs of the following paper:
🍅 TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
Ziyao Shangguan*1,
Chuhan Li*1,
Yuxuan Ding1,
Yanan Zheng1,
Yilun Zhao1,
Tesca Fitzgerald1,
Arman Cohan12
*Equal contribution.
1Yale University 2Allen Institute of AI
TOMATO - A Visual Temporal Reasoning Benchmark
Introduction
Our study of existing benchmarks… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/TOMATO.ChinaHeritaQA
Images
This folder contains visual data for the ChinaHeritaQA benchmark: https://arxiv.org/abs/2606.08959
Contents
Folder
Description
Image_data/
Chinese UNESCO World Heritage Site images (2,279 images from 51 sites)
worlds_data/
Non-Chinese World Heritage Site images (133 images from 23 sites)
Overview
The image dataset includes a comprehensive collection of photographs from both Chinese and international UNESCO World Heritage… See the full description on the dataset page: https://huggingface.co/datasets/Multilingual-NLP/ChinaHeritaQA.VQA-princeton-nlp-CharXiv-clean
Description
French translation of the princeton-nlp/CharXiv dataset that we processed.
Citation
@article{wang2024charxiv,
title={CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs},
author={Wang, Zirui and Xia, Mengzhou and He, Luxi and Chen, Howard and Liu, Yitao and Zhu, Richard and Liang, Kaiqu and Wu, Xindi and Liu, Haotian and Malladi, Sadhika and Chevalier, Alexis and Arora, Sanjeev and Chen, Danqi},
journal={arXiv preprint… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/VQA-princeton-nlp-CharXiv-clean.neurips_glotocr
GlotOCR-bench
GlotOCR-bench is a dataset of 16375 images covering 158 writing systems (+2000 languages), designed to evaluate the fundamental OCR capabilities required to support diverse writing systems and languages.
Quick links:
🏆 Leaderboard
📝 License
ru_paradetox
ParaDetox: Text Detoxification with Parallel Data (Russian)
This repository contains information about Russian Paradetox dataset -- the first parallel corpus for the detoxification task -- as well as models for the detoxification of Russian texts.
📰 Updates
[2025] !!!NOW OPEN!!! TextDetox CLEF2025 shared task: for even more -- 15 languages! website 🤗Starter Kit
[2025] COLNG2025: Daryna Dementieva, Nikolay Babakov, Amit Ronen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider… See the full description on the dataset page: https://huggingface.co/datasets/s-nlp/ru_paradetox.K-prism
K-Prism
Korean diagnostic benchmark data for evaluating hallucination in vision-language
models. Evaluation code and protocol documentation are available at
alsgur0720/K-Prism.
Files
File
Contents
Text_track.json
504 text-track questions
Image_track.json
498 image-track questions
images/
165 original images referenced by the image track
The release contains 1,002 questions and approximately 303 MB of annotations and
images. Keep both JSON files… See the full description on the dataset page: https://huggingface.co/datasets/KETI-NLP/K-prism.M3SciQA
🧑🔬 M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark For Evaluating Foundatio Models
EMNLP 2024 Findings
🖥️ Code
Introduction
In the realm of foundation models for scientific research, current benchmarks predominantly focus on single-document, text-only tasks and fail to adequately represent the complex workflow of such research. These benchmarks lack the $\textit{multi-modal}$, $\textit{multi-document}$ nature of scientific research, where comprehension… See the full description on the dataset page: https://huggingface.co/datasets/yale-nlp/M3SciQA.retriever-princeton-nlp-CharXiv-clean
Description
princeton-nlp/CharXiv dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@article{wang2024charxiv,
title={CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs},
author={Wang, Zirui and Xia, Mengzhou and He, Luxi and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-princeton-nlp-CharXiv-clean.KnowRecall
KnowRecall
This repository contains the KnowRecall benchmark, introduced in Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs.
Dataset Description
Imagine a French tourist visiting Tokyo Tower, snapping a photo and asking an MLLM about the tower’s height.
Naturally, they would expect a correct response in their native language.
However, if the model provides the right answer in Japanese but fails to do so in French, it… See the full description on the dataset page: https://huggingface.co/datasets/nlp-waseda/KnowRecall.GOAT-Bench
The GOAT Benchmark (HomePage)
We introduce the GOAT-Bench, a comprehensive and specialized dataset designed to evaluate large multimodal models through meme-based multimodal social abuse. GOAT-Bench comprises over 6K diverse memes, encompassing a range of themes including hate speech and offensive content. Our focus is to assess the ability of LMMs to accurately identify online abuse, specifically in terms of hatefulness, misogyny, offensiveness, sarcasm, and harmfulness. We… See the full description on the dataset page: https://huggingface.co/datasets/HKBU-NLP/GOAT-Bench.nlpdlvladtst
