datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jat-dataset
JAT Dataset
Dataset Description
The Jack of All Trades (JAT) dataset combines a wide range of individual datasets. It includes expert demonstrations by expert RL agents, image and caption pairs, textual data and more. The JAT dataset is part of the JAT project, which aims to build a multimodal generalist agent.
Paper: https://huggingface.co/papers/2402.09844
Usage
>>> from datasets import load_dataset
>>> dataset =… See the full description on the dataset page: https://huggingface.co/datasets/jat-project/jat-dataset.EarthDial-Dataset
🌍 EarthDial-Dataset
The EarthDial-Dataset is a curated collection of evaluation-only datasets focused on remote sensing and Earth observation downstream tasks. It is designed to benchmark vision-language models (VLMs) and multimodal reasoning systems on real-world scenarios involving satellite and aerial imagery.
📚 Key Features
Evaluation-focused: All datasets are for inference/testing only — no train/val splits.
Diverse Tasks:
Classification
Object Detection
Change… See the full description on the dataset page: https://huggingface.co/datasets/akshaydudhane/EarthDial-Dataset.TransCity-VLM-dataset
TransCity-VLM Dataset
The TransCity-VLM Dataset provides multimodal smart-city data for traffic, energy, mobility, grid operation, urban context understanding, and map-grounded question answering. It supports the training and evaluation of vision-language models for urban prediction, decision support, conversational QA, and reasoning tasks.
The training data are available at this Hugging Face dataset repository.
Dataset Summary
Split
Rows / Files
test JSONL… See the full description on the dataset page: https://huggingface.co/datasets/TransCity-VLM/TransCity-VLM-dataset.STRIDE-QA-Dataset-Mini
STRIDE-QA-Dataset-Mini
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
⚠️ Note: STRIDE-QA-Dataset-Mini is provided as a preliminary version and does not fully match the format of the… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset-Mini.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.CaReCoS
CaReCoS
A medical acoustic question-answering dataset for reasoning over mel spectrograms
of heart, lung, and cough sounds. Each record provides a clinical question, the
mel-spectrogram image of a recording, a ground-truth answer, and the
recording's clinical metadata.
The task is purely visual: a model receives the spectrogram image together with the
question and must reason over the spectrogram to produce the answer. The raw audio is
not used as model input - the original .wav… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-dataset-1/CaReCoS.Inst-It-Dataset
Inst-IT Dataset: An Instruction Tuning Dataset with Multi-level Fine-Grained Annotations
introduced in the paper Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning
🌐 Homepage | Code | 🤗 Paper | 📖 arXiv
Inst-IT Dataset Overview
We create a large-scale instruction tuning dataset, the Inst-it Dataset. To the best of our knowledge, this is the first dataset that provides fine-grained annotations centric on specific… See the full description on the dataset page: https://huggingface.co/datasets/Inst-IT/Inst-It-Dataset.ContPhy_Dataset
ContPhy Dataset Repository
ContPhy: Continuum Physical Concept Learning and Reasoning from Videos
Zhicheng Zheng*, Xin Yan*, Zhenfang Chen*, Jingzhou Wang, Qin Zhi Eddie Lim, Joshua B. Tenenbaum, and Chuang Gan (* denotes equal contributions)
ICML 2024Links | Project Page | Paper (Arxiv) | Codebase | Cite ContPhy
Structure
Mini Dataset
contphy_mini.zip: 20 videos each scenario with annotations
Full Dataset… See the full description on the dataset page: https://huggingface.co/datasets/zzcnewly/ContPhy_Dataset.vecombot-dataset
VECOM — Bộ dữ liệu thị trường Thương mại điện tử Việt Nam (VEComBot)
Bộ dữ liệu phụ lục cho đồ án tốt nghiệp VEComBot — hệ thống Đa tác tử (Multi-Agent
System) phân tích và tổng hợp thị trường Thương mại điện tử Việt Nam (VECOM). Đây là
kho tài liệu nguồn và corpus đã qua xử lý (figure-aware) được nạp vào PostgreSQL/pgvector
để phục vụ cả nhánh MAS lẫn nhánh baseline naive RAG.
Mục đích: dùng cho nghiên cứu học thuật và tái lập kết quả đồ án. Các báo cáo gốc là
ấn phẩm công… See the full description on the dataset page: https://huggingface.co/datasets/binhtran23/vecombot-dataset.PaperAudit_Dataset
PaperAudit Origin Data
This directory contains the original paper data downloaded and preprocessed for the PaperAudit project. The data includes papers from top-tier machine learning conferences with their parsed content, metadata, synthetic error annotations, and review information.
PaperAudit Dataset Overview
This repository is part of the full PaperAudit Dataset, which includes:
PaperAudit_Dataset/
├── PaperAudit_Origin_Data/ # Original paper data (raw + preprocessed)… See the full description on the dataset page: https://huggingface.co/datasets/mayiwen/PaperAudit_Dataset.atspm-dataset
ATSPM QA Dataset
Dataset Description
The ATSPM QA Dataset is a collection of synthetic question-and-answer pairs designed to train and evaluate large language models on their ability to interpret and analyze Automated Traffic Signal Performance Measures (ATSPM) charts. This dataset is a crucial component for developing agentic AI systems that can automate the analysis of traffic signal data.
The data was generated synthetically by posing questions to a large language… See the full description on the dataset page: https://huggingface.co/datasets/grhone/atspm-dataset.Turkish-medical-visual-question-answering-LLaVa-dataset
Türkçe Radyoloji Görüntüleme Veri Seti - data_RAD
data_RAD veri seti, radyoloji görüntüleri üzerinde görsel soru-cevaplama (VQA) araştırmaları yapmak amacıyla Türkçeye çevrilmiş ve LLaVa mimarisiyle uyumlu hale getirilmiştir. Bu veri seti, tıbbi görüntü analizi ve yapay zeka destekli radyoloji uygulamalarını geliştirmek için kullanılabilir.
Veri Seti İçeriği
Toplam Görüntü Sayısı: 316
Veri Yapısı: DatasetDict({ train: Dataset({ features: ['image'], num_rows: 316 }) })
Özellikler:… See the full description on the dataset page: https://huggingface.co/datasets/nezahatkorkmaz/Turkish-medical-visual-question-answering-LLaVa-dataset.universal-preference-hijacking-datasets
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image.
Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images.
This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.VGAVS-Datasetbase_paper:
arxiv
MedMASLab_dataset
MedMASLab Dataset
Paper | GitHub
📋 Overview
MedMASLab is the unified, comprehensive benchmarking platform specifically designed for medical vision-language multi-agent systems. It addresses critical challenges in the medical AI field by providing standardized infrastructure, rigorous evaluation metrics, and extensive empirical insights.
Dataset Summary
MedMASLab provides the most extensive benchmark to date for medical vision-language agents, standardizing… See the full description on the dataset page: https://huggingface.co/datasets/qyhhhhh/MedMASLab_dataset.tubitak-olimpiyat-dataset
TUBITAK Science Olympiad Dataset
This dataset contains multiple-choice and open-ended scientific questions sourced from the TUBITAK (The Scientific and Technological Research Council of Turkey) Science Olympiads spanning various years. It is intended to serve as a benchmark for evaluating the advanced analytical, mathematical, and computational reasoning capabilities of Large Language Models (LLMs) in the Turkish language.
The dataset comprises approximately 2700 problems across… See the full description on the dataset page: https://huggingface.co/datasets/sghosts/tubitak-olimpiyat-dataset.Chinese-Image-Text-Corpus-dataset
REILX/Chinese-Image-Text-Corpus-dataset
[ English | 中文 ]
Introduction
The REILX/Chinese-Image-Text-Corpus-dataset is a multimodal dataset that pairs Chinese textual data with corresponding images. This dataset is derived from the Chinese-Xinhua Dictionary Database, which includes idioms, single characters, words, and aphorisms.
Dataset Structure
The dataset is organized into the following categories:
Idioms: Traditional Chinese idioms with explanations and… See the full description on the dataset page: https://huggingface.co/datasets/REILX/Chinese-Image-Text-Corpus-dataset.Flickr-Dataset-5k
Flickr1k
This dataset is a subset of the Original Flickr30k Dataset, containing [Total number of samples, e.g., 5000] image-caption pairs. It has been specifically created for [Briefly state the purpose, e.g., faster experimentation with image captioning models or a specific research focus].
Dataset Details
Original Dataset: Original Flickr30k dataset on Hugging Face Hub
Subset Size: 5000
Data Format: Each example contains the following fields:
image: The raw bytes of… See the full description on the dataset page: https://huggingface.co/datasets/Vishva007/Flickr-Dataset-5k.Deaftest_datasetOfficial Deaftest dataset for the paper "AV-Odyssey: Can Your Multimodal LLMs Really Understand Audio-Visual Information?".
🌟 For more details, please refer to the project page with data examples: https://av-odyssey.github.io/.
[🌐 Webpage] [📖 Paper] [🤗 Huggingface AV-Odyssey Dataset] [🤗 Huggingface Deaftest Dataset] [🏆 Leaderboard]
🔥 News
2024.11.24 🌟 We release AV-Odyssey, the first-ever comprehensive evaluation benchmark to explore whether MLLMs really understand… See the full description on the dataset page: https://huggingface.co/datasets/AV-Odyssey/Deaftest_dataset.thai_famous_people_images_dataset
Thai Famous People Image Dataset
Dataset Description
The Thai Famous People Image Dataset is a collection of images and descriptions of famous Thai personalities. This dataset is designed to provide a comprehensive resource for researchers, developers, and enthusiasts interested in Thai culture, history, and notable figures. The data was extracted from the Thai Wikipedia dump in September 2024, ensuring up-to-date and relevant information.
Maintainer
Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/iapp/thai_famous_people_images_dataset.myawady-raw-dataset
Myawady Raw News Corpus 🇲🇲
This dataset contains over 59,000 full-text Burmese news articles scraped from the Myawady News Portal, the official media outlet of the Myanmar military government.
Unlike the title-only version, this dataset includes complete article content, with metadata fields such as category, publication date, and image URLs. It is intended for use in Myanmar NLP and AI research, including:
🧠 Language modeling
📰 Text summarization
🏷️ Named entity… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myawady-raw-dataset.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.tubitak-olimpiyat-dataset
TUBITAK Science Olympiad Dataset
This dataset contains multiple-choice and open-ended scientific questions sourced from the TUBITAK (The Scientific and Technological Research Council of Turkey) Science Olympiads spanning various years. It is intended to serve as a benchmark for evaluating the advanced analytical, mathematical, and computational reasoning capabilities of Large Language Models (LLMs) in the Turkish language.
The dataset comprises approximately 2700 problems across… See the full description on the dataset page: https://huggingface.co/datasets/alpsahin/tubitak-olimpiyat-dataset.phy-dataset-1k
Physics Dataset (1K)
Dataset Description
This dataset contains 908 physics problems with accompanying images, captions, and detailed answers. It is designed for training and evaluating multimodal language models on physics knowledge and problem-solving tasks.
Dataset Structure
Data Fields
id (int64): Unique identifier for each problem
problem (string): The physics question or problem statement
caption (List[string]): List of… See the full description on the dataset page: https://huggingface.co/datasets/ttlynne/phy-dataset-1k.SuperQA-Dataset-TestRepo
SuperQA-Dataset
1. Introduction
The SuperQA-Dataset represents a major advancement in question-answering benchmark datasets. In this latest release, we have significantly improved data quality through enhanced curation pipelines, rigorous validation processes, and comprehensive quality assurance measures. The dataset demonstrates exceptional performance across various data quality metrics, making it ideal for training and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/toolevalxm/SuperQA-Dataset-TestRepo.charecterization_paper_datasetVideo-Games-Behavioral-Addiction-Datasetspe_mcq_dataset
SPE MCQ Dataset
SPE MCQ, which stands for Society of Petroleum Engineers Multiple Choice Question, is a dataset which consists of 100 MCQs of petroleum engineering topic. This MCQ bank is originally from the Study Guide for the SPE Petroleum Engineering Certification Examination (4th ed) (2011)
The MCQ has diverse topics, such as common knowledge, drilling, completion and production, and reservoir engineering.
The dataset has 5 columns:
number: Question number
question: Question… See the full description on the dataset page: https://huggingface.co/datasets/ynuwara/spe_mcq_dataset.SRPO_RL_datasets
SRPO Dataset: Reflection-Aware RL Training Data
This repository provides the multimodal reasoning dataset used in the paper:
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
We release two versions of the dataset:
39K version (modified_39Krelease.jsonl + images.zip)
Enhanced 47K+ version (47K_release_plus.jsonl + 47K_release_plus.zip)
Both follow the same unified format, containing multimodal (image–text) reasoning data with self-reflection… See the full description on the dataset page: https://huggingface.co/datasets/bruce360568/SRPO_RL_datasets.Dataset_Large_test
PPU-Bench
