datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TrainSet
SIQA TrainSet
A standardized TrainSet for the Scientific Image Quality Assessment (SIQA), designed to train multimodal models on two core tasks SIQA-U & SIQA-S.
NOTICE:Due to Hugging Face Datasets' automatic caching mechanism, the local disk usage can be up to ~20× larger than the actual dataset size.
You only need to download the 'images/' folder to run inference or evaluation, please ignore the .parquet in 'SIQA-S/' and 'SIQA-U'
To Quick Start, you download by git:
git clone… See the full description on the dataset page: https://huggingface.co/datasets/SIQA/TrainSet.VisOnlyQA_Train
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information".
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception tasks on 4… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_Train.mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 4B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k.mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.MetaVQA-Trainmhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.TSUMM-Suite_Training
🧱 TSUMM-Suite Training Set
This repository provides the training set of TSUMM-Suite, a multimodal dataset for unified time series understanding and generation. It contains forecasting and imputation tasks for TS-image generation. By aligning understanding tasks with generation tasks, TSUMM-Suite enables temporal understanding to improve time series generation.
🎨 Task Illustration
TSUMM-Suite… See the full description on the dataset page: https://huggingface.co/datasets/TimeOmni-VL/TSUMM-Suite_Training.madqa-training
Chrisyichuan/madqa-training
MADQA document QA contrastive training data with hard negatives.
Contents
madqa_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 1840
unique images: 3598
avg… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/madqa-training.GeoQA-train-Vision-R1-cot-rewrite
Dataset Card for GeoQA-train-Vision-R1-cot-rewrite
This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models.
Dataset Details
Dataset Description
The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features detailed and… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-train-Vision-R1-cot-rewrite.refcocoplus_trainGeoQA-PLUS-aug-train-Vision-R1-cot-rewrite
Dataset Card for GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite
This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA-PLUS subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models.
Dataset Details
Dataset Description
The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite.BACE-V-Train
BACE-V-SMILES Train Dataset
Dataset Description
This dataset contains molecular data with visual representations for BACE related compounds.
Features
Question: Question related to the molecule
Answer: Corresponding answer
TargetMolecule: SMILES representation of the target molecule
SampleMethod: Method used for sampling
SampleNum: Sample number
SampleRep: Sample repetition
image: Generated molecular structure image from SMILES
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/molvision/BACE-V-Train.moca-colpali-training
Chrisyichuan/moca-colpali-training
MOCA ColPali contrastive training data with hard negatives.
Contents
moca_colpali_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 118195
unique… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-colpali-training.moca-visrag-ind-training
Chrisyichuan/moca-visrag-ind-training
MOCA VisRAG independent-split contrastive training data with hard negatives.
Contents
moca_visrag_ind_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 122752… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-ind-training.BBBP-V-Train
BBBP-V-SMILES Train Dataset
Dataset Description
This dataset contains molecular data with visual representations for BBBP related compounds.
Features
Question: Question related to the molecule
Answer: Corresponding answer
TargetMolecule: SMILES representation of the target molecule
SampleMethod: Method used for sampling
SampleNum: Sample number
SampleRep: Sample repetition
image: Generated molecular structure image from SMILES
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/molvision/BBBP-V-Train.HIV-V-Train
HIV-V-SMILES Train Dataset
Dataset Description
This dataset contains molecular data with visual representations for HIV related compounds.
Features
Question: Question related to the molecule
Answer: Corresponding answer
TargetMolecule: SMILES representation of the target molecule
SampleMethod: Method used for sampling
SampleNum: Sample number
SampleRep: Sample repetition
image: Generated molecular structure image from SMILES
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/molvision/HIV-V-Train.moca-visrag-syn-training
Chrisyichuan/moca-visrag-syn-training
MOCA VisRAG synthetic-split contrastive training data with hard negatives.
Contents
moca_visrag_syn_converted.jsonl — query-image pairs with hard negatives
images/ — all referenced images
Each metadata row:
{
"query": "...",
"chunk_path": "images/...",
"neg_chunk_paths": ["images/...", "images/..."],
"source_positive_rank": 0,
"source_positive_score": 0.0,
"source_dataset": "moca"
}
Summary
rows: 239206… See the full description on the dataset page: https://huggingface.co/datasets/Chrisyichuan/moca-visrag-syn-training.Tox21-V-Train
Tox21-V-SMILES Train Dataset
Dataset Description
This dataset contains molecular data with visual representations for Tox21 related compounds.
Features
Question: Question related to the molecule
Answer: Corresponding answer
TargetMolecule: SMILES representation of the target molecule
SampleMethod: Method used for sampling
SampleNum: Sample number
SampleRep: Sample repetition
image: Generated molecular structure image from SMILES
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/molvision/Tox21-V-Train.Clintox-V-Train
Clintox-V-SMILES Train Dataset
Dataset Description
This dataset contains molecular data with visual representations for Clintox related compounds.
Features
Question: Question related to the molecule
Answer: Corresponding answer
TargetMolecule: SMILES representation of the target molecule
SampleMethod: Method used for sampling
SampleNum: Sample number
SampleRep: Sample repetition
image: Generated molecular structure image from SMILES
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/molvision/Clintox-V-Train.
