CoolFace
Datasetpublic

Kaij00/MSVQA

This is a multimodal cross-scenario dataset for continual learning with MLLMs. We provide a simple script to split the dataset in multiple ways. The dataset format has been adjusted for Qwen. The coordinates in 'train_annfiles.json' and 'val_annfiles.json' are adjusted to Qwen2.5VL format. And 'train_annfiles_ori.json' and 'val_annfiles_ori.json' retain the original coordinates of the bounding box. You need to adjust the coordinates fit your format. Detailed information can refer to… See the full description on the dataset page: https://huggingface.co/datasets/Kaij00/MSVQA.

sourceHugging Facemitupdated 9mo agoView on Hugging Face
2likes3.5kdownloads
Dataset Card

This is a multimodal cross-scenario dataset for continual learning with MLLMs. We provide a simple script to split the dataset in multiple ways.

The dataset format has been adjusted for Qwen. The coordinates in 'trainannfiles.json' and 'valannfiles.json' are adjusted to Qwen2.5VL format. And 'trainannfilesori.json' and 'valannfilesori.json' retain the original coordinates of the bounding box.

You need to adjust the coordinates fit your format.

Detailed information can refer to http://arxiv.org/abs/2511.18507

If you use this datasets, please cite:

@misc{jiang2025multimodalcontinuallearningmllms, title={Multimodal Continual Learning with MLLMs from Multi-scenario Perspectives}, author={Kai Jiang and Siqi Huang and Xiangyu Chen and Jiawei Shao and Hongyuan Zhang and Xuelong Li}, year={2025}, eprint={2511.18507}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2511.18507}, }

Source data is from 4 public datasets:

@article{2022FAIR1M, title={FAIR1M: A benchmark dataset for fine-grained object recognition in high-resolution remote sensing imagery}, author={ Sun, Xian and Wang, Peijin and Yan, Zhiyuan and Xu, Feng and Wang, Ruiping and Diao, Wenhui and Chen, Jin and Li, Jihao and Feng, Yingchao and Xu, Tao }, journal={ISPRS Journal of Photogrammetry and Remote Sensing}, volume={184}, year={2022}, }

@article{RUOD23, title={Rethinking general underwater object detection: Datasets, challenges, and solutions}, author={Chenping, Fu and Risheng, Liu and Xin, Fan and Puyang, Chen and Hao, Fu and Wanqi, Yuan and Ming, Zhu and Zhongxuan, Luo}, journal={Neurocomputing}, year={2023} }

@article{zhu2021detection, title={Detection and tracking meet drones challenge}, author={Zhu, Pengfei and Wen, Longyin and Du, Dawei and Bian, Xiao and Fan, Heng and Hu, Qinghua and Ling, Haibin}, journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, volume={44}, number={11}, pages={7380--7399}, year={2021}, publisher={IEEE} }

@INPROCEEDINGS{Damen2018EPICKITCHENS, title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset}, author={Damen, Dima and Doughty, Hazel and Farinella, Giovanni Maria and Fidler, Sanja and Furnari, Antonino and Kazakos, Evangelos and Moltisanti, Davide and Munro, Jonathan and Perrett, Toby and Price, Will and Wray, Michael}, booktitle={European Conference on Computer Vision (ECCV)}, year={2018} }