visual-instruction
food-visual-instructions
Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025)
This repos contains the food visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models.
The main project page is: Adapt-MLLM-to-Domains
Data Information
Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from extended Recipe1M+ dataset. These synthetic… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/food-visual-instructions.pandagpt_visual_instruction_dataset[Dataset Details] This dataset is constructed by combining LLaVA Visual Instruct 150K and the dataset released by MiniGPT-4.
[License] Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Intended use
Primary intended uses: The primary use of this dataset is research on large multimodal models and chatbots.
Primary intended users: The primary intended users of the model are researchers and hobbyists in… See the full description on the dataset page: https://huggingface.co/datasets/openllmplayground/pandagpt_visual_instruction_dataset.remote-sensing-visual-instructions
Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025)
This repos contains the remote-sensing visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models.
The main project page is: Adapt-MLLM-to-Domains
Data Information
Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from NWPU-Captions, RSICD, RSITMD… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/remote-sensing-visual-instructions.biomed-visual-instructions
Adapting Multimodal Large Language Models to Domains via Post-Training (EMNLP 2025)
This repos contains the biomedicine visual instructions for post-training MLLMs in our paper: On Domain-Specific Post-Training for Multimodal Large Language Models.
The main project page is: Adapt-MLLM-to-Domains
Data Information
Using our visual instruction synthesizer, we generate visual instruction tasks based on the image-caption pairs from PubMedVision (referred to as PMC_refined… See the full description on the dataset page: https://huggingface.co/datasets/AdaptLLM/biomed-visual-instructions.DOD-Instruction-5040-02-Visual-Information
DoD Visual Information Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 250 document-grounded question-and-answer records based on DoD Instruction 5040.02, “Visual Information (VI),” dated October 27, 2011, and incorporating Change 2 effective April 20, 2018.
The source establishes Department of Defense policy, responsibilities, and procedures for managing visual-information records… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DOD-Instruction-5040-02-Visual-Information.visual_instruction_tuning_ID_reference
