datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-4v-distribution-shift
License
This repository is licensed under the MIT License.
Description
This Hugging Face repository hosts the random case dataset utilized in our research project, detailed in the GitHub repository gpt-4v-distribution-shift.
These datasets are crucial for evaluating the performance of multimodal foundation models under various distribution shift scenarios.
Using the Dataset
For detailed instructions on how to use this dataset to reproduce the results presented… See the full description on the dataset page: https://huggingface.co/datasets/jameszhou-gl/gpt-4v-distribution-shift.GPT4O_Image_T2Ipokemon-gpt4-captions
Dataset Card for "pokemon-gpt4-captions"
This dataset is just lambdalabs/pokemon-blip-captions but the captions come from GPT-4 (Turbo).
Code used to generate the captions:
import base64
from io import BytesIO
import requests
from PIL import Image
def encode_image(image):
buffered = BytesIO()
image.save(buffered, format="JPEG")
img_str = base64.b64encode(buffered.getvalue())
returnimg_str.decode("utf-8")
def create_payload(image_string):
payload = {… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/pokemon-gpt4-captions.220k-GPT4Vision-captions-from-LIVIS
220k-GPT4Vision-captions-from-LVIS
by: Christoph Schuhmann, Peter Bevan, 21 Nov, 2023
This dataset comprises 220,000 captioned images from the LVIS dataset. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted into captions using Mistral-7B-OpenOrca.
PROMPT
"""<<SYS>> You are a highly intelligent, empathic, helpful, respectful, and honest assistant with high emotional intelligence. Always… See the full description on the dataset page: https://huggingface.co/datasets/laion/220k-GPT4Vision-captions-from-LIVIS.anime-with-gpt4v-caption-for-lora
Anime style image - text by GPT4V small dataset
The text is as follows:
This is a charming anime-style illustration featuring a young girl as the main subject. The image predominantly uses a soft, pastel color palette, creating a gentle and whimsical ambiance. The main character has light blonde hair styled in two low twintails, secured with what could be interpreted as dark-colored hair ties or ribbons. She has large expressive blue eyes and a demure expression, with… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/anime-with-gpt4v-caption-for-lora.gpt-4o-minicoco-gpt4oTest
MMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.GPT-4o-RestoreThis dataset is associated with the paper: https://arxiv.org/abs/2505.05621.
dreamlip-gpt4v-500kmm-interp-CompCap-gpt4-data
CompCap-GPT4: A GPT-4 Captioned Version of CompCap-118K
Dataset Sources
Paper: CompCap: Improving Multimodal Large Language Models with Composite Captions
Dataset Structure
Download Options
Direct Download:The repository includes CI_type.zip and CI_type.json. The JSON file follows the Llava format:
{
"id": ID,
"image": IMAGE_PATH,
"conversations": [
{"from": "human", "value": QUESTION},
{"from": "gpt", "value": ANSWER}
]
}
Using… See the full description on the dataset page: https://huggingface.co/datasets/htlou/mm-interp-CompCap-gpt4-data.CompCap-gpt4
CompCap-GPT4: A GPT-4 Captioned Version of CompCap-118K
Dataset Sources
Paper: CompCap: Improving Multimodal Large Language Models with Composite Captions
Dataset Structure
Download Options
Direct Download:The repository includes CI_type.zip and CI_type.json. The JSON file follows the Llava format:
{
"id": ID,
"image": IMAGE_PATH,
"conversations": [
{"from": "human", "value": QUESTION},
{"from": "gpt", "value": ANSWER}
]
}
Using… See the full description on the dataset page: https://huggingface.co/datasets/xchen16/CompCap-gpt4.allava_instruct_gpt4v_zhgpt4v-datasetcomicstrips-gpt4o-blip3
Comic Strips
Dataset Details
Dataset Description
This dataset contains indie comics from Reddit, then captioned with GPT4o and BLIP3.
Currently, only the GPT4o captions are available in this repository. The BLIP3 captions will be uploaded soon.
Roughly 1400 images were captioned at a cost of ~$11 using GPT4o (25 May 2024 version).
Curated by: @pseudoterminalx
Funded by @pseudoterminalx
License: MIT
Dataset Sources
Unlike other free-to-use… See the full description on the dataset page: https://huggingface.co/datasets/bghira/comicstrips-gpt4o-blip3.laion-gpt4v-from-lavisgpt4v-raw-chunkscrag-mm-single-turn-public-v0.1.2-with-images-gpt4.1-v2gpt4o-receipt
GPT4o-Receipt: AI-Generated Receipt Dataset
This directory contains the AI-generated receipts from the
GPT4o-Receipt benchmark, introduced in:
GPT4o-Receipt: A Dataset and Human Study for AI-Generated Document ForensicsYan Zhang*, Simiao Ren*†, Ankit Raj, En Wei, Dennis Ng, Alex Shen, Jiayu Xue, Yuxin Zhang, Evelyn MarottaarXiv:2603.11442 · March 2026 · CC BY-NC-SA 4.0*Equal contribution. †Corresponding author: benren@scam.ai
What Is GPT4o-Receipt?
GPT4o-Receipt is… See the full description on the dataset page: https://huggingface.co/datasets/Scam-AI/gpt4o-receipt.GPT4O_Image_I2Ipokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
GPT4Scene-All
Validation Dataset of GPT4Scene
🏠 Overview
This dataset card is for the GPT4Scene project. You can see the more information below.
Github Code: Link to Github
Arxiv Paper: Link to Arxiv
Project Page: Link to Project
🤗 Hugging Face
Function
Huggingface Link
Validation Dataset
alexzyqi/GPT4Scene-Val-Dataset
Validation Annotations
alexzyqi/GPT4Scene-Val-Annotation
Pretrain Models
Qwen/Qwen2-VL-7B-Instruct
Trained Weights… See the full description on the dataset page: https://huggingface.co/datasets/alexzyqi/GPT4Scene-All.flickr30k-transformed-captions-gpt4oThis is a "de-biased" version of https://huggingface.co/datasets/nlphuji/flickr30k dataset.
The new alt_text column was produced by GPT-4o using the following script : https://github.com/mozilla/distilvit/blob/main/distilvit/curate_gpt.py
Learn more about why and how we did it here : https://github.com/mozilla/distilvit/blob/main/docs/fighting_bias.md
See the code here : https://github.com/mozilla/distilvit/blob/main/distilvit/curate_gpt.py
For the licence, see the original dataset.
airoboros-gpt4-m2.0
Overview
This is a merge of https://hf.co/datasets/jondurbin/airoboros-gpt4-1.4.1 and https://hf.co/datasets/jondurbin/airoboros-gpt4-2.0
Category breakdown
Licence and usage restrictions
The data was generated by gpt-4 via OpenAI API calls.
The ToS for OpenAI API usage has a clause preventing the output from being used to train a model that competes with OpenAI
what does compete actually mean here?
these small open source models will not produce output… See the full description on the dataset page: https://huggingface.co/datasets/jondurbin/airoboros-gpt4-m2.0.pexels-gpt4oImages collected from Pexels, using 1000 images following 5 categories:
nudes
war
group
animals
smoking
See https://www.pexels.com/license/ for the license
They were then annotated using gpt4-o, see https://github.com/mozilla/distilvit/blob/main/distilvit/gpt4.py
dreamsim_crop_cosine-gpt4_DE_diverse_promptslaion-14k-GPT4V-LIVIS-Captionspokemon-gpt4-1kThis dataset was modified from diffusers/pokemon-gpt4-captions and contains 1k Pokémon-related image-captioning instruction data points.
You can organize content in the dataset_info.json in LLaMA Factory like this:
"pokemon_1k": {
"hf_hub_url": "BUAADreamer/pokemon-gpt4-1k",
"formatting": "sharegpt",
"columns": {
"messages": "messages",
"images": "images"
},
"tags": {
"role_tag": "role",
"content_tag": "content",
"user_tag": "user",
"assistant_tag":… See the full description on the dataset page: https://huggingface.co/datasets/BUAADreamer/pokemon-gpt4-1k.dreamsim_crop_cosine-gpt4_diverse_prompts_NG_NT_NKE5_NKCO50_L2B5dreamsim_crop_cosine-gpt4_DE_diverse_prompts_NG
