datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpt-4v-distribution-shift
License
This repository is licensed under the MIT License.
Description
This Hugging Face repository hosts the random case dataset utilized in our research project, detailed in the GitHub repository gpt-4v-distribution-shift.
These datasets are crucial for evaluating the performance of multimodal foundation models under various distribution shift scenarios.
Using the Dataset
For detailed instructions on how to use this dataset to reproduce the results presented… See the full description on the dataset page: https://huggingface.co/datasets/jameszhou-gl/gpt-4v-distribution-shift.220k-GPT4Vision-captions-from-LIVIS
220k-GPT4Vision-captions-from-LVIS
by: Christoph Schuhmann, Peter Bevan, 21 Nov, 2023
This dataset comprises 220,000 captioned images from the LVIS dataset. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted into captions using Mistral-7B-OpenOrca.
PROMPT
"""<<SYS>> You are a highly intelligent, empathic, helpful, respectful, and honest assistant with high emotional intelligence. Always… See the full description on the dataset page: https://huggingface.co/datasets/laion/220k-GPT4Vision-captions-from-LIVIS.anime-with-gpt4v-caption-for-lora
Anime style image - text by GPT4V small dataset
The text is as follows:
This is a charming anime-style illustration featuring a young girl as the main subject. The image predominantly uses a soft, pastel color palette, creating a gentle and whimsical ambiance. The main character has light blonde hair styled in two low twintails, secured with what could be interpreted as dark-colored hair ties or ribbons. She has large expressive blue eyes and a demure expression, with… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/anime-with-gpt4v-caption-for-lora.GameplayCaptions-GPT-4Vdreamlip-gpt4v-500kMMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.GameplayCaptions-GPT-4V-V2allava_instruct_gpt4v_zhgpt4v-datasetlaion-gpt4v-from-lavisgpt4v-raw-chunksMMInstruct-GPT4V_mistral-7b_l0_cutMMInstruct-GPT4V_mistral-7b_cosi_cutMMInstruct-GPT4V_mistral-7b_cooccur_cutlaion-14k-GPT4V-LIVIS-CaptionsHidden-Flaws-GPT-4VModels, data, and code posted here reflect the research conducted in the Computational Biology Branch, NCBI/NLM. The information produced on this website is not intended for direct diagnostic use or medical decision-making without review and oversight by a clinical professional. Individuals should not change their health behavior solely on the basis of information produced on this website. NIH does not independently verify the validity or utility of the information produced by this tool. If… See the full description on the dataset page: https://huggingface.co/datasets/ncbi/Hidden-Flaws-GPT-4V.laion-14k-GPT4V-LIVIS-Captions_Malayalam
Malayalam translated version of laion-14k-GPT4V-LIVIS-Captions
Translated using indictrans2
Translation Code : code
gpt4v-LAION-discord
Dataset Card for "gpt4v-LAION-discord"
More Information needed
textocr-gpt4v
Dataset Card for TextOCR-GPT4V
Dataset Summary
TextOCR-GPT4V is Meta's TextOCR dataset dataset captioned with emphasis on text OCR using GPT4V. To get the image, you will need to agree to their terms of service.
Supported Tasks
The TextOCR-GPT4V dataset is intended for generating benchmarks for comparison of an MLLM to GPT4v.
Languages
The caption languages are in English, while various texts in images are in many languages such as Spanish, Japanese… See the full description on the dataset page: https://huggingface.co/datasets/jimmycarter/textocr-gpt4v.llama3base_rewritert0_v3_ablt-gpt4vsm0-wofilter_transformedsftgpt-4v-eval-samples
GPT-4V Eval samples
This is a hand curated images from the web and questions asked by myself to GPT-4V to understand its ability and limits.
I am mainly focus in localization, OCR ability and understanding of GPT-4V vision module. So the language part is skipped as we already seen in GPT-4. As long as GPT-4V can extract the required information in text, the rest of the LLM shouldn't have any issue answering the rest of the questions.
The numbers of examples is still pretty tiny… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/gpt-4v-eval-samples.gpt4v-briefingsgpt4v-emotion-dataset
Dataset Card for "gpt4v-emotion-dataset"
More Information needed
llama3base_rewritert0_v3_ablt-gpt4vsm0-wfilter_transformedsfttextocr_gpt4v_cleaned
textocr_gpt4v_cleaned
The textocr(gpt4v)__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
49,207
QA turns
195,643
answers rewritten by the cleaning pass
0
QA created by the cleaning pass (new_qa)
not measured for this family
shards
86
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers… See the full description on the dataset page: https://huggingface.co/datasets/Elliot-Data/textocr_gpt4v_cleaned.GPT-4V-ChatsMMInstruct-GPT4V_mistral-7b_cooccur_fullDesign2Code_human_eval_reference_vs_gpt4vFind more details in our paper.
gpt4v_cleaned
gpt4v_cleaned
The gpt4v__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
10
QA turns
43
answers rewritten by the cleaning pass
1
QA created by the cleaning pass (new_qa)
35 (81.4%)
shards
1
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong but salvageable,
drops what it… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/gpt4v_cleaned.GPT-4V-DescribeChangesCutscene
Dataset Card for "GPT-4V-DescribeChangesCutscene"
More Information needed
