datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.gpt4v-raw-chunkstextocr-gpt4v
Dataset Card for TextOCR-GPT4V
Dataset Summary
TextOCR-GPT4V is Meta's TextOCR dataset dataset captioned with emphasis on text OCR using GPT4V. To get the image, you will need to agree to their terms of service.
Supported Tasks
The TextOCR-GPT4V dataset is intended for generating benchmarks for comparison of an MLLM to GPT4v.
Languages
The caption languages are in English, while various texts in images are in many languages such as Spanish, Japanese… See the full description on the dataset page: https://huggingface.co/datasets/jimmycarter/textocr-gpt4v.gpt4vsent4arenahard_gpt4vsllama3
Dataset Card for radm/arenahard_gpt4vsllama3
The dataset was created for fine-tuning Llama-3-70B-Instruct as a judge on Arena Hard (https://github.com/lm-sys/arena-hard-auto)
Dataset Info
question_id: question id from Arena Hard
instruction: original instruction from Arena Hard
model: model whose responses are evaluated against the baseline model (gpt-4-0314) - gpt-4-turbo-2024-04-09 (score: 82.6) and Llama-2-70b-chat-hf (score: 11.6)
input: responses of the evaluated… See the full description on the dataset page: https://huggingface.co/datasets/radm/arenahard_gpt4vsllama3.gpt4vsent2.0
