CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mvp-lab /LLaVA-OneVision-2-Data LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family. The release contains large-scale video data at several duration ranges, video captions and source mappings, and spatial-reasoning data used for mid-training. At a Glance The dataset is split across two Hugging Face repositories because of its size: Repository What it contains Part 1 (this repository) ~60-second video shards, captions for all duration ranges… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-2-Data.imagevideo-text-to-textn<1K39 likes125k downloads24d agoHugging Face02mvp-lab /LLaVA-OneVision-1.5-Instruct-Data LLaVA-OneVision-1.5 Instruction Data Paper | Code 📌 Introduction This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the development of LLaVA-OneVision-1.5. LLaVA-OneVision-1.5 is a novel family of Large Multimodal Models (LMMs) that achieve state-of-the-art performance with significantly reduced computational and financial costs. This meticulously curated 22M instruction dataset (LLaVA-OneVision-1.5-Instruct) is part of a… See the full description on the dataset page: https://huggingface.co/datasets/mvp-lab/LLaVA-OneVision-1.5-Instruct-Data.imageimage-text-to-text10M<n<100M81 likes61k downloads2mo agoHugging Face03lmms-lab /LLaVA-Video-178K Dataset Card for LLaVA-Video-178K Uses This dataset is used for the training of the LLaVA-Video model. We only allow the use of this dataset for academic research and education purpose. For OpenAI GPT-4 generated data, we recommend the users to check the OpenAI Usage Policy. Data Sources For the training of LLaVA-Video, we utilized video-language data from five primary sources: LLaVA-Video-178K: This dataset includes 178,510 caption entries, 960,792 open-ended… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-Video-178K.textvisual-question-answering1M<n<10M202 likes36k downloads2y agoHugging Face04lmms-lab /LLaVA-OneVision-Data Dataset Card for LLaVA-OneVision [2024-09-01]: Uploaded VisualWebInstruct(filtered), it's used in OneVision Stage almost all subsets are uploaded with HF's required format and you can use the recommended interface to download them and follow our code below to convert them. the subset of ureader_kg and ureader_qa are uploaded with the processed jsons and tar.gz of image folders. You may directly download them from the following url.… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Data.image1M<n<10M238 likes24k downloads1y agoHugging Face05lmms-lab /LLaVA-ReCap-CC12Mimage1M<n<10M9 likes16k downloads2y agoHugging Face06ackermans26 /LLaVA-OneVision-1.5-Instruct-Data-qwen-formattext1M<n<10M0 likes6.7k downloads2mo agoHugging Face07lmms-lab-encoder /llava-bench-in-the-wild Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of LLaVA-Bench(wild) that is used in LLaVA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @misc{liu2023improvedllava, author={Liu, Haotian and Li, Chunyuan… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/llava-bench-in-the-wild.imagen<1K10 likes5.6k downloads3y agoHugging Face08d0rj /LLaVA-OneVision-Data-ru LLaVA-OneVision-Data-ru Translated lmms-lab/LLaVA-OneVision-Data dataset into Russian language using Google translate. Almost all datasets have been translated, except for the following: ["tallyqa(cauldron,llava_format)", "clevr(cauldron,llava_format)", "VisualWebInstruct(filtered)", "figureqa(cauldron,llava_format)", "magpie_pro(l3_80b_mt)", "magpie_pro(qwen2_72b_st)", "rendered_text(cauldron)", "ureader_ie"] Usage import datasets data =… See the full description on the dataset page: https://huggingface.co/datasets/d0rj/LLaVA-OneVision-Data-ru.imagetext-generation1M<n<10M4 likes5k downloads2y agoHugging Face09lmms-lab /LLaVA-NeXT-Data Dataset Card for LLaVA-NeXT We provide the whole details of LLaVA-NeXT Dataset. In this dataset, we include the data that was used in the instruction tuning stage for LLaVA-NeXT and LLaVA-NeXT(stronger). Aug 30, 2024: We update the dataset with raw format (de-compress it for json file and images with structured folder), you can directly download them if you are familiar with LLaVA data format. Dataset Sources Compared to the instruction data mixture for LLaVA-1.5… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-NeXT-Data.image100K<n<1M47 likes3.8k downloads2y agoHugging Face10Icey444 /llava_v1_5_mix665k LLaVA v1.5 Mix 665K Dataset This dataset contains 665,298 multimodal instruction-following samples used for fine-tuning the LLaVA v1.5 model. Dataset Structure id: Unique identifier for the sample model: Model name (if applicable) conversations: JSON string containing conversation turns in original format image: List of PIL Image objects (embedded in parquet) image_path: List of strings containing original relative paths to images Load the Dataset from… See the full description on the dataset page: https://huggingface.co/datasets/Icey444/llava_v1_5_mix665k.imageimage-text-to-text100K<n<1M1 likes3.1k downloads11mo agoHugging Face11lmms-lab-encoder /LLaVA-NeXT-Interleave-Bench LLaVA-Interleave Bench Dataset Card Dataset details Dataset type: LLaVA-Interleave Bench is a comprehensive set of multi-image datasets that are collected from public datasets or generated by the GPT-4V API. It is constructed for evaluating the interleaved multi-image reaoning capbilities of LMMs. Dataset date: LLaVA-Interleave Bench was collected in April 2024, and released in June 2024. Paper or resources for more information: Blog:… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/LLaVA-NeXT-Interleave-Bench.imagevisual-question-answering10K<n<100K15 likes2.5k downloads2y agoHugging Face12theblackcat102 /llava-instruct-mix LLaVA Instruct Mix Added OCR and Chart QA dataset into this for more text extraction questions imagevisual-question-answering100K<n<1M12 likes2.5k downloads3y agoHugging Face13Project-AgML /Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset Agri-LLaVA Agri-LLaVA is a large multimodal instruction dataset for agriculture, pairing crop/leaf images with multi-turn diagnostic conversations about plant diseases, pests, and nutrient deficiencies. It is compiled from 16 public source datasets (see the license table below). This dataset has been converted to Parquet format with image bytes embedded directly, standardized to the HF image_text_to_text format with a single conversational messages schema. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/Agri-LLaVA_Agricultural_Pests_And_Diseases_Feature_Alignment_Dataset.imageimage-text-to-text100K<n<1M0 likes2.3k downloads2mo agoHugging Face14lmms-lab /LLaVA-ReCap-CC3Mimage1M<n<10M21 likes2k downloads2y agoHugging Face15Xkev /LLaVA-CoT-100k Dataset Card for LLaVA-CoT The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.textvisual-question-answering10K<n<100K106 likes1.9k downloads9mo agoHugging Face16Leonardo6 /llava-finetuneimage100K<n<1M0 likes1.9k downloads1y agoHugging Face17HuggingFaceH4 /llava-instruct-mix-vsfttheblackcat102/llava-instruct-mix reformated for VSFT with TRL's SFT Trainer. See https://github.com/huggingface/trl/blob/main/examples/scripts/vsft_llava.py. image100K<n<1M49 likes1.7k downloads2y agoHugging Face18Ethlake /llava-665k LLaVA-v1.5 Mix665K — Arrow (images embedded) The LLaVA-v1.5 visual instruction-tuning mixture (llava_v1_5_mix665k) converted to a 🤗 datasets Arrow dataset with image bytes embedded. 665,298 examples (624,610 image–text + 40,688 text-only). ⚠️ This repo is a raw save_to_disk Arrow snapshot. The dataset viewer and load_dataset() do not work here — load it with load_from_disk as shown below. Loading from huggingface_hub import snapshot_download from datasets… See the full description on the dataset page: https://huggingface.co/datasets/Ethlake/llava-665k.imagevisual-question-answering100K<n<1M0 likes1.5k downloads3mo agoHugging Face19KaiChen1998 /coda-lm-llava-format CODA-LM Dataset Card CODA-LM is the multi-modal version of the CODA dataset, used in the CODA-LM paper. Both English and Chinese annotations are available. Check detailed usage in our Github repo. This repo contains the CODA-LM dataset, which has been reorganized in the LLaVA data format. You are also welcome to check the original CODA-LM data which contains more metadata vanilla annotations. Usage from datasets import load_dataset # name can be selected from… See the full description on the dataset page: https://huggingface.co/datasets/KaiChen1998/coda-lm-llava-format.imageimage-to-text10K<n<100K3 likes1.4k downloads2y agoHugging Face20BUAADreamer /llava-en-zh-300kThis dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English Visual Instruction Data from openbmb. You can use it in LLaMA Factory by specifying --dataset llava_150k_en,llava_150k_zh. imagetext-generation100K<n<1M36 likes1.3k downloads2y agoHugging Face21spatial-reason /llava_trajectoriesimage1K<n<10K0 likes1.2k downloads7mo agoHugging Face22brivangl /midjourney-v6-llavaThis dataset based on https://huggingface.co/datasets/CortexLM/midjourney-v6 dataset, captioned with LLava-1.6 model. This dataset was released as is. By accessing and using this dataset, you acknowledge and agree that Cortex Foundation and the author of this repo are not responsible for any copyright violations or legal consequences that may arise from the use of these images. imagetext-to-image100K<n<1M17 likes1.2k downloads2y agoHugging Face23opencsg /LLaVA-Instruct-600K-Chinese 仿照 LLaVA-Instruct-150K ,使用 Qwen2.5-VL-32B-Instruct 合成的用于微调中文VLM的数据;也可以与英文数据集混合使用,训练多语言VLM 任务类型为基于单张图片的问答和对话,每个样本都对应一张不同的图片,其中大部分图片包含中文字符,更适合中文场景下视觉语言模型的训练。 图片从各类中文网站上爬取 包含3类任务:日常对话、复杂推理、描述图片。日常对话通常是5轮对话,其余任务是1轮对话。 每种任务的数量如下: 任务类型 数量 日常对话 247,431 复杂推理 194,646 描述图片 199,791 用于生成对话数据的prompt如下 日常对话 设计一个你和一个询问这张照片的人之间的对话。答案应该是视觉AI助手看到图像并回答问题的语气。 你需要提出不同的问题并给出相应的答案。问题可以包括询问图像视觉内容的问题,包括对象类型、对象计数、对象动作、对象位置、对象之间的相对位置等。必须是有明确答案的问题,即 (1) 人们可以在图像中明确看到问题所问的内容,并且可以自信地回答; (2)… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/LLaVA-Instruct-600K-Chinese.imagevisual-question-answering100K<n<1M9 likes1k downloads1y agoHugging Face24trl-lib /llava-instruct-mix LLaVA Instruct Mix Summary The LLaVA Instruct Mix dataset is a processed version of LLaVA Instruct Mix. Data Structure Format: Conversational Type: Language-modeling Columns: "images": The image associated with the text. "prompt": A list of messages that form the context for the conversation. "completion": The last message in the conversation, which is the model's response. This structure allows models to learn from the context of the conversation… See the full description on the dataset page: https://huggingface.co/datasets/trl-lib/llava-instruct-mix.image100K<n<1M4 likes945 downloads1y agoHugging Face25myeongkyunkang /LLaVA-Med-60K-IM-text LLaVA-Med-60K-IM-text This dataset is a text format of llava_med_instruct_60k_inline_mention.json. We built this dataset using the Meta-Llama-3-70B-Instruct, and the instruction we used is: Rewrite the question-answer pairs into a paragraph format (Do not use the words 'question' and 'answer' in your responses):. PMC articles that failed to download are excluded. Non-medical images (e.g., diagrams) are excluded in an automatic way. Despite these efforts, this dataset is not… See the full description on the dataset page: https://huggingface.co/datasets/myeongkyunkang/LLaVA-Med-60K-IM-text.image10K<n<100K0 likes914 downloads2y agoHugging Face26asadfgglie /LLaVA-Instruct-150K-zh-tw使用openCC將BUAADreamer/llava-en-zh-300k翻譯成繁體中文 image100K<n<1M2 likes891 downloads2y agoHugging Face27CaptionEmporium /coyo-hd-11m-llavanext Dataset Card for coyo-hd-11m-llavanext Dataset Summary This is a data of 22,794,288 synthetic captions for 11,397,144 images from coyo-700m. The "hd" in the title refers to two aspects: high density and high definition. While large alt-text image pair datasets have many images, only a very small proportion of these images are in higher resolutions and have substantial concept density. For example, many of these datasets consist of more than 50% thumbnail sized or very… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/coyo-hd-11m-llavanext.imagetext-to-image10M<n<100M29 likes750 downloads2y agoHugging Face28lmms-lab /llava-critic-113k Dataset Card for LLaVA-Critic-113k 🪐 Project Page: https://llava-vl.github.io/blog/2024-10-03-llava-critic/ 📰 Paper: https://arxiv.org/abs/2410.02712 🤗 Huggingface Collection: https://huggingface.co/collections/lmms-lab/llava-critic-66fe3ef8c6e586d8435b4af8 👋 Point of Contact: Tianyi Xiong Dataset Summary LLaVA-Critic-113k is a high quality critic instruction-following dataset tailored to follow instructions in complex evaluation setting, providing both… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/llava-critic-113k.image100K<n<1M28 likes726 downloads2y agoHugging Face29tejasvaidhya /llava-cc3m-pretrain-595Ktext100K<n<1M0 likes716 downloads3y agoHugging Face30acthuan /LLaVA-Med HealthGPT : A Medical Large Vision-Language Model for Unifying Comprehension and Generation via Heterogeneous Knowledge Adaptation Tianwei Lin1, Wenqiao Zhang1, Sijing Li1, Yuqian Yuan1, Binhe Yu2, Haoyuan Li3, Wanggui He3, Hao Jiang3, Mengze Li4, Xiaohui Song1, Siliang Tang1, Jun Xiao1, Hui Lin1, Yueting Zhuang1, Beng Chin Ooi5 1Zhejiang University, 2University of Electronic Science and Technology of China, 3Alibaba, 4The Hong Kong University of Science and Technology, 5National… See the full description on the dataset page: https://huggingface.co/datasets/acthuan/LLaVA-Med.text0 likes700 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.