datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ScaleEdit-12M
ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework
📌 Overview
The largest open-source instruction-based image editing dataset to date.
ScaleEdit-12M contains 11.8 million rigorously verified instruction–image pairs spanning 23 task families across diverse real and synthetic visual domains. It was constructed using ScaleEditor, a fully open-source hierarchical multi-agent framework that eliminates… See the full description on the dataset page: https://huggingface.co/datasets/InternVL-U/ScaleEdit-12M.InternVL-PerformanceInternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption
pexels-568k-internvl2
Dataset Card for pexels-568k-internvl2
Dataset Summary
This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed.
Languages
The text is in English, but occasionally text in images in other languages is transcribed.
Intended Usage
Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.InternVL-evalInternVL-Chat-V1-2-SFT-Data
Data Card for InternVL-Chat-V1-2-SFT-Data
Overview
Inspired by LLaVA-NeXT, we adopted a data-efficient SFT strategy to train InternVL-Chat-V1-2, utilizing approximately 1.2M of visual instruction tuning samples in total, all of which are fully open-source. In a macro sense, we build upon ShareGPT-4V and additionally integrate LLaVA-ZH, DVQA, ChartQA, AI2D, DocVQA, GeoQA+, and SynthDoG-EN. Most of the data remains consistent with LLaVA-NeXT.
Citation
If you use… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-Chat-V1-2-SFT-Data.InternVL_Chat_V12_SFT_Datainternvl3-2b-coco-apgd-eps8internvl3-2b-coco-apgd-eps16internvl-auditor-v2internvl3-2b-coco-apgd-eps4MathCanvas_internvl3_5_math_origin_predflickr-megalith-10m-internvl2-multi-caption
Dataset Card for flickr-megalith-10m-internvl2-multi-caption
Dataset Summary
This is approximately 57.3 million synthetic captions for the images found in madebyollin/megalith-10m.
It includes the following captions:
InternVL2 8B long captions (by CaptionEmporium)
InternVL2 8B short captions (by CaptionEmporium)
Florence2 long captions (by aipicasso)
Florence2 short captions (by CaptionEmporium)
ShareCaptioner long captions (by drawthingsai)
ShareCaptioner short… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/flickr-megalith-10m-internvl2-multi-caption.InternVL-SA-1B-Caption
Dataset Card for InternVL-SA-1B-Caption
Overview
The InternVL-SA-1B-Caption Dataset is a bilingual dataset created using the InternVL2-Llama3-76B model. The dataset contains 12 million image-caption pairs in both English and Chinese. All images are sourced from Meta’s SA-1B dataset, and captions were generated using specific prompts designed to minimize hallucinations and ensure accurate descriptions based on visible image content. The dataset is intended for use in tasks… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption.dreambench_eval_results_internvl2_5_78b_mpo_awq_init_1_prompt_collect_training_datainternvl3_5_eval-v4MemBench-InternVL3.5-Eval
MemBench-InternVL3.5-Eval
Evaluation dataset for image editing experiments on ppr10k, comparing four methods under the same selection protocol.Each method folder contains one dataset.jsonl and corresponding edited/source image pairs.
This repo is for reproduction and inspection-only purposes. To learn how to use it, visit the official codebase laitifranz/MemCoach#reproducing-paper-results.
[!NOTE]
Compact download: A single zip archive MemBench-InternVL3.5-Eval-Artifacts.zip… See the full description on the dataset page: https://huggingface.co/datasets/laitifranz/MemBench-InternVL3.5-Eval.InternVL-SA-1B-Caption-512robo2vlm-1-internvlRefCOCO-plus_testA-100_internvl3_5_detr_bboxonly_1231RefCOCO-g_testA-100_internvl3_5_detr_token_fpn_1231internvl3_5_eval-v3internvl3_5_eval-v1internvl3_5_eval-v0InternVL3_featinternvl_moneydatatwofranka_pick_and_place_12_9_internvlainternvl3_5_eval-v2Mono-InternVL-2B-Synthetic-Data
Mono-InternVL-2B Synthetic Data
This dataset is used for training the S1.2 stage of Mono-InternVL-2B, as described in the paper Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models.
Project Page: https://internvl.github.io/blog/2024-10-10-Mono-InternVL/
Code: https://github.com/OpenGVLab/Mono-InternVL
Dataset Description
Purpose
This dataset is used for training the S1.2 stage of Mono-InternVL-2B.
Data… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/Mono-InternVL-2B-Synthetic-Data.franka_pick_and_place_12_6_internvla
