datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MAmmoTH-VL-Instruct-12M
MAmmoTH-VL-Instruct-12M
🏠 Homepage | 🤖 MAmmoTH-VL-8B | 💻 Code | 📄 Arxiv | 📕 PDF | 🖥️ Demo
Introduction
Our simple yet scalable visual instruction data rewriting pipeline consists of three steps: manual data source collection, rewriting using MLLMs/LLMs, and filtering via the same MLLM as a judge. Examples below illustrate transformations in math and science categories, showcasing detailed, step-by-step responses.
The data distribution of… See the full description on the dataset page: https://huggingface.co/datasets/MAmmoTH-VL/MAmmoTH-VL-Instruct-12M.ET-Instruct-164K
E.T. Instruct 164K
arXiv | Project Page | GitHub
E.T. Instruct 164K is a large-scale instruction-tuning dataset tailored for fine-grained event-level and time-sensitive video understanding. It contains 101K meticulously collected videos under diverse domains and 9 event-level understanding tasks with well-designed instruction-response pairs. The average video length is around 146 seconds.
📦 Download Dataset
You may download the dataset using the following command.… See the full description on the dataset page: https://huggingface.co/datasets/PolyU-ChenLab/ET-Instruct-164K.llava-1.5-665k-instructionsThis dataset repository, LLaVA-1.5-665K-Instructions, is notably utilized in the paper Zero-Shot Vision Encoder Grafting via LLM Surrogates.
The official code repository for the paper can be found here: https://github.com/kaiyuyue/zero
LLaVA-1.5-665K-Instructions
This dataset repo contains the entire LLaVA-1.5-665K-Instructions in one place, including images and text sequences.
The images are in train_split/*.tars and the text sequences are in jsons:
llava_v1_5_mix665k.json is the… See the full description on the dataset page: https://huggingface.co/datasets/kaiyuyue/llava-1.5-665k-instructions.ureader-instruction-1.0open-lm-instruction-dataInnovator-VL-Instruct-Sciencefivl-instruct
FiVL-Instruct Dataset
FiVL: A Frameword for Improved Vision-Language Alignment introduces grounded datasets for both training and evaluation, building upon existing vision-question-answer and instruction datasets
Each sample in the original datasets was augmented with key expressions, along with their corresponding bounding box indices and segmentation masks within the images.
Dataset Details
Creators: Intel Labs
Version: 1.0 (Updated: 2024-12-18)
License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/Intel/fivl-instruct.GPQA_verifications_GenRM-Base_Llama-3.3-70B-InstructMATH128_verifications_GenRM-FT_Llama-3.1-8B-InstructMATH128_verifications_GenRM-FT_Qwen-2.5-7B-InstructChinese-Dialogue-180k-Instruct-AudioMATH128_Solutions_Llama-3.1-8B-InstructMATH128_Solutions_Qwen-2.5-7B-InstructMATH128_Solutions_Llama-3.3-70B-InstructX2I-instruct
X2I-instruct (WebP-compressed)
This dataset is a WebP q=80 re-encoded version of yzwang/X2I-subject-driven, packed into plain .tar shards (each <= 30 GiB).
All images (PNG / JPEG / WebP) were re-encoded as WebP at quality 80.
Total size shrunk from ~1.79 TB -> ~140 GB (~13x compression).
JSONL metadata files are rewritten to point at the new .webp paths (see *.webp.jsonl).
All instructions, sample structure, and the per-subdir directory layout are preserved.
Layout… See the full description on the dataset page: https://huggingface.co/datasets/ymhao/X2I-instruct.M4-Instruct-Multi0504_combination_instruction_wikihowMATH128_verifications_Llama-3.3-70B-Instruct_GenRM-Baseinstructs2s-webdatasetlcb128_llama3-8B-instruct_256samples_ver32_temp0-7aime24_qwen2.5-7b_ver_Llama-3.1-8B-Instruct_data-qwen_25_7b_gpt_4o_verify_train_e3_LR-5e-7_7Klenmmlu-chat-Llama-3.2-Instructbenchmark-chat-Llama-3.2-InstructInstructIR
