datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Leopard-Instruct
Leopard-Instruct
Paper | Github | Models-LLaVA | Models-Idefics2
Summaries
Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [checkpoint] and Leopard-Idefics2 [checkpoint].
Loading dataset
to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0)… See the full description on the dataset page: https://huggingface.co/datasets/wyu1/Leopard-Instruct.Innovator-VL-Instruct-46M
Innovator-VL-Instruct-46M
Paper | Code
🤗🤗 The data is being uploaded continuously
Introduction
To further enhance the model’s ability to handle a broad range of visual tasks with accurate, grounded, and instruction-aligned responses, we perform full-parameter visual instruction supervised fine-tuning (SFT).This SFT stage serves as a critical bridge between multimodal pretraining and subsequent reinforcement learning, providing both general capability coverage and a… See the full description on the dataset page: https://huggingface.co/datasets/InnovatorLab/Innovator-VL-Instruct-46M.self-oss-instruct-sc2-exec-filter-50kFinal self-alignment training dataset for StarCoder2-Instruct.
seed: Contains the seed Python function
concepts: Contains the concepts generated from the seed
instruction: Contains the instruction generated from the concepts
response: Contains the execution-validated response to the instruction
This dataset utilizes seed Python functions derived from the MultiPL-T pipeline.
python_code_instructions_18k_alpaca
Dataset Card for python_code_instructions_18k_alpaca
The dataset contains problem descriptions and code in python language.
This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here.
AoPS-InstructReproduction of AoPS-Instruct training set using code here: https://github.com/DSL-Lab/aops
molecule_property_instruction
Dataset Card for "molecule_property_instruction"
More Information needed
tulu-3-sft-personas-instruction-following
Dataset Descriptions
This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset.
To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper.
Curated by: Allen Institute for AI
Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.testing_self_instruct_small
Dataset Card for "testing_self_instruct_small"
More Information needed
MMEB_Test_InstructEvol-Instruct-Python-26k
Evol-Instruct-Python-26k
Filtered version of the nickrosh/Evol-Instruct-Code-80k-v1 dataset that only keeps Python code (26,588 samples). You can find a smaller version of it here mlabonne/Evol-Instruct-Python-1k.
Here is the distribution of the number of tokens in each row (instruction + output) using Llama's tokenizer:
instruction_following
Dataset Card for "livebench/instruction_following"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions… See the full description on the dataset page: https://huggingface.co/datasets/livebench/instruction_following.Mantis-Instruct
Mantis-Instruct
Paper | Website | Github | Models | Demo
Summaries
Mantis-Instruct is a fully text-image interleaved multimodal instruction tuning dataset,
containing 721K examples from 14 subsets and covering multi-image skills including co-reference, reasoning, comparing, temporal understanding.
It's been used to train Mantis Model families
Mantis-Instruct has a total of 721K instances, consisting of 14 subsets to cover all the multi-image skills.
Among the… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Instruct.LLaVA-OneVision-1.5-Instruct-Data-qwen-formatdetails_meta-llama__Llama-3.1-8B-Instruct_private
Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct
Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct.
The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task.
The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.instructpix2pix-clip-filtered
Dataset Card for InstructPix2Pix CLIP-filtered
Dataset Summary
The dataset can be used to train models to follow edit instructions. Edit instructions
are available in the edit_prompt. original_image can be used with the edit_prompt and
edited_image denotes the image after applying the edit_prompt on the original_image.
Refer to the GitHub repository to know more about
how this dataset can be used to train a model that can follow instructions.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.ramanv-image-vlm-instructionhelpful-instructions
Dataset Card for Helpful Instructions
Dataset Summary
Helpful Instructions is a dataset of (instruction, demonstration) pairs that are derived from public datasets. As the name suggests, it focuses on instructions that are "helpful", i.e. the kind of questions or tasks a human user might instruct an AI assistant to perform. You can load the dataset as follows:
from datasets import load_dataset
# Load all subsets
helpful_instructions =… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/helpful-instructions.arxiv-titles-instructorxl-embeddings
arxiv-titles-instructorxl-embeddings
This dataset contains 768-dimensional embeddings generated from the arxiv
paper titles using InstructorXL model. Each
vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The
dataset was created using precomputed embeddings exposed by the Alexandria Index.
Generation process
The embeddings have been generated using the following instruction:
Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.Infinity-Instruct
Infinity Instruct
Beijing Academy of Artificial Intelligence (BAAI)
[Paper][Code][🤗]
The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and high-quality instruction… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Infinity-Instruct.Dolci-Instruct-SFT
Dolci Instruct SFT Mixture
Note that this collection licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Instruct SFT mixture was used to train Olmo 3 7B Instruct SFT.
It contains 2,152,112 samples from the following sets:
Sources include a mixture of existing prompts:
OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT.Instruction-Following-IFEval
SEA-IFEval
SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. It is based on IFEval and was manually translated by native speakers for Indonesian, Javanese, Sundanese, Thai, Tagalog, and Vietnamese.
Supported Tasks and Leaderboards
SEA-IFEval is designed for evaluating chat or instruction-tuned large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Instruction-Following-IFEval.llava-instruct-mix
LLaVA Instruct Mix
Added OCR and Chart QA dataset into this for more text extraction questions
details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge
Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.tasksource-instruct-v0
Dataset Card for "tasksource-instruct-v0" (TSI)
Multi-task instruction-tuning data recasted from 485 of the tasksource datasets.
Dataset size is capped at 30k examples per task to foster task diversity.
!pip install tasksource, pandit
import tasksource, pandit
df = tasksource.list_tasks(instruct=True).sieve(id=lambda x: 'mmlu' not in x)
for tasks in df.id:
yield tasksource.load_task(task,instruct=True,max_rows=30_000,max_rows_eval=200)
https://github.com/sileod/tasksource… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct-v0.Dolci-Instruct-DPO
Dolci Instruct DPO Mixture
This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
The Dolci Instruct DPO mixture was used to preference tune Olmo 3 Instruct 7B. It contains 260,000 preference pairs in total, including:
125,000 pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025)
125,000 pairs created with a delta-aware Ultrafeedback-esque GPT-judge pipeline… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-DPO.Pt-Corpus-Instruct
Portuguese-Corpus Instruct
Dataset Summary
Portuguese-Corpus Instruct is a concatenation of several portions of Brazilian Portuguese datasets found in the Hub.
In a tokenized format, the dataset (uncompressed) weighs 80 GB and has approximately 6.2B tokens. This version of the corpus (Pt-Corpus-Instruct) includes several instances of conversational and general instructional data, allowing trained models to go through preference pre-training during their initial… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct.llava-instruct-mix-vsfttheblackcat102/llava-instruct-mix reformated for VSFT with TRL's SFT Trainer.
See https://github.com/huggingface/trl/blob/main/examples/scripts/vsft_llava.py.
Infinity-Instruct
Infinity Instruct
Beijing Academy of Artificial Intelligence (BAAI)
[Paper][Code][🤗] (would be released soon)
The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and… See the full description on the dataset page: https://huggingface.co/datasets/manifoldlabs/Infinity-Instruct.KoLLaVA-v1.5-Instruct-581k
KoLLaVA-v1.5-Instruct-581k
한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다.
데이터셋 정보
총 샘플 수: 435,093개
형식: ChatML 형식 (role: user/assistant, content: 텍스트)
이미지: COCO + GQA + Visual Genome 데이터셋
언어: 한국어
포함된 데이터셋
COCO 데이터: 362,953개 샘플
MS COCO 2017 이미지 기반
한국어 대화 데이터
GQA 데이터: 72,140개 샘플
GQA (Visual Question Answering) 이미지 기반
한국어 대화 데이터
Visual Genome 데이터: 포함
Visual Genome 이미지 기반
한국어 대화 데이터
제외된 데이터셋
EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.full-math-private-n256-Qwen2.5-3B-Instruct-bon
