CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wyu1 /Leopard-Instruct Leopard-Instruct Paper | Github | Models-LLaVA | Models-Idefics2 Summaries Leopard-Instruct is a large instruction-tuning dataset, comprising 925K instances, with 739K specifically designed for text-rich, multiimage scenarios. It's been used to train Leopard-LLaVA [checkpoint] and Leopard-Idefics2 [checkpoint]. Loading dataset to load the dataset without automatically downloading and process the images (Please run the following codes with datasets==2.18.0)… See the full description on the dataset page: https://huggingface.co/datasets/wyu1/Leopard-Instruct.image1M<n<10M64 likes193k downloads2y agoHugging Face02InnovatorLab /Innovator-VL-Instruct-46M Innovator-VL-Instruct-46M Paper | Code 🤗🤗 The data is being uploaded continuously Introduction To further enhance the model’s ability to handle a broad range of visual tasks with accurate, grounded, and instruction-aligned responses, we perform full-parameter visual instruction supervised fine-tuning (SFT).This SFT stage serves as a critical bridge between multimodal pretraining and subsequent reinforcement learning, providing both general capability coverage and a… See the full description on the dataset page: https://huggingface.co/datasets/InnovatorLab/Innovator-VL-Instruct-46M.imageimage-text-to-text10M<n<100M9 likes67k downloads8mo agoHugging Face03bigcode /self-oss-instruct-sc2-exec-filter-50kFinal self-alignment training dataset for StarCoder2-Instruct. seed: Contains the seed Python function concepts: Contains the concepts generated from the seed instruction: Contains the instruction generated from the concepts response: Contains the execution-validated response to the instruction This dataset utilizes seed Python functions derived from the MultiPL-T pipeline. text10K<n<100K108 likes41k downloads2y agoHugging Face04iamtarun /python_code_instructions_18k_alpaca Dataset Card for python_code_instructions_18k_alpaca The dataset contains problem descriptions and code in python language. This dataset is taken from sahil2801/code_instructions_120k, which adds a prompt column in alpaca style. Refer to the source here. textquestion-answering10K<n<100K349 likes37k downloads3y agoHugging Face05DeepStudentLlama /AoPS-InstructReproduction of AoPS-Instruct training set using code here: https://github.com/DSL-Lab/aops text1M<n<10M18 likes31k downloads2y agoHugging Face06haitengzhao /molecule_property_instruction Dataset Card for "molecule_property_instruction" More Information needed textquestion-answering10M<n<100M20 likes23k downloads3y agoHugging Face07allenai /tulu-3-sft-personas-instruction-following Dataset Descriptions This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset. To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper. Curated by: Allen Institute for AI Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.texttext-generation10K<n<100K68 likes15k downloads2y agoHugging Face08HuggingFaceH4 /testing_self_instruct_small Dataset Card for "testing_self_instruct_small" More Information needed textn<1K2 likes14k downloads3y agoHugging Face09ziyjiang /MMEB_Test_Instructtext10K<n<100K0 likes12k downloads2y agoHugging Face10mlabonne /Evol-Instruct-Python-26k Evol-Instruct-Python-26k Filtered version of the nickrosh/Evol-Instruct-Code-80k-v1 dataset that only keeps Python code (26,588 samples). You can find a smaller version of it here mlabonne/Evol-Instruct-Python-1k. Here is the distribution of the number of tokens in each row (instruction + output) using Llama's tokenizer: text10K<n<100K15 likes8.6k downloads3y agoHugging Face11livebench /instruction_following Dataset Card for "livebench/instruction_following" LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties: LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses. Each question has verifiable, objective ground-truth answers, allowing hard questions… See the full description on the dataset page: https://huggingface.co/datasets/livebench/instruction_following.textn<1K6 likes8k downloads1y agoHugging Face12TIGER-Lab /Mantis-Instruct Mantis-Instruct Paper | Website | Github | Models | Demo Summaries Mantis-Instruct is a fully text-image interleaved multimodal instruction tuning dataset, containing 721K examples from 14 subsets and covering multi-image skills including co-reference, reasoning, comparing, temporal understanding. It's been used to train Mantis Model families Mantis-Instruct has a total of 721K instances, consisting of 14 subsets to cover all the multi-image skills. Among the… See the full description on the dataset page: https://huggingface.co/datasets/TIGER-Lab/Mantis-Instruct.text100K<n<1M41 likes7.9k downloads2y agoHugging Face13ackermans26 /LLaVA-OneVision-1.5-Instruct-Data-qwen-formattext1M<n<10M0 likes6.8k downloads2mo agoHugging Face14SaylorTwift /details_meta-llama__Llama-3.1-8B-Instruct_private Dataset Card for Evaluation run of meta-llama/Llama-3.1-8B-Instruct Dataset automatically created during the evaluation run of model meta-llama/Llama-3.1-8B-Instruct. The dataset is composed of 78 configuration, each one corresponding to one of the evaluated task. The dataset has been created from 20 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/SaylorTwift/details_meta-llama__Llama-3.1-8B-Instruct_private.textn<1K0 likes4.5k downloads1y agoHugging Face15timbrooks /instructpix2pix-clip-filtered Dataset Card for InstructPix2Pix CLIP-filtered Dataset Summary The dataset can be used to train models to follow edit instructions. Edit instructions are available in the edit_prompt. original_image can be used with the edit_prompt and edited_image denotes the image after applying the edit_prompt on the original_image. Refer to the GitHub repository to know more about how this dataset can be used to train a model that can follow instructions. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/timbrooks/instructpix2pix-clip-filtered.image100K<n<1M48 likes4.3k downloads4y agoHugging Face16lingamvamshikrishnareddy /ramanv-image-vlm-instructiontext100K<n<1M0 likes3.6k downloads24d agoHugging Face17HuggingFaceH4 /helpful-instructions Dataset Card for Helpful Instructions Dataset Summary Helpful Instructions is a dataset of (instruction, demonstration) pairs that are derived from public datasets. As the name suggests, it focuses on instructions that are "helpful", i.e. the kind of questions or tasks a human user might instruct an AI assistant to perform. You can load the dataset as follows: from datasets import load_dataset # Load all subsets helpful_instructions =… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/helpful-instructions.text100K<n<1M24 likes3.4k downloads4y agoHugging Face18Qdrant /arxiv-titles-instructorxl-embeddings arxiv-titles-instructorxl-embeddings This dataset contains 768-dimensional embeddings generated from the arxiv paper titles using InstructorXL model. Each vector has an abstract used to create it, along with the DOI (Digital Object Identifier). The dataset was created using precomputed embeddings exposed by the Alexandria Index. Generation process The embeddings have been generated using the following instruction: Represent the Research Paper title for retrieval;… See the full description on the dataset page: https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings.textsentence-similarity1M<n<10M5 likes3.2k downloads3y agoHugging Face19BAAI /Infinity-Instructgated Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and high-quality instruction… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Infinity-Instruct.tabulartext-generation10M<n<100M765 likes2.9k downloads10mo agoHugging Face20allenai /Dolci-Instruct-SFT Dolci Instruct SFT Mixture Note that this collection licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. The Dolci Instruct SFT mixture was used to train Olmo 3 7B Instruct SFT. It contains 2,152,112 samples from the following sets: Sources include a mixture of existing prompts: OpenThoughts 3 (Apache 2.0): Extended to 32K context length and downsampled code prompts to 16X multiple, to 941,166 total prompts… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT.textother1M<n<10M62 likes2.8k downloads8mo agoHugging Face21aisingapore /Instruction-Following-IFEvalgated SEA-IFEval SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. It is based on IFEval and was manually translated by native speakers for Indonesian, Javanese, Sundanese, Thai, Tagalog, and Vietnamese. Supported Tasks and Leaderboards SEA-IFEval is designed for evaluating chat or instruction-tuned large language models (LLMs).… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/Instruction-Following-IFEval.texttext-generation1K<n<10K0 likes2.6k downloads9mo agoHugging Face22theblackcat102 /llava-instruct-mix LLaVA Instruct Mix Added OCR and Chart QA dataset into this for more text extraction questions imagevisual-question-answering100K<n<1M12 likes2.5k downloads3y agoHugging Face23OALL /details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge Dataset Card for Evaluation run of grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge Dataset automatically created during the evaluation run of model grimjim/Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge. The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_grimjim__Llama-3-Instruct-8B-SimPO-SPPO-Iter3-merge.tabular100K<n<1M0 likes2.2k downloads2y agoHugging Face24tasksource /tasksource-instruct-v0 Dataset Card for "tasksource-instruct-v0" (TSI) Multi-task instruction-tuning data recasted from 485 of the tasksource datasets. Dataset size is capped at 30k examples per task to foster task diversity. !pip install tasksource, pandit import tasksource, pandit df = tasksource.list_tasks(instruct=True).sieve(id=lambda x: 'mmlu' not in x) for tasks in df.id: yield tasksource.load_task(task,instruct=True,max_rows=30_000,max_rows_eval=200) https://github.com/sileod/tasksource… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/tasksource-instruct-v0.texttext-generation1M<n<10M24 likes2k downloads3mo agoHugging Face25allenai /Dolci-Instruct-DPO Dolci Instruct DPO Mixture This dataset is licensed under ODC-BY. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines. The Dolci Instruct DPO mixture was used to preference tune Olmo 3 Instruct 7B. It contains 260,000 preference pairs in total, including: 125,000 pairs created with the preference heuristic described in Delta Learning (Geng et al. 2025) 125,000 pairs created with a delta-aware Ultrafeedback-esque GPT-judge pipeline… See the full description on the dataset page: https://huggingface.co/datasets/allenai/Dolci-Instruct-DPO.text100K<n<1M16 likes2k downloads7mo agoHugging Face26nicholasKluge /Pt-Corpus-Instruct Portuguese-Corpus Instruct Dataset Summary Portuguese-Corpus Instruct is a concatenation of several portions of Brazilian Portuguese datasets found in the Hub. In a tokenized format, the dataset (uncompressed) weighs 80 GB and has approximately 6.2B tokens. This version of the corpus (Pt-Corpus-Instruct) includes several instances of conversational and general instructional data, allowing trained models to go through preference pre-training during their initial… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/Pt-Corpus-Instruct.texttext-generation10M<n<100M3 likes2k downloads2y agoHugging Face27HuggingFaceH4 /llava-instruct-mix-vsfttheblackcat102/llava-instruct-mix reformated for VSFT with TRL's SFT Trainer. See https://github.com/huggingface/trl/blob/main/examples/scripts/vsft_llava.py. image100K<n<1M49 likes2k downloads2y agoHugging Face28manifoldlabs /Infinity-Instruct Infinity Instruct Beijing Academy of Artificial Intelligence (BAAI) [Paper][Code][🤗] (would be released soon) The quality and scale of instruction data are crucial for model performance. Recently, open-source models have increasingly relied on fine-tuning datasets comprising millions of instances, necessitating both high quality and large scale. However, the open-source community has long been constrained by the high costs associated with building such extensive and… See the full description on the dataset page: https://huggingface.co/datasets/manifoldlabs/Infinity-Instruct.texttext-generation10M<n<100M5 likes1.9k downloads2y agoHugging Face29ko-vlm /KoLLaVA-v1.5-Instruct-581k KoLLaVA-v1.5-Instruct-581k 한국어 Vision-Language 모델을 위한 instruction tuning 데이터셋입니다. 데이터셋 정보 총 샘플 수: 435,093개 형식: ChatML 형식 (role: user/assistant, content: 텍스트) 이미지: COCO + GQA + Visual Genome 데이터셋 언어: 한국어 포함된 데이터셋 COCO 데이터: 362,953개 샘플 MS COCO 2017 이미지 기반 한국어 대화 데이터 GQA 데이터: 72,140개 샘플 GQA (Visual Question Answering) 이미지 기반 한국어 대화 데이터 Visual Genome 데이터: 포함 Visual Genome 이미지 기반 한국어 대화 데이터 제외된 데이터셋 EKVQA 데이터: AI Hub 라이선스로 인해 공개 불가… See the full description on the dataset page: https://huggingface.co/datasets/ko-vlm/KoLLaVA-v1.5-Instruct-581k.image100K<n<1M0 likes1.7k downloads1y agoHugging Face30ENSEONG /full-math-private-n256-Qwen2.5-3B-Instruct-bontabular100K<n<1M0 likes1.7k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.