datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-gpt4
Dataset Card for alpaca-gpt4
This dataset originates from this repository.
The alpaca-gpt4 dataset is specifically used for fine-tuning LLMs based on the instruction generated by GPT-4 using Alpaca prompts.
Dataset Details
Dataset Description
Each sample is comprised of four columns: instruction, input, output and text.
Language(s): English
License: Creative Commons NonCommercial (CC BY-NC 4.0)
Dataset Sources
The code from the original repository… See the full description on the dataset page: https://huggingface.co/datasets/flwrlabs/alpaca-gpt4.alpaca-gpt4
Dataset Card for "alpaca-gpt4"
This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/vicgalle/alpaca-gpt4.GPT-4o-evaluation-biases
A database to support the evaluation of gender biases in GPT-4o output
The database and its construction process are described in the paper "A database to support the evaluation of gender biases in GPT-4o output" by Mehner et al., presented at the 1st ISCA/ITG Workshop on Diversity in Large Speech and Language Models (Berlin, Februar 20, 2025).
Introduction
This is a database of prompts and answers generated with GPT-4o-mini and GPT-4o in a pretest and a main test… See the full description on the dataset page: https://huggingface.co/datasets/mtec-TUB/GPT-4o-evaluation-biases.alpaca_gpt4_zhBorrowed from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
Removed 6,103 mistruncated examples.
You can use it in LLaMA Factory by specifying dataset: alpaca_gpt4_zh.
alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Ichsan2895/alpaca-gpt4-indonesian.distill-gpt4-eng-chat
Description
Introducing dataset consisting of gpt4 answers to users requests. Queries were taken from allenai/WildChat-1M and causal-lm/instructions. Texts (requests and responses) were deleted in 3 cases:
either has non-english letters and special symbols
either has http-links
either has html blocks
either has perplexity more than 1.5*IQR + third quantile ( in some cases average perplexity value of sentences or maximum value was used )
GPT-4-PromptsMulti-Turn Conversational Prompts from ChatGPT-4 (10K+ Tokens)
Abstract:
This dataset offers a valuable collection of multi-turn conversational prompts generated by ChatGPT-4, carefully curated for diverse prompt styles (chatml, gemma, llama). Each prompt exceeds 10,000 tokens, providing ample context and inspiration for training and evaluating large language models. Ideal for researchers and developers interested in exploring advanced conversational AI capabilities.
Table of Contents:… See the full description on the dataset page: https://huggingface.co/datasets/erfanzar/GPT-4-Prompts.HuatuoGPT2-SFT-GPT4-140K
HuatuoGPT2-SFT-GPT4-140K
140K Chinese medical instructions generated by GPT-4, based on questions from HuatuoGPT Dataset.
This dataset contains supervised fine-tuning instructions for HuatuoGPT2, designed to enhance the model's ability to follow instructions in real medical scenarios. We have made all the data (142,248 entries) in this dataset publicly available.
Repository
Github: https://github.com/FreedomIntelligence/HuatuoGPT-II
Citation… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/HuatuoGPT2-SFT-GPT4-140K.lightblue-tagengo-gpt4
lightblue/tagengo-gpt4
An unofficial, reformatted version of lightblue/tagengo-gpt4.
Tagengo is described by its author as the world's largest high-quality multilingual chat dataset - containing over 75,000 single-turn conversations between humans and GPT‑4 (gpt-4-0125-preview) across 74 languages. It fills a major gap in multilingual chat data, which has so far been limited compared to English.
Additional Processing
Split by language
Kept only entries with both… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/lightblue-tagengo-gpt4.alpaca_gpt4_enBorrowed from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
You can use it in LLaMA Factory by specifying dataset: alpaca_gpt4_en.
MMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.legal-contract-gpt41-redlining-10k
legal-contract-gpt41-redlining-10k
Dataset Description
This dataset contains 9977 synthetic legal contract redlines generated using GPT-4.1 model mix (base, mini, nano) with structured outputs. It is designed for fine-tuning LLMs (including OpenAI GPT-3.5/4, Llama, and other models) to assist with legal document redlining and clause revision.
Key Features
🤖 9977 synthetic redlines generated by GPT-4.1 model mix (base, mini, nano)
📋 Multiple training formats:… See the full description on the dataset page: https://huggingface.co/datasets/UmaiTech/legal-contract-gpt41-redlining-10k.Vietnamese-alpaca-gpt4-gg-translatedpokemon-gpt4o-captionsBorrowed from: https://huggingface.co/datasets/jugg1024/pokemon-gpt4o-captions
You can use it in LLaMA Factory by specifying dataset: pokemon_cap.
Magpie-Pro-10K-GPT4o-minialpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/Faishal-Anwar/alpaca-gpt4-indonesian.alpaca-data-gpt4-chinese-zhtw
Dataset Card for "alpaca-data-gpt4-chinese-zhtw"
This dataset contains Chinese (zh-tw) Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This dataset is a translation from English to Chinese.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/alpaca-data-gpt4-chinese-zhtw.alpaca-gpt4-indonesianBase model : FreedomIntelligence/alpaca-gpt4-indonesian
We wrangled the original dataset format to 'input' & 'output' format. For example:
BEFORE:
[ { "from": "human",
"value": "Saranlah slogan untuk kampanye daur ulang\n" },
{ "from": "gpt",
"value": "1. \"Kurangi, gunakan kembali, daur ulang: Bersama untuk masa depan yang lebih hijau.\"\n2. \
"Daur ulanglah hari ini, untuk masa depan yang lebih baik.\"\n3. \"Ubah sampahmu menjadi harta karun - Daur ulang!\"\n4. \
"Daur… See the full description on the dataset page: https://huggingface.co/datasets/glhpradipta/alpaca-gpt4-indonesian.alpaca-gpt4-hun
Dataset Card for "alpaca-gpt4"
This dataset contains Hungarian (translated from English) Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs. Original model: https://huggingface.co/datasets/vicgalle/alpaca-gpt4
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K… See the full description on the dataset page: https://huggingface.co/datasets/Bazsalanszky/alpaca-gpt4-hun.alpaca_gpt4_dialogue_vi
Description
The dataset is from 5CD-AI/Vietnamese-c-s-ale-alpaca-gpt4-data-gg-translated, formatted as dialogues for speed and ease of use. Many thanks to 5CD-AI for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/alpaca_gpt4_dialogue_vi.alpaca_gpt4_dialogue_en
Description
The dataset is from 5CD-AI/Vietnamese-c-s-ale-alpaca-gpt4-data-gg-translated, formatted as dialogues for speed and ease of use. Many thanks to 5CD-AI for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/alpaca_gpt4_dialogue_en.alpaca-gpt4
Dataset Card for "alpaca-gpt4"
This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/foxespri/alpaca-gpt4.alpaca-gpt4
Dataset Card for "alpaca-gpt4"
This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has the… See the full description on the dataset page: https://huggingface.co/datasets/doyancha/alpaca-gpt4.GPT-4ualpaca-gpt4
Dataset Card for "alpaca-gpt4-cleaned"
This dataset contains Ukrainian Instruction-Following translated by facebook/nllb-200-3.3B
The dataset was originaly shared in this repository: https://github.com/tloen/alpaca-lora
Licensing Information
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
ResourceGuard-gpt4Created by: TrungKien Dang, with research all source from internet
gpt-4o-qa
Quardo/gpt-4o-qa
Description
This dataset is generated by OpenAI's GPT-4O (gpt-4o-2024-08-06). It includes a large collection of question-answer pairs created and evaluated by the AI model, providing a comprehensive resource for various natural language processing tasks.
Warning
Please note that this dataset may contain errors or inconsistencies as it is fully generated by an AI model. It is highly recommended to check and edit the data before usage, as AI can… See the full description on the dataset page: https://huggingface.co/datasets/Quardo/gpt-4o-qa.gpt4-prompts
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/gooberdooberloober/gpt4-prompts.ro-alpaca-gpt4This dataset is the translated vicgalle/alpaca-gpt4 instruct dataset using LLMic, a bilingual Romanian-English LLM.
The alpaca-gpt4 is an English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset is available under the Creative Commons NonCommercial (CC BY-NC 4.0).
@article{peng2023instruction,
title={Instruction Tuning with GPT-4},
author={Peng, Baolin and Li, Chunyuan and He, Pengcheng and Galley, Michel and Gao, Jianfeng},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/faur-ai/ro-alpaca-gpt4.alpaca-gpt4-mat
Dataset Card for "alpaca-gpt4"
This dataset contains English Instruction-Following generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
The dataset was originaly shared in this repository: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM. This is just a wraper for compatibility with huggingface's datasets library.
Dataset structure
It contains 52K instruction-following data generated by GPT-4 using the same prompts as in Alpaca.
The dataset has the same… See the full description on the dataset page: https://huggingface.co/datasets/ismatazmy/alpaca-gpt4-mat.
