datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.alpaca-cleaned-ru
alpaca-cleaned-ru
Translated version of yahma/alpaca-cleaned into Russian.
alpaca-cleaned-pt
Data Description
This HF data repository contains the Portuguese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Portuguese.
Usage
This data is intended to be used for Portuguese instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-pt.alpaca-cleaned-italian
Dataset Card for Alpaca-Cleaned-Italian
About the translation and the original data
The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here).
The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English.
Additional notes on the translation
Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.alpaca_cleaned_ja_json
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.alpaca-id-cleaned
Dataset Card for Indonesian Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is the Indonesian translated version of the cleaned original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an… See the full description on the dataset page: https://huggingface.co/datasets/cahya/alpaca-id-cleaned.alpaca-cleaned-es
Data Description
This HF data repository contains the Spanish Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Spanish.
Usage
This data is intended to be used for Spanish instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-es.alpaca-cleaned-dutch
Dataset Card for Alpaca Cleaned Dutch
Dataset Summary
This dataset contains 51,712 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch. They are translations of Alpaca Cleaned Dataset.
☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$57.99💸. If you like this dataset, please consider buying me a coffee to offset a portion of this cost, I appreciate it a lot! ☕
If you use this dataset or refer to… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/alpaca-cleaned-dutch.marathi-alpaca-cleaned-translated
Marathi Alpaca Cleaned Translated
A Marathi translation of the 51,760-row Alpaca-Cleaned instruction-tuning dataset — Unsloth's hosted fork of yahma/alpaca-cleaned, which fixes hallucinations, empty outputs, and formatting errors found in the original Stanford Alpaca-52k dataset.
Translated using Meta's facebook/nllb-200-distilled-600M model. Built to reproduce and evaluate the Marathi instruction-tuning experiment from Khade et al., CHiPSAL 2025. The original paper translated… See the full description on the dataset page: https://huggingface.co/datasets/lubzo/marathi-alpaca-cleaned-translated.alpaca-cleaned-cs
Data Description
This HF data repository contains the Czech Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Czech.
Usage
This data is intended to be used for Czech instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-cs.alpaca-cleaned-indonesian
🦙🛁 Cleaned Alpaca Dataset (INDONESIAN)
Welcome to the Cleaned Alpaca Dataset repository! This repository hosts a cleaned and curated version of a dataset used to train the Alpaca LLM (Large Language Model). The original dataset had several issues that are addressed in this cleaned version.
On April 8, 2023 the remaining uncurated instructions (~50,000) were replaced with data from the GPT-4-LLM dataset. Curation of the incoming GPT-4 data is ongoing.
A 7b Lora model (trained on… See the full description on the dataset page: https://huggingface.co/datasets/ilhamfadheel/alpaca-cleaned-indonesian.alpaca-cleaned-fr
Data Description
This HF data repository contains the French Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into French.
Usage
This data is intended to be used for French instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-fr.alpaca-cleaned-ru
Data Description
This HF data repository contains the Russian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Russian.
Usage
This data is intended to be used for Russian instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru.stanford-alpaca-cleaned-turkish-translated09/04/2023 Update:
New instructions added from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM
Original Version: https://github.com/tatsu-lab/stanford_alpaca#data-release
AI BASED TRANSLATION RESULTS OF STANFORD ALPACA EN TO TR
For academic only, please cite before you use it.
Taşar, D. E. T. (2023). stanford-alpaca-cleaned-turkish-translated [Dataset]. In Stanford Alpaca TR (1.0.1.a). https://huggingface.co/datasets/emre/stanford-alpaca-cleaned-turkish-translated… See the full description on the dataset page: https://huggingface.co/datasets/emre/stanford-alpaca-cleaned-turkish-translated.Alpaca_Evol_Instruct_CleanedAlpaca Evol Instruct cleaned of refusals, scrubbed of overly repetitive responses, aggresively deduplicated, and all URLs removed from the output. The final dataset has aproximately 54k instructions.
Base dataset https://huggingface.co/datasets/victor123/evol_instruct_70k
alpaca-cleaned-ru
alpaca-cleaned-ru
converter for autotrain from d0rj/alpaca-cleaned-ru
Translated version of yahma/alpaca-cleaned into Russian.
alpaca-cleaned-bg
Data Description
This HF data repository contains the Bulgarian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Bulgarian.
Usage
This data is intended to be used for Bulgarian instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-bg.alpaca-cleaned-de
Data Description
This HF data repository contains the German Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into German.
Usage
This data is intended to be used for German instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-de.alpaca-cleaned-fi
Data Description
This HF data repository contains the Finnish Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Finnish.
Usage
This data is intended to be used for Finnish instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-fi.alpaca-cleaned-serbian-full
Serbian Alpaca Cleaned Dataset
Original Repository: https://github.com/gururise/AlpacaDataCleaned
Original HF Repository: https://huggingface.co/datasets/yahma/alpaca-cleaned
Dataset Description
This is a serbian cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the… See the full description on the dataset page: https://huggingface.co/datasets/datatab/alpaca-cleaned-serbian-full.alpaca-cleaned-zh
Data Description
This HF data repository contains the Chinese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Chinese.
Usage
This data is intended to be used for Chinese instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-zh.alpaca-cleaned-hinglish
Alpaca Cleaned (Hinglish Version)
Dataset Description
This is a high-quality Hinglish (Hindi written in Latin script) translation of the yahma/alpaca-cleaned dataset.
It is designed for instruction fine-tuning large language models to make them conversational in Indian contexts.
Dataset Summary
Original Source: yahma/alpaca-cleaned (51,760 rows)
Language: Hinglish (Code mixed Hindi-English)
Translation Method: High-precision batch translation using… See the full description on the dataset page: https://huggingface.co/datasets/hbpkillerX/alpaca-cleaned-hinglish.alpaca-cleaned-bn
Dataset Card for Alpaca-Cleaned-bn
This is a cleaned bengali translated version of the original Alpaca Dataset released by Stanford.
Uses
import datasets
dataset = datasets.load_dataset("abrarfahim/alpaca-cleaned-bn")
print(dataset[0])
Dataset Structure
{'system_prompt': 'You are a virtual assistant, deliver a comprehensive response.',
'qas_id': 'YY9S5K',
'question_text': '"সন্দেহ" শব্দের সঠিক প্রতিশব্দ নির্বাচন করুন।',
'orig_answer_texts': '"সন্দেহ"… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/alpaca-cleaned-bn.alpaca-cleaned
Dataset Card for Alpaca-Cleaned
Repository: https://github.com/gururise/AlpacaDataCleaned
Dataset Description
This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset:
Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer.
"instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/melvindave/alpaca-cleaned.alpaca-cleaned-zh-cn
Data Description
This HF data repository contains the Chinese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning.
GitHub
Paper
Creation
Machine-translated from yahma/alpaca-cleaned into Chinese.
Usage
This data is intended to be used for Chinese instruction tuning.
The dataset has roughly 52K instances in the JSON format.
Each instance has an instruction, an output, and an optional input. An example is shown below:
{… See the full description on the dataset page: https://huggingface.co/datasets/leo009/alpaca-cleaned-zh-cn.alpaca-cleaned-fralpaca-cleaned-chatml
ChatML Reformat of yahma/alpaca-cleaned
I'd like to try instruction-tuning dataset with chat-tuning format.
Usage
dataset = load_dataset("pacozaa/alpaca-cleaned-chatml", split = "train")
print(dataset[0]["text"])
QA-text-generation-alpaca-data-cleaned
Dataset Card for mBART-QA-Processed
This dataset consists of tokenized pairs of instructions and contexts designed for fine-tuning Sequence-to-Sequence models (like mBART or T5) on Question Answering tasks.
Dataset Details
Dataset Description
The dataset is a processed version of a Question Answering corpus (SQuAD-like). It has been formatted to follow a specific prompt structure: instruction: {question} input: {context}. The targets (labels) are the direct… See the full description on the dataset page: https://huggingface.co/datasets/SOULAMA/QA-text-generation-alpaca-data-cleaned.alpaca-cleaned_uz
Dataset Summary
This dataset is a translation of the alpaca-cleaned dataset into Uzbek (Latin), using the GPT-4o mini API.
