CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yahma /alpaca-cleaned Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer. "instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.texttext-generation10K<n<100K888 likes24k downloads3y agoHugging Face02unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K25 likes3.4k downloads9mo agoHugging Face03d0rj /alpaca-cleaned-ru alpaca-cleaned-ru Translated version of yahma/alpaca-cleaned into Russian. texttext-generation10K<n<100K22 likes332 downloads3y agoHugging Face04pinzhenchen /alpaca-cleaned-pt Data Description This HF data repository contains the Portuguese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Portuguese. Usage This data is intended to be used for Portuguese instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-pt.texttext-generation10K<n<100K5 likes153 downloads3y agoHugging Face05DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes148 downloads2y agoHugging Face06shi3z /alpaca_cleaned_ja_json Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.texttext-generation100K<n<1M13 likes111 downloads3y agoHugging Face07cahya /alpaca-id-cleaned Dataset Card for Indonesian Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is the Indonesian translated version of the cleaned original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an… See the full description on the dataset page: https://huggingface.co/datasets/cahya/alpaca-id-cleaned.texttext-generation10K<n<100K9 likes108 downloads3y agoHugging Face08pinzhenchen /alpaca-cleaned-es Data Description This HF data repository contains the Spanish Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Spanish. Usage This data is intended to be used for Spanish instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-es.texttext-generation10K<n<100K4 likes85 downloads3y agoHugging Face09BramVanroy /alpaca-cleaned-dutch Dataset Card for Alpaca Cleaned Dutch Dataset Summary This dataset contains 51,712 conversations between een AI assistant and a (fake) "Human" (generated) in Dutch. They are translations of Alpaca Cleaned Dataset. ☕ Want to help me out? Translating the data with the OpenAI API, and prompt testing, cost me 💸$57.99💸. If you like this dataset, please consider buying me a coffee to offset a portion of this cost, I appreciate it a lot! ☕ If you use this dataset or refer to… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/alpaca-cleaned-dutch.textquestion-answering10K<n<100K10 likes73 downloads3y agoHugging Face10lubzo /marathi-alpaca-cleaned-translated Marathi Alpaca Cleaned Translated A Marathi translation of the 51,760-row Alpaca-Cleaned instruction-tuning dataset — Unsloth's hosted fork of yahma/alpaca-cleaned, which fixes hallucinations, empty outputs, and formatting errors found in the original Stanford Alpaca-52k dataset. Translated using Meta's facebook/nllb-200-distilled-600M model. Built to reproduce and evaluate the Marathi instruction-tuning experiment from Khade et al., CHiPSAL 2025. The original paper translated… See the full description on the dataset page: https://huggingface.co/datasets/lubzo/marathi-alpaca-cleaned-translated.texttext-generation10K<n<100K0 likes71 downloads7d agoHugging Face11pinzhenchen /alpaca-cleaned-cs Data Description This HF data repository contains the Czech Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Czech. Usage This data is intended to be used for Czech instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-cs.texttext-generation10K<n<100K0 likes69 downloads3y agoHugging Face12ilhamfadheel /alpaca-cleaned-indonesian 🦙🛁 Cleaned Alpaca Dataset (INDONESIAN) Welcome to the Cleaned Alpaca Dataset repository! This repository hosts a cleaned and curated version of a dataset used to train the Alpaca LLM (Large Language Model). The original dataset had several issues that are addressed in this cleaned version. On April 8, 2023 the remaining uncurated instructions (~50,000) were replaced with data from the GPT-4-LLM dataset. Curation of the incoming GPT-4 data is ongoing. A 7b Lora model (trained on… See the full description on the dataset page: https://huggingface.co/datasets/ilhamfadheel/alpaca-cleaned-indonesian.texttext-generation10K<n<100K1 likes69 downloads2y agoHugging Face13pinzhenchen /alpaca-cleaned-fr Data Description This HF data repository contains the French Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into French. Usage This data is intended to be used for French instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-fr.texttext-generation10K<n<100K1 likes59 downloads3y agoHugging Face14pinzhenchen /alpaca-cleaned-ru Data Description This HF data repository contains the Russian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Russian. Usage This data is intended to be used for Russian instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-ru.texttext-generation10K<n<100K1 likes54 downloads3y agoHugging Face15emre /stanford-alpaca-cleaned-turkish-translated09/04/2023 Update: New instructions added from: https://github.com/Instruction-Tuning-with-GPT-4/GPT-4-LLM Original Version: https://github.com/tatsu-lab/stanford_alpaca#data-release AI BASED TRANSLATION RESULTS OF STANFORD ALPACA EN TO TR For academic only, please cite before you use it. Taşar, D. E. T. (2023). stanford-alpaca-cleaned-turkish-translated [Dataset]. In Stanford Alpaca TR (1.0.1.a). https://huggingface.co/datasets/emre/stanford-alpaca-cleaned-turkish-translated… See the full description on the dataset page: https://huggingface.co/datasets/emre/stanford-alpaca-cleaned-turkish-translated.text-generation10K<n<100K27 likes53 downloads3y agoHugging Face16Dans-DiscountModels /Alpaca_Evol_Instruct_CleanedAlpaca Evol Instruct cleaned of refusals, scrubbed of overly repetitive responses, aggresively deduplicated, and all URLs removed from the output. The final dataset has aproximately 54k instructions. Base dataset https://huggingface.co/datasets/victor123/evol_instruct_70k texttext-generation100K<n<1M6 likes52 downloads3y agoHugging Face17ASIDS /alpaca-cleaned-ru alpaca-cleaned-ru converter for autotrain from d0rj/alpaca-cleaned-ru Translated version of yahma/alpaca-cleaned into Russian. text-generation10K<n<100K0 likes50 downloads3y agoHugging Face18pinzhenchen /alpaca-cleaned-bg Data Description This HF data repository contains the Bulgarian Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Bulgarian. Usage This data is intended to be used for Bulgarian instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below:… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-bg.texttext-generation10K<n<100K2 likes36 downloads3y agoHugging Face19pinzhenchen /alpaca-cleaned-de Data Description This HF data repository contains the German Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into German. Usage This data is intended to be used for German instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-de.texttext-generation10K<n<100K1 likes34 downloads3y agoHugging Face20pinzhenchen /alpaca-cleaned-fi Data Description This HF data repository contains the Finnish Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Finnish. Usage This data is intended to be used for Finnish instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-fi.texttext-generation10K<n<100K0 likes34 downloads3y agoHugging Face21datatab /alpaca-cleaned-serbian-full Serbian Alpaca Cleaned Dataset Original Repository: https://github.com/gururise/AlpacaDataCleaned Original HF Repository: https://huggingface.co/datasets/yahma/alpaca-cleaned Dataset Description This is a serbian cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the… See the full description on the dataset page: https://huggingface.co/datasets/datatab/alpaca-cleaned-serbian-full.texttext-generation10K<n<100K1 likes31 downloads3y agoHugging Face22pinzhenchen /alpaca-cleaned-zh Data Description This HF data repository contains the Chinese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Chinese. Usage This data is intended to be used for Chinese instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-zh.texttext-generation10K<n<100K1 likes30 downloads3y agoHugging Face23hbpkillerX /alpaca-cleaned-hinglish Alpaca Cleaned (Hinglish Version) Dataset Description This is a high-quality Hinglish (Hindi written in Latin script) translation of the yahma/alpaca-cleaned dataset. It is designed for instruction fine-tuning large language models to make them conversational in Indian contexts. Dataset Summary Original Source: yahma/alpaca-cleaned (51,760 rows) Language: Hinglish (Code mixed Hindi-English) Translation Method: High-precision batch translation using… See the full description on the dataset page: https://huggingface.co/datasets/hbpkillerX/alpaca-cleaned-hinglish.texttext-generation10K<n<100K1 likes28 downloads8mo agoHugging Face24abrarfahim /alpaca-cleaned-bn Dataset Card for Alpaca-Cleaned-bn This is a cleaned bengali translated version of the original Alpaca Dataset released by Stanford. Uses import datasets dataset = datasets.load_dataset("abrarfahim/alpaca-cleaned-bn") print(dataset[0]) Dataset Structure {'system_prompt': 'You are a virtual assistant, deliver a comprehensive response.', 'qas_id': 'YY9S5K', 'question_text': '"সন্দেহ" শব্দের সঠিক প্রতিশব্দ নির্বাচন করুন।', 'orig_answer_texts': '"সন্দেহ"… See the full description on the dataset page: https://huggingface.co/datasets/abrarfahim/alpaca-cleaned-bn.textquestion-answering10K<n<100K0 likes24 downloads3y agoHugging Face25melvindave /alpaca-cleaned Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer. "instruction":"Summarize the… See the full description on the dataset page: https://huggingface.co/datasets/melvindave/alpaca-cleaned.texttext-generation10K<n<100K0 likes21 downloads10mo agoHugging Face26leo009 /alpaca-cleaned-zh-cn Data Description This HF data repository contains the Chinese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Chinese. Usage This data is intended to be used for Chinese instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown below: {… See the full description on the dataset page: https://huggingface.co/datasets/leo009/alpaca-cleaned-zh-cn.texttext-generation10K<n<100K6 likes19 downloads2y agoHugging Face27TPM-28 /alpaca-cleaned-frtexttext-generation10K<n<100K0 likes16 downloads2y agoHugging Face28pacozaa /alpaca-cleaned-chatml ChatML Reformat of yahma/alpaca-cleaned I'd like to try instruction-tuning dataset with chat-tuning format. Usage dataset = load_dataset("pacozaa/alpaca-cleaned-chatml", split = "train") print(dataset[0]["text"]) texttext-generation10K<n<100K0 likes16 downloads2y agoHugging Face29SOULAMA /QA-text-generation-alpaca-data-cleaned Dataset Card for mBART-QA-Processed This dataset consists of tokenized pairs of instructions and contexts designed for fine-tuning Sequence-to-Sequence models (like mBART or T5) on Question Answering tasks. Dataset Details Dataset Description The dataset is a processed version of a Question Answering corpus (SQuAD-like). It has been formatted to follow a specific prompt structure: instruction: {question} input: {context}. The targets (labels) are the direct… See the full description on the dataset page: https://huggingface.co/datasets/SOULAMA/QA-text-generation-alpaca-data-cleaned.text-generation10K<n<100K0 likes16 downloads8mo agoHugging Face30bizb0630 /alpaca-cleaned_uz Dataset Summary This dataset is a translation of the alpaca-cleaned dataset into Uzbek (Latin), using the GPT-4o mini API. texttext-generation10K<n<100K1 likes15 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.