CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01yahma /alpaca-cleaned Dataset Card for Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an answer. "instruction":"Summarize… See the full description on the dataset page: https://huggingface.co/datasets/yahma/alpaca-cleaned.texttext-generation10K<n<100K888 likes24k downloads3y agoHugging Face02unsloth /alpaca-cleaned Dataset Card for Alpaca-Cleaned Forked from https://huggingface.co/datasets/yahma/alpaca-cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is a cleaned version of the original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused… See the full description on the dataset page: https://huggingface.co/datasets/unsloth/alpaca-cleaned.texttext-generation10K<n<100K25 likes3.4k downloads9mo agoHugging Face03d0rj /alpaca-cleaned-ru alpaca-cleaned-ru Translated version of yahma/alpaca-cleaned into Russian. texttext-generation10K<n<100K22 likes332 downloads3y agoHugging Face04saillab /alpaca-korean-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-korean-cleaned.text10K<n<100K0 likes238 downloads2y agoHugging Face05saillab /alpaca-japanese-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-japanese-cleaned.text10K<n<100K0 likes223 downloads2y agoHugging Face06saillab /alpaca-english-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-english-cleaned.text10K<n<100K0 likes222 downloads2y agoHugging Face07saillab /alpaca-portuguese-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-portuguese-cleaned.text10K<n<100K0 likes218 downloads2y agoHugging Face08saillab /alpaca-russian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-russian-cleaned.text10K<n<100K2 likes211 downloads2y agoHugging Face09yuekai /belle_1M_and_alpaca_cleaned dataset Alpaca and Belle 1M for InstructGLM github project training 2 likes192 downloads3y agoHugging Face10saillab /alpaca-vietnamese-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-vietnamese-cleaned.text10K<n<100K0 likes192 downloads2y agoHugging Face11saillab /alpaca-german-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-german-cleaned.text10K<n<100K0 likes179 downloads2y agoHugging Face12saillab /alpaca-spanish-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-spanish-cleaned.text10K<n<100K0 likes179 downloads2y agoHugging Face13saillab /alpaca-chinesesimplified-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-chinesesimplified-cleaned.text10K<n<100K1 likes164 downloads2y agoHugging Face14saillab /alpaca-thai-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-thai-cleaned.text10K<n<100K1 likes162 downloads2y agoHugging Face15saillab /alpaca-hindi-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-hindi-cleaned.text10K<n<100K0 likes158 downloads2y agoHugging Face16pinzhenchen /alpaca-cleaned-pt Data Description This HF data repository contains the Portuguese Alpaca dataset used in our study of monolingual versus multilingual instruction tuning. GitHub Paper Creation Machine-translated from yahma/alpaca-cleaned into Portuguese. Usage This data is intended to be used for Portuguese instruction tuning. The dataset has roughly 52K instances in the JSON format. Each instance has an instruction, an output, and an optional input. An example is shown… See the full description on the dataset page: https://huggingface.co/datasets/pinzhenchen/alpaca-cleaned-pt.texttext-generation10K<n<100K5 likes153 downloads3y agoHugging Face17DanielSc4 /alpaca-cleaned-italian Dataset Card for Alpaca-Cleaned-Italian About the translation and the original data The translation was done with X-ALMA, a 13-billion-parameter model that surpasses state-of-the-art open-source multilingual LLMs (as of Q1 2025, paper here). The original alpaca-cleaned dataset is also kept here so that there is parallel data for Italian and English. Additional notes on the translation Despite the good quality of the translation, errors, though rare, are… See the full description on the dataset page: https://huggingface.co/datasets/DanielSc4/alpaca-cleaned-italian.texttext-generation100K<n<1M7 likes148 downloads2y agoHugging Face18sarpba /alpaca-cleaned-gemini-hunBazsalanszky/alpaca-cleaned-gemini-hun alpaca fordításának a szűrése llama3.1 segítségével. A szűrő prompt: "Egy profi adatelemző vagy, aki a user - assistant interakciót elemzi. Az aszisztant válasza mennyire felelt meg a felhaszálói kérésnek vagy kérdésnek 1-10 közt. Elemezd a választ, légy alapos. Az 1-es érték azt jelenti, hogy teljesen helytelen a válasz a 10-es érték azt jelenti, hogy a válasz teljesen megfelel a user kérésének vagy kérdésének. Csak egy az elemzésednek megfelelő számot… See the full description on the dataset page: https://huggingface.co/datasets/sarpba/alpaca-cleaned-gemini-hun.text10K<n<100K2 likes143 downloads2y agoHugging Face19Thaweewat /alpaca-cleaned-52k-th Summary This is a Thai 🇹🇭-instructed dataset translated from cleaned version of the original Alpaca Dataset released by Stanford using Google Cloud Translation, contain 52,000 instructions and demonstrations generated by OpenAI's text-davinci-003 engine. This instruction data can be used to conduct instruction-tuning for language models and make the language model follow instruction better. The following issues have been identified in the original release and fixed in this… See the full description on the dataset page: https://huggingface.co/datasets/Thaweewat/alpaca-cleaned-52k-th.textquestion-answering10K<n<100K17 likes138 downloads3y agoHugging Face20iamshnoo /alpaca-cleaned-bengaliTranslated from yahma/alpaca-cleaned using NLLB-1.3B Dataset Card for "alpaca-cleaned-bengali" More Information needed text10K<n<100K13 likes130 downloads3y agoHugging Face21shi3z /alpaca_cleaned_ja_json Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/shi3z/alpaca_cleaned_ja_json.texttext-generation100K<n<1M13 likes111 downloads3y agoHugging Face22iamshnoo /alpaca-cleaned-persianTranslated from yahma/alpaca-cleaned using NLLB-1.3B Dataset Card for "alpaca-cleaned-persian" More Information needed text10K<n<100K8 likes109 downloads3y agoHugging Face23cahya /alpaca-id-cleaned Dataset Card for Indonesian Alpaca-Cleaned Repository: https://github.com/gururise/AlpacaDataCleaned Dataset Description This is the Indonesian translated version of the cleaned original Alpaca Dataset released by Stanford. The following issues have been identified in the original release and fixed in this dataset: Hallucinations: Many instructions in the original dataset had instructions referencing data on the internet, which just caused GPT3 to hallucinate an… See the full description on the dataset page: https://huggingface.co/datasets/cahya/alpaca-id-cleaned.texttext-generation10K<n<100K9 likes108 downloads3y agoHugging Face24jayasuryajsk /alpaca_cleaned_activations_layer_16text1K<n<10K1 likes107 downloads2y agoHugging Face25saillab /alpaca-nepali-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-nepali-cleaned.text10K<n<100K2 likes104 downloads2y agoHugging Face26iamshnoo /alpaca-cleaned-hindiTranslated from yahma/alpaca-cleaned using NLLB-1.3B Dataset Card for "alpaca-cleaned-hindi" More Information needed text10K<n<100K4 likes99 downloads3y agoHugging Face27mmosiolek /pl_alpaca_data_cleaned Polpaca: The Polish Alpaca Please find the model here: https://huggingface.co/mmosiolek/polpaca-lora-7b This repository contains the polish translations of the datasets for constructing and evaluating instruction following models: Alpaca. Training The following dataset was translated: https://github.com/gururise/AlpacaDataCleaned It might be also found here: https://huggingface.co/datasets/yahma/alpaca-cleaned For the translation process, I relied on GPT-3.5-Turbo and the… See the full description on the dataset page: https://huggingface.co/datasets/mmosiolek/pl_alpaca_data_cleaned.7 likes90 downloads3y agoHugging Face28iamshnoo /alpaca-cleaned-chineseTranslated from yahma/alpaca-cleaned using NLLB-1.3B Dataset Card for "alpaca-cleaned-chinese" More Information needed text10K<n<100K4 likes88 downloads3y agoHugging Face29saillab /alpaca-javanese-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-javanese-cleaned.text10K<n<100K0 likes87 downloads2y agoHugging Face30open-llm-leaderboard-old /details_dsvv-cair__alpaca-cleaned-llama-30b-bf16 Dataset Card for Evaluation run of dsvv-cair/alpaca-cleaned-llama-30b-bf16 Dataset Summary Dataset automatically created during the evaluation run of model dsvv-cair/alpaca-cleaned-llama-30b-bf16 on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_dsvv-cair__alpaca-cleaned-llama-30b-bf16.0 likes86 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.