AlpacaCOT
Alpaca-CoT
Instruction-Finetuning Dataset Collection (Alpaca-CoT)
This repository will continuously collect various instruction tuning datasets. And we standardize different datasets into the same format, which can be directly loaded by the code of Alpaca model.
We also have conducted empirical study on various instruction-tuning datasets based on the Alpaca model, as shown in https://github.com/PhoebusSi/alpaca-CoT.
If you think this dataset collection is helpful to you, please like… See the full description on the dataset page: https://huggingface.co/datasets/QingyiSi/Alpaca-CoT.alpaca-cot-collectionalpaca-cot-zh-refined-by-data-juicer
Alpaca-CoT -- ZH (refined by Data-Juicer)
A refined Chinese version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to fine-tune a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 18.7GB).
Dataset Information
Number of samples: 9,873,214 (Keep ~46.58% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-zh-refined-by-data-juicer.alpaca-cot-en-refined-by-data-juicer
Alpaca-CoT -- EN (refined by Data-Juicer)
A refined English version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to fine-tune a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 226GB).
Dataset Information
Number of samples: 72,855,345 (Keep ~54.48% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-en-refined-by-data-juicer.alpaca-cot-collection-jsonifizealpaca-cot-collection_stringified-jsonifize
