datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alpaca-cot-collectionalpaca-cot-zh-refined-by-data-juicer
Alpaca-CoT -- ZH (refined by Data-Juicer)
A refined Chinese version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to fine-tune a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 18.7GB).
Dataset Information
Number of samples: 9,873,214 (Keep ~46.58% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-zh-refined-by-data-juicer.alpaca-cot-en-refined-by-data-juicer
Alpaca-CoT -- EN (refined by Data-Juicer)
A refined English version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality.
This dataset is usually used to fine-tune a Large Language Model.
Notice: Here is a small subset for previewing. The whole dataset is available here (About 226GB).
Dataset Information
Number of samples: 72,855,345 (Keep ~54.48% from the original dataset)
Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-en-refined-by-data-juicer.alpaca-cot-collection-jsonifizealpaca-cot-collection_stringified-jsonifize
