slim
Datasets
All datasets matching “slim”SlimPajama-6BSampled version of cerebras/SlimPajama-627B.
Since the original data was shuffled before chunking, I only downloaded train/chunk1 (of 10 total) and further sampled 10%. This should result in roughly 6B tokens, hence SlimPajama-6B.
The dataset is 24GBs in storage size when decompressed (original dataset is over 2TBs) and has 5489000 rows.
The validation set and test set were sampled as well.
Data source proportions for SlimPajama-627B and SlimPajama-6B
For sanity purpose, I… See the full description on the dataset page: https://huggingface.co/datasets/DKYoon/SlimPajama-6B.SlimPajama-627B_ReuploadAs datasets puts limits on the number of calls to huggingface, downloading SlimPajama-627B is problematic as it's composed of a ton of small files.
I have reuploaded it here as larger chunks to easily download the dataset without having to do anything hacky.
The original dataset can be found here https://huggingface.co/datasets/cerebras/SlimPajama-627B
SlimPajama-Meta-rater
Annotated SlimPajama Dataset
Dataset Description
This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions.
Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.slimetaoshite300nenshiranaiuchinilevelmaxninattemashitasononi
Bangumi Image Base of Slime Taoshite 300-nen, Shiranai Uchi Ni Level Max Ni Nattemashita: Sono Ni
This is the image base of bangumi Slime Taoshite 300-nen, Shiranai Uchi ni Level Max ni Nattemashita: Sono Ni, we detected 66 characters, 7416 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/slimetaoshite300nenshiranaiuchinilevelmaxninattemashitasononi.SlimOrca
Overview
This is a new curated subset of our OpenOrca data. This release provides an efficient means of reaching performance on-par with using larger slices of our data, while only including ~500k GPT-4 completions.
The key change in this dataset is that we've done an additional pass, using GPT-4 to remove answers which appear wrong based on the human annotations from the FLAN dataset.
This reduces the dataset size to only ~500k entries, allowing training to a similar quality level… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca.slimpajama_Llama2_Tokenizer
slimpajama_Llama2_Tokenizer
The original slimpajama_Llama2_Tokenizer.tar.gz archive (≈794 GB on Ubuntu) was split into smaller 40 GB chunks for easier upload to Hugging Face.
sudo apt install git-lfs
pip install -U huggingface_hub # `hf version`==1.1.4
tar cvf - slimpajama_Llama2_Tokenizer/ | pigz -p 16 > slimpajama_Llama2_Tokenizer.tar.gz
split -b 40G -d -a 3 slimpajama_Llama2_Tokenizer.tar.gz slimpajama_Llama2_Tokenizer/slimpajama_Llama2_Tokenizer_part_
# Upload files… See the full description on the dataset page: https://huggingface.co/datasets/jsun/slimpajama_Llama2_Tokenizer.
