IFM/K2Datasets
K2 Dataset Card The following data mix was used to train K2 and achieve results in line with Llama 2 70B. Dataset Details K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below. Dataset Description: Stage 1 Dataset Starting Tokens Multiplier Total Tokens % of Total dm-math 4.33B 3x 13B 1% pubmed-abstracts (from the Pile) 4.77B 3x 14.3B 1.1% uspto (from the Pile) 4.77B 3x… See the full description on the dataset page: https://huggingface.co/datasets/IFM/K2Datasets.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face