IFM/K2Datasets
K2 Dataset Card The following data mix was used to train K2 and achieve results in line with Llama 2 70B. Dataset Details K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below. Dataset Description: Stage 1 Dataset Starting Tokens Multiplier Total Tokens % of Total dm-math 4.33B 3x 13B 1% pubmed-abstracts (from the Pile) 4.77B 3x 14.3B 1.1% uspto (from the Pile) 4.77B 3x… See the full description on the dataset page: https://huggingface.co/datasets/IFM/K2Datasets.
K2 Dataset Card
<!-- Provide a quick summary of the dataset. -->
The following data mix was used to train K2 and achieve results in line with Llama 2 70B.
Dataset Details
K2 was trained on 1.4T tokens across two stages. The data sources and data mix for each stage are listed below.
Dataset Description: Stage 1
<!-- Provide a longer summary of what this dataset is. -->
Dataset Description: Stage 2
*rounding
Data Collection and Processing
<!-- This section describes the data collection and processing process such as data selection criteria, filtering and normalization methods, tools and libraries used, etc. -->
A step-by-step tutorial for reproducing the K2's data preperation can be found in the LLM360 Pretraining Suite here
Bias, Risks, and Limitations
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
Citation
BibTeX:
@misc{
title={LLM360 K2-65B: Scaling Up Open and Transparent Language Models},
author={The LLM360 Team},
year={2024},
}