CoolFace
Datasetpublic

m-a-p/Matrix

Matrix An open-source pretraining dataset containing 4690 billion tokens, this bilingual dataset with both English and Chinese texts is used for training neo models. Dataset Composition The dataset consists of several components, each originating from different sources and serving various purposes in language modeling and processing. Below is a brief overview of each component: Common Crawl Extracts from the Common Crawl project, featuring a rich diversity… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/Matrix.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
176likes18kdownloads
code_code.0042.jsonl4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:c4f1eb4f9c28aa9ee2b5b0ec6c6d35b255584e9eb4b33e965f5fa9c6ba6cb7ce3size 429468503164