CoolFace
Datasetpublic

firdavsus/LLM_multimodal

Dataset Card for LLM_multimodal LLM_multimodal is the official training corpus designed for the LLM D6 model series. It contains a massive, high-quality collection of 300 Billion tokens, carefully curated to balance linguistic diversity, mathematical reasoning, and programming capabilities. This repository hosts both the raw/processed pre-training data and the instruction-following datasets used for supervised fine-tuning (SFT). tokenizer is located at LLM_D6 in my… See the full description on the dataset page: https://huggingface.co/datasets/firdavsus/LLM_multimodal.

sourceHugging Faceapache-2.0updated 2mo agoView on Hugging Face
0likes11downloads

firdavsus/LLM_multimodal · main · files are served by the source, never re-hosted here