firdavsus/LLM_multimodal
Dataset Card for LLM_multimodal LLM_multimodal is the official training corpus designed for the LLM D6 model series. It contains a massive, high-quality collection of 300 Billion tokens, carefully curated to balance linguistic diversity, mathematical reasoning, and programming capabilities. This repository hosts both the raw/processed pre-training data and the instruction-following datasets used for supervised fine-tuning (SFT). tokenizer is located at LLM_D6 in my… See the full description on the dataset page: https://huggingface.co/datasets/firdavsus/LLM_multimodal.
011
