dignity045/Collective-Corpus
🧠Collective Corpus — Universal Pretraining + Finetuning Dataset (500B+ Tokens) Collective-Corpus is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains — all in one place. 📚 Dataset Scope This dataset aims to cover the full LLM lifecycle, from raw pretraining to domain-specialized finetuning. 1. Pretraining Corpus Large-scale, diverse… See the full description on the dataset page: https://huggingface.co/datasets/dignity045/Collective-Corpus.
This repository belongs to dignity045 on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
