CoolFace
Datasetpublic

dignity045/Collective-Corpus

🧠 Collective Corpus — Universal Pretraining + Finetuning Dataset (500B+ Tokens) Collective-Corpus is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains — all in one place. 📚 Dataset Scope This dataset aims to cover the full LLM lifecycle, from raw pretraining to domain-specialized finetuning. 1. Pretraining Corpus Large-scale, diverse… See the full description on the dataset page: https://huggingface.co/datasets/dignity045/Collective-Corpus.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
2likes959downloads
settings

This repository belongs to dignity045 on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

nameCollective-Corpus
visibilitypublic
licenceapache-2.0
gatedno
ownerdignity045
Account settings
dignity045/Collective-Corpus · CoolFace