CoolFace
Datasetpublic

dignity045/Collective-Corpus

🧠 Collective Corpus β€” Universal Pretraining + Finetuning Dataset (500B+ Tokens) Collective-Corpus is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains β€” all in one place. πŸ“š Dataset Scope This dataset aims to cover the full LLM lifecycle, from raw pretraining to domain-specialized finetuning. 1. Pretraining Corpus Large-scale, diverse… See the full description on the dataset page: https://huggingface.co/datasets/dignity045/Collective-Corpus.

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
2likes959downloads
Dataset Card

🧠 Collective Corpus β€” Universal Pretraining + Finetuning Dataset (500B+ Tokens)

![Hugging Face](https://huggingface.co/datasets/dignity045/Collective-Corpus) ![License: Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) ![Status](#-current-status)

`Collective-Corpus` is a massive-scale, multi-domain dataset designed to train Transformer-based language models from scratch and finetune them across a wide variety of domains β€” all in one place.

πŸ“š Dataset Scope

This dataset aims to cover the full LLM lifecycle, from raw pretraining to domain-specialized finetuning.

1. Pretraining Corpus

  • β€”Large-scale, diverse multilingual text sources
  • β€”Cleaned, deduplicated, and filtered for quality
  • β€”Inspired by datasets like C4 and FineWeb

2. Domain-Specific Finetuning

  • β€”Instruction Following & Dialogue β€” Chatbots, multi-turn conversations
  • β€”Code β€” Python, JavaScript, Java, C++, and more
  • β€”Math & Logical Reasoning
  • β€”Specialized Fields β€” Research papers, technical documentation

πŸ“Š Scale

  • β€”Total Tokens: 500B+
  • β€”Estimated Text Samples: 700M+
  • β€”Target Model Size: Suitable for training large models from scratch
  • β€”Covers general-purpose and domain-specific training needs

🎯 Goals

  1. 1.Build a unified corpus for full-stack LLM development.
  2. 2.Enable open and reproducible large-scale language model research.
  3. 3.Support finetuning for high-impact domains like code, math, and dialogue.

🚧 Current Status

  • β€”Model Pretraining: Currently training a Transformer model from scratch on the full 500B+ token dataset.
  • β€”Public Release: Planned after model training completes.

🀝 Collaboration

We are actively seeking open-source collaborators to:

  • β€”Contribute to dataset cleaning, filtering, and deduplication
  • β€”Assist in large-scale model training and evaluation
  • β€”Provide expertise for specialized domain corpora

We also offer free guidance on:

  • β€”Dataset curation best practices
  • β€”Efficient large-scale LLM training pipelines
  • β€”Transformer architecture optimization

πŸ’Ό Open for Collaboration

I’m actively looking to connect with researchers, engineers, and organizations passionate about dataset engineering, large-scale model training, and applied NLP. Whether it’s open-source projects, research collaborations, or large-scale AI initiatives β€” let’s build something impactful together.

πŸ”— GitHub: Dhiraj309 πŸ”— LinkedIn: Dhiraj Patil


πŸ“… Release Timeline

StageStatus
Data Curation🚧 In Progress
Model Pretraining🚧 In Progress
Dataset Public Release⏳ Post-training

πŸ“œ License

Released under the Apache License 2.0 β€” you are free to use, modify, and distribute this dataset in compliance with the full license text.


🌍 Let’s build the next generation of open-source LLMs β€” together.