bf369/BeejX-Agriculture-DAPT-Corpus
BeejX LLM: DAPT Corpus Curating ""Grade A+ Clean Text for Indian Agriculture AI. Dataset Overview The BeejX DAPT Corpus (dapt_train_final.txt) is a highly curated Domain-Adapted Pre-Training (DAPT) dataset designed to teach Large Language Models the deep, technical nuances of Indian Agriculture. Our goal was to transform raw, noisy agricultural documents (textbooks, market reports, scientific PDFs) into "Grade A+" clean text suitable for continuously training base models… See the full description on the dataset page: https://huggingface.co/datasets/bf369/BeejX-Agriculture-DAPT-Corpus.
<h1 align="center" style="color: #4CAF50; font-weight: bold;">BeejX LLM: DAPT Corpus </h1>
Curating ""Grade A+ Clean Text for Indian Agriculture AI.
Dataset Overview
The BeejX DAPT Corpus (dapt_train_final.txt) is a highly curated Domain-Adapted Pre-Training (DAPT) dataset designed to teach Large Language Models the deep, technical nuances of Indian Agriculture.
Our goal was to transform raw, noisy agricultural documents (textbooks, market reports, scientific PDFs) into "Grade A+" clean text suitable for continuously training base models like Gemma 2B, converting them into specialized agricultural assistants.
Dataset Structure
The dataset consists of highly structured bilingual text (English and Hindi). Data is separated by standard <doc> tags to help models distinguish between different contexts during continuous pre-training.
The BeejX Cleaning Pipeline
Raw agricultural data is extremely noisy. To achieve "Grade A+" status, this dataset went through a strict 3-Stage Refining Process:
1. Harvest (Extraction)
- Challenges Handled: Multi-column layouts, hidden text layers, and complex Hindi numerals extracted from diverse formats (PDFs, Images, Scans).
2. Refine (Cleaning)
- Actions:
- Removed "digital noise", headers/footers ("Page 1", "RBSE Class 12").
- Collapsed whitespace and fixed line breaks.
- Stripped irrelevant quizzes, tables of content, and diagrams.
3. Polish (Semantic Fixing)
- "4 vs 1" OCR Curse: Fixed the systematic error where OCR read the Hindi '1' as '4' (e.g., correcting
42% moisture->12%). - Ghost Brackets: Removed lingering
[and]artifact fragments. - Contextual Typos: Fixed domain-specific terms like
किसमें(in what) ->किस्में(varieties).
Project Artifacts & Quality Grading
We graded the subsets of this corpus on a strict scale before approving them for training:
Key Agricultural Sources & Citations
This dataset was painstakingly compiled, translated, and extracted from the following authoritative resources. If you use this dataset, please adhere to the original source licenses:
Author & Maintenance
- Lead Data Work: ME It's Me (BhashkarFulara369)
