datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swiss-legal-dapt-datasetDAPTSL_fMoWBeejX-Agriculture-DAPT-CorpusBeejX LLM: DAPT Corpus
Curating ""Grade A+ Clean Text for Indian Agriculture AI.
Dataset Overview
The BeejX DAPT Corpus (dapt_train_final.txt) is a highly curated Domain-Adapted Pre-Training (DAPT) dataset designed to teach Large Language Models the deep, technical nuances of Indian Agriculture.
Our goal was to transform raw, noisy agricultural documents (textbooks, market reports, scientific PDFs) into "Grade A+" clean text suitable for continuously training base models… See the full description on the dataset page: https://huggingface.co/datasets/bf369/BeejX-Agriculture-DAPT-Corpus.dapt-2020-pcapsvazhi-dapt-sources-v2_0vazhi-dapt-tamil-v1_1vazhi-dapt-tamil-v1_0DAPT-Counselling-Conversations-2vazhi-dapt-tamil-v2_1DAPT-Counselling-Conversationsvazhi-dapt-tamil-v2_0MOMENT-DAPT-Scalingmlm-dapt-aeroastro
