datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ham10kbengal-dharma-corpus
Bengal Dharma Corpus
Evolution of Bengali Devotional Language: a multi-tradition corpus spanning
Old Bengali, Sanskrit, and modern Bengali across Buddhist, Shakta, and Vaishnava
traditions, 8th to 19th century.
Assembled and curated by Joy Bose (joyboseroy), June 2026.
Code and analysis: https://github.com/joyboseroy/bengal-dharma-corpus
Related dataset: joyboseroy/darshana-graph (arXiv:2606.18222)
What this corpus is
This is a curated collection of 75 texts from… See the full description on the dataset page: https://huggingface.co/datasets/joyboseroy/bengal-dharma-corpus.CT-RATE-Dataset-cleaned
