dharma
Datasets
All datasets matching “dharma”ham10kDharmaOCR-Benchmark
DharmaOCR-Benchmark
Overview
DharmaOCR-Benchmark is a 496-instance evaluation suite for OCR models focused on Brazilian Portuguese documents. It covers printed text, handwritten text, and legal/administrative documents — domains underrepresented in existing benchmarks like OCRBench and olmOCR-Bench.
This benchmark evaluates not only transcription quality, but also text degeneration rate and unit inference cost as first-class metrics.
Released alongside the… See the full description on the dataset page: https://huggingface.co/datasets/Dharma-AI/DharmaOCR-Benchmark.defining-dharma
In Search of Dharma — research corpus
Dharma here means how human societies build, transmit and enforce ethics —
not metaphysics. A traceably-sourced research corpus: every claim cited, drawn
from anthropology, evolutionary biology, history and comparative ethics.
Written by Gary Dean (Biksu Okusi).
The corpus has three configs serving different purposes:
notes (52 records, 1562 cited sources) — Stage-1
research notes. Each note answers one registry question and carries… See the full description on the dataset page: https://huggingface.co/datasets/garydean/defining-dharma.dharma-1
"Dharma-1"
A new carefully curated benchmark set, designed for a new era where the true end user uses LLM's for zero-shot and one-shot tasks, for a vast majority of the time.
Stop training your models on mindless targets (eval_loss, train_loss), start training your LLM on lightweight Dharma as an eval target.
A mix of all the top benchmarks.
Formed to have an equal distribution of some of the most trusted benchmarks used by those developing SOTA LLMs, comprised of only 3,000… See the full description on the dataset page: https://huggingface.co/datasets/pharaouk/dharma-1.dharma-2
"dharma_g1i5 Dataset"
A dharma evaluation dataset with the following configuration:
||| Subject: MMLU, Size: 38 |||
||| Subject: ARC-Challenge, Size: 38 |||
||| Subject: ARC-Easy, Size: 38 |||
||| Subject: BoolQ, Size: 35 |||
||| Subject: winogrande, Size: 38 |||
||| Subject: openbookqa, Size: 38 |||
||| Subject: truthful_qa, Size: 38 |||
||| Subject: agieval, Size: 37 |||
Made with https://github.com/pharaouk/dharma 🚀
bengal-dharma-corpus
Bengal Dharma Corpus
Evolution of Bengali Devotional Language: a multi-tradition corpus spanning
Old Bengali, Sanskrit, and modern Bengali across Buddhist, Shakta, and Vaishnava
traditions, 8th to 19th century.
Assembled and curated by Joy Bose (joyboseroy), June 2026.
Code and analysis: https://github.com/joyboseroy/bengal-dharma-corpus
Related dataset: joyboseroy/darshana-graph (arXiv:2606.18222)
What this corpus is
This is a curated collection of 75 texts from… See the full description on the dataset page: https://huggingface.co/datasets/joyboseroy/bengal-dharma-corpus.
