datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brahma-sutras
Brahma Sūtras (ब्रह्मसूत्राणि) — Complete Śrī Madhvācārya Dvaita Sarvamūla Canon
The complete computational and hermeneutic dataset of all 555 Brahma Sūtras of Bādarāyaṇa Vyāsa with the Dvaita Vedānta (Tattvavāda) commentaries of Śrī Madhvācārya (Ānandatīrtha).
🏛️ Corpus Structure & Sarvamūla Works
This dataset includes:
555 Canonical Sūtras structured across 4 Adhyāyas (Samanvaya, Avirodha, Sādhana, Phala) and 16 Pādas.
Pāṇinian Padaccheda: Exact… See the full description on the dataset page: https://huggingface.co/datasets/gnumanth/brahma-sutras.gnu-prolog-adaptation-corpus
GNU Prolog adaptation corpus — AutoScientist Challenge (Math & Code)
~1,200 execution-verified GNU Prolog task/completion pairs plus a frozen
175-task holdout (holdout.jsonl, hash-pinned before any training run).
Every completion was verified by executing it against the task's checks;
no completion entered the corpus on an LLM's word alone. Generator, seeds
and manifest included. Used to train
AryaGarg23/llama-3.2-3b-gnu-prolog-lora (13.7% -> 76.0%
executable pass@1 at 3B).
gnu-coreutils-small
