CoolFace
20 results

sutra

EmmaLeonhart /sutra-w2c-corpus sutra-w2c-corpus A weights↔code training corpus for Sutra weight→code decompilation: generated Sutra programs whose behavior is carried by matrices, paired with those matrices (the "weights") and the program's substrate input→output behavior. The long-term goal is a model that maps weights → code (recovering the program from its learned parameters). This dataset is generated by experiments/weight_to_code_corpus.py in the Sutra repo (where it is pinned as the corpus/ submodule)… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/sutra-w2c-corpus.textn<1K0 likes1.6k downloads3mo agoHugging Facecodelion /sutra-1B Sutra 1B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational patterns Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-1B.tabulartext-generation100K<n<1M2 likes1.2k downloads7mo agoHugging Facecodelion /sutra-10B Sutra 10B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 10,193,029 educational entries totaling over 10 billion tokens. This is the largest dataset in the Sutra series, designed to demonstrate that dense, curated datasets can provide best-in-class pretraining performance for small language models. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-10B.tabulartext-generation1M<n<10M12 likes570 downloads7mo agoHugging Facegnumanth /brahma-sutras Brahma Sūtras (ब्रह्मसूत्राणि) — Complete Śrī Madhvācārya Dvaita Sarvamūla Canon The complete computational and hermeneutic dataset of all 555 Brahma Sūtras of Bādarāyaṇa Vyāsa with the Dvaita Vedānta (Tattvavāda) commentaries of Śrī Madhvācārya (Ānandatīrtha). 🏛️ Corpus Structure & Sarvamūla Works This dataset includes: 555 Canonical Sūtras structured across 4 Adhyāyas (Samanvaya, Avirodha, Sādhana, Phala) and 16 Pādas. Pāṇinian Padaccheda: Exact… See the full description on the dataset page: https://huggingface.co/datasets/gnumanth/brahma-sutras.tabulartext-generationn<1K1 likes95 downloads23d agoHugging Facecodelion /sutra-magpie-sft Sutra Magpie SFT Dataset A high-quality dataset of 20,682 instruction-response pairs for supervised fine-tuning (SFT) of language models. Generated using seed prompts from the Sutra framework with Magpie-style response generation. Dataset Description This dataset provides diverse, high-quality instruction-response pairs suitable for training instruction-following language models. Generation Method Seed Prompts: Started with 30K diverse seed prompts from… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-magpie-sft.texttext-generation10K<n<100K2 likes82 downloads7mo agoHugging Facecodelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes67 downloads7mo agoHugging Face