CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EmmaLeonhart /sutra-w2c-corpus sutra-w2c-corpus A weights↔code training corpus for Sutra weight→code decompilation: generated Sutra programs whose behavior is carried by matrices, paired with those matrices (the "weights") and the program's substrate input→output behavior. The long-term goal is a model that maps weights → code (recovering the program from its learned parameters). This dataset is generated by experiments/weight_to_code_corpus.py in the Sutra repo (where it is pinned as the corpus/ submodule)… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/sutra-w2c-corpus.textn<1K0 likes1.6k downloads3mo agoHugging Face02codelion /sutra-1B Sutra 1B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 948,709 educational entries totaling over 1 billion tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational patterns Cross-domain… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-1B.tabulartext-generation100K<n<1M2 likes1.2k downloads7mo agoHugging Face03codelion /sutra-10B Sutra 10B Pretraining Dataset A high-quality pedagogical dataset designed for LLM pretraining, containing 10,193,029 educational entries totaling over 10 billion tokens. This is the largest dataset in the Sutra series, designed to demonstrate that dense, curated datasets can provide best-in-class pretraining performance for small language models. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-10B.tabulartext-generation1M<n<10M12 likes570 downloads7mo agoHugging Face04gnumanth /brahma-sutras Brahma Sūtras (ब्रह्मसूत्राणि) — Complete Śrī Madhvācārya Dvaita Sarvamūla Canon The complete computational and hermeneutic dataset of all 555 Brahma Sūtras of Bādarāyaṇa Vyāsa with the Dvaita Vedānta (Tattvavāda) commentaries of Śrī Madhvācārya (Ānandatīrtha). 🏛️ Corpus Structure & Sarvamūla Works This dataset includes: 555 Canonical Sūtras structured across 4 Adhyāyas (Samanvaya, Avirodha, Sādhana, Phala) and 16 Pādas. Pāṇinian Padaccheda: Exact… See the full description on the dataset page: https://huggingface.co/datasets/gnumanth/brahma-sutras.tabulartext-generationn<1K1 likes95 downloads23d agoHugging Face05codelion /sutra-magpie-sft Sutra Magpie SFT Dataset A high-quality dataset of 20,682 instruction-response pairs for supervised fine-tuning (SFT) of language models. Generated using seed prompts from the Sutra framework with Magpie-style response generation. Dataset Description This dataset provides diverse, high-quality instruction-response pairs suitable for training instruction-following language models. Generation Method Seed Prompts: Started with 30K diverse seed prompts from… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-magpie-sft.texttext-generation10K<n<100K2 likes82 downloads7mo agoHugging Face06codelion /sutra-100M Sutra 100M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 70,435 educational entries totaling approximately 100 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-100M.tabulartext-generation10K<n<100K3 likes67 downloads7mo agoHugging Face07codelion /sutra-improved-100M Sutra Improved 100M A self-improved pedagogical dataset for LLM pretraining, containing 413,899 entries totaling 110,038,011 tokens (~110 million). This dataset was created by applying an iterative self-improvement process to the Sutra-10B dataset, where each sample was rewritten using Gemma-3-4B-IT and only the better version (original or rewritten) was kept, followed by comprehensive deduplication and quality filtering. Dataset Description This dataset explores… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-improved-100M.texttext-generation100K<n<1M2 likes37 downloads6mo agoHugging Face08codelion /sutra-10M Sutra 10M Pretraining Dataset A high-quality synthetic pedagogical dataset designed for LLM pretraining, containing 7,252 educational entries totaling approximately 10 million tokens. Dataset Description This dataset was generated using the Sutra framework, which creates structured educational content optimized for language model pretraining. Each entry is designed to maximize learning efficiency through: Clear pedagogical structure: Content follows proven educational… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-10M.tabulartext-generation1K<n<10K3 likes32 downloads7mo agoHugging Face09Abhisingh-18 /Sutra-1.3B-Data Sutra-1.3B — Training Data Recipe Every dataset used to build Sutra-1.3B, a 1.32B MoE model trained from scratch — with the exact config, split, text field and token share for each, plus the code that turns them into the corpus. This repo is the recipe, not the ingredients. The tokenized corpus is 93 GB of uint16 shards derived from other people's datasets, each under its own licence. Rather than redistribute that, this gives you the specification and the script — run it and you… See the full description on the dataset page: https://huggingface.co/datasets/Abhisingh-18/Sutra-1.3B-Data.tabulartext-generationn<1K0 likes29 downloads1mo agoHugging Face10malharinamdar /marathi-sutra-full-tokenizedtext1M<n<10M0 likes27 downloads1y agoHugging Face11codelion /sutra-30k-seeds Sutra 30K Seeds A curated dataset of 30,320 diverse instruction prompts designed for generating high-quality SFT (Supervised Fine-Tuning) datasets. These seeds serve as the foundation for creating instruction-response pairs for post-training language models. Dataset Description This dataset contains seed prompts across 4 primary capabilities and 18 sub-capabilities, designed to cover the core competencies needed for instruction-following models. Generation Method… See the full description on the dataset page: https://huggingface.co/datasets/codelion/sutra-30k-seeds.texttext-generation10K<n<100K2 likes25 downloads7mo agoHugging Face12shethjenil /Jain-Sutra-Arthtext1K<n<10K0 likes12 downloads2y agoHugging Face13dataspoof /Vedanga-Sutra-Literature-Dataset 📜 Vedāṅga & Sutra Literature Dataset (Cleaned CSV Format) This repository contains cleaned and structured CSV datasets of important Vedāṅga and Sutra literature, including: Gṛhya Sutras Śrauta Sutras Śulba Sutras Prātiśākhyas Pariśiṣṭas All texts are primarily sourced from: GRETIL – Göttingen Register of Electronic Texts in Indian Languages🔗 https://gretil.sub.uni-goettingen.de/gretil.html The raw Sanskrit texts were extracted from GRETIL and transformed into structured… See the full description on the dataset page: https://huggingface.co/datasets/dataspoof/Vedanga-Sutra-Literature-Dataset.0 likes12 downloads7mo agoHugging Face14gowribharat /sutras_casestext10K<n<100K0 likes11 downloads1y agoHugging Face15Ashoka74 /sutras_testimagen<1K0 likes10 downloads2y agoHugging Face16gowribharat /sutrastext10K<n<100K0 likes7 downloads1y agoHugging Face17kilanisainikhil /sutra-aerial-datasetimage1K<n<10K0 likes7 downloads4mo agoHugging Face18malharinamdar /hindi-sutra-scaled-tokenizedtext1M<n<10M0 likes4 downloads1y agoHugging Face19kurniapratiwi061 /humanoid-sutradara-datatextn<1K0 likes4 downloads8mo agoHugging Face20Sutranix /dsa-db2tabularn<1K0 likes2 downloads7mo agoHugging Face21malharinamdar /hindi-sutra-full-tokenizedtext1M<n<10M0 likes1 downloads1y agoHugging Face22Sutranix /dsa-dbtabularn<1K0 likes1 downloads7mo agoHugging Face23Sutranix /dsa-db-tinytabularn<1K0 likes1 downloads7mo agoHugging Face24Raavi-Patil-Founder /sutra-ai0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.