CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WithinUsAI /CitationGround-1M CitationGround-1M (Platinum) Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z What this dataset is CitationGround-1M is a citation-locked grounded QA/RAG dataset: Answer using only the provided contexts Provide span-level citations (doc_id + offsets) Includes answerable=false hard negatives for abstention behavior Features / schema (JSONL) example_id (string) question (string) contexts (list of docs)… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/CitationGround-1M.question-answeringn<1K1 likes482 downloads9mo agoHugging Face02WithinUsAI /GOD_Coder_Complete_DataSet GOD_Coder_Complete_DataSet Subtitle A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants. Dataset Summary GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder. The dataset focuses on teaching models how to: diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.text-generation100K<n<1M5 likes353 downloads6mo agoHugging Face03WithinUsAI /Genesis_AI_Code_50k Genesis AI Code 50K (Expert) Developed by: Within Us AI Expert dataset with diff supervision, failure→reflection→correction, MoE labels, and governance flags. Splits train: 49,000 validation: 1,000 Highlights Tests-as-truth supervision patterns Diff-first patching Agentic loops (plan→edit→test→reflect) with bounded budgets Tool-call trace supervision (where present) Governance/audit & policy-gate awareness Storage format Parquet… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_50k.texttext-generation10K<n<100K0 likes85 downloads9mo agoHugging Face04WithinUsAI /Python_GOD_Coder_Omniforge_AI_12k Python GOD Coder Omniforge AI 12k Creator: Within Us AI A 12,000-row mixed-format Python coding dataset designed as a sharpening corpus for building a small but dangerous Python specialist. This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model: implementation with tests strict code-only instruction following debugging and repair refactoring for readability and production readiness next-token code completion… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Python_GOD_Coder_Omniforge_AI_12k.texttext-generation10K<n<100K1 likes84 downloads7mo agoHugging Face05WithinUsAI /Genesis_AI_Code_10k Genesis AI Code 10K Developed by: Within Us AI Foundation dataset emphasizing tests-as-truth, agentic loops, and evaluation thinking. Splits train: 9,800 validation: 200 Highlights Tests-as-truth supervision patterns Diff-first patching Agentic loops (plan→edit→test→reflect) with bounded budgets Tool-call trace supervision (where present) Governance/audit & policy-gate awareness Storage format Parquet unavailable (No module named… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_10k.text-generation10K<n<100K0 likes67 downloads9mo agoHugging Face06WithinUsAI /Genesis_AI_Code_100k Genesis AI Code 100K (Frontier) Developed by: Within Us AI Frontier dataset with tool-call traces, self-grading, budgets, and audit orientation. Splits train: 98,000 validation: 2,000 Highlights Tests-as-truth supervision patterns Diff-first patching Agentic loops (plan→edit→test→reflect) with bounded budgets Tool-call trace supervision (where present) Governance/audit & policy-gate awareness Storage format Parquet unavailable (No… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_100k.texttext-generation10K<n<100K1 likes62 downloads9mo agoHugging Face07WithinUsAI /Genesis_AI_Code_1k_Demo Genesis AI Code (Demo) 1K Developed by: Within Us AI Best-of demo subset for instant evaluation and fast adoption. Splits train: 1,000 validation: 1,000 Highlights Tests-as-truth supervision patterns Diff-first patching Agentic loops (plan→edit→test→reflect) with bounded budgets Tool-call trace supervision (where present) Governance/audit & policy-gate awareness Storage format Parquet unavailable (No module named 'pyarrow'); JSONL… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_1k_Demo.texttext-generation1K<n<10K0 likes50 downloads9mo agoHugging Face08WithinUsAI /OpenToolTrace-X OpenToolTrace-X (Platinum) Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z What this dataset is OpenToolTrace-X is a replayable, verifiable corpus of tool-using agent trajectories. Each record contains: A user goal (prompt) and constraints An initial_state describing the starting environment/repo snapshot A trajectory (tool calls + observations) A final_state (artifacts/diff/output) verification (tests, checksums, exit… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/OpenToolTrace-X.texttext-generationn<1K1 likes48 downloads9mo agoHugging Face09WithinUsAI /Royal_Ghost_Coder_10M Royal Ghost Coder 10M A large-scale, synthetic instruction-tuning corpus designed to train code-capable, agentic models on structured “instruction → input → output” workflows at high volume. The dataset ships as a single JSONL file and is auto-converted to Parquet by Hugging Face for faster streaming. Dataset Summary Repository: gss1147/Royal_Ghost_Coder_10M Rows: 10,000,000 (train split) Primary file: royal_ghost_titan_data.jsonl Format: JSON Lines (one JSON… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Royal_Ghost_Coder_10M.text-generation10M<n<100M0 likes29 downloads8mo agoHugging Face10WithinUsAI /All_Known_Species_Index_25k WithinUsAI/All_Known_Species_Index_25k — Master Scholars Academics (25k) A real-species index dataset sampled from NCBI Taxonomy (taxdump) at the species rank, formatted for Tiny-Recursive-Model-friendly fine-tuning. What this is (and is not) This dataset contains 25,000 real species records sampled from the NCBI taxonomy graph (rank=species). It does not enumerate every described species name (the full list is far larger and updated continuously).If you want… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/All_Known_Species_Index_25k.text-generation10K<n<100K0 likes26 downloads9mo agoHugging Face11WithinUsAI /Sports_25k WithinUsAI/Sports_25k — Master Scholars Academics (25k) This dataset is designed for academic-grade fine-tuning of LLMs on sports rules, sports science, and quantitative sports analytics with a Tiny-Recursive-Model-friendly structure. What’s inside (25,000 examples) Task mix (fixed): 7,000 Fact-check / verification items (wrapper=verify_true_false, truth_mode=verifiable_fact) 10,000 Self-contained quantitative reasoning items (wrapper=minimal_chain… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Sports_25k.question-answering10K<n<100K1 likes25 downloads9mo agoHugging Face12WithinUsAI /Human_25k WithinUsAI/Human_25k — Master Scholars Academics (25k) Academic-grade dataset for fine-tuning LLMs on human anatomy, physiology, clinical concepts, and biomedical quantitative reasoning in a Tiny-Recursive-Model-friendly format. What’s inside (25,000 examples) Task mix (fixed): 8,000 Fact-check / verification (wrapper=verify_true_false, truth_mode=verifiable_fact) 9,000 Self-contained quantitative reasoning (wrapper=minimal_chain, truth_mode=self_contained_math)… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Human_25k.question-answering10K<n<100K1 likes22 downloads9mo agoHugging Face13WithinUsAI /gpt2_to_gpt5.5_distilled_25k GPT-2 to GPT-5.5 Advanced Reasoning Distillation (25k) Dataset Description 25,000 unique, high-quality instruction-response pairs designed for knowledge distillation and supervised fine-tuning. The dataset elevates GPT-2 Medium toward GPT-5.5-level performance on complex reasoning tasks. Core goal: Transfer frontier reasoning capabilities (multi-step CoT, cross-domain synthesis, edge-case analysis, novel insights) from a hypothetical GPT-5.5 teacher into smaller… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gpt2_to_gpt5.5_distilled_25k.texttext-generation10K<n<100K1 likes22 downloads4mo agoHugging Face14WithinUsAI /DEEPMIND_Alpha_Distilled AlphaMirror-25k: DeepMind Alpha-Inspired Reasoning Dataset for LLM Fine-Tuning Dataset Summary AlphaMirror-25k is a high-quality synthetic dataset containing exactly 25,000 instruction-response pairs designed to fine-tune any large language model to mirror the advanced reasoning, scientific discovery, and problem-solving style of Google DeepMind's Alpha series (AlphaFold, AlphaGo/AlphaZero, AlphaEvolve, AlphaGeometry, AlphaTensor, and related systems). The dataset… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DEEPMIND_Alpha_Distilled.texttext-generation1K<n<10K0 likes18 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.