withinusai
claude_mythos_distilled_25k
Claude Mythos Distilled 25K
A high-quality synthetic supervised fine-tuning (SFT) dataset designed to train and fine-tune any LLM to mirror the capabilities, reasoning style, agentic behavior, and technical depth of Anthropic's Claude Mythos (distilled frontier model).
Dataset Summary
Size: 25,000 high-quality examples
Format: JSONL with chat messages (user/assistant pairs) + rich metadata
Categories (balanced for general + specialized capability):
Cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/claude_mythos_distilled_25k.CitationGround-1M
CitationGround-1M (Platinum)
Developer/Publisher: Within US AIVersion: 0.1.0 (sample pack)Created: 2025-12-30T16:53:41Z
What this dataset is
CitationGround-1M is a citation-locked grounded QA/RAG dataset:
Answer using only the provided contexts
Provide span-level citations (doc_id + offsets)
Includes answerable=false hard negatives for abstention behavior
Features / schema (JSONL)
example_id (string)
question (string)
contexts (list of docs)… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/CitationGround-1M.GOD_Coder_Complete_DataSet
GOD_Coder_Complete_DataSet
Subtitle
A large-scale complete-project coding dataset by gss1147 / WithIn Us AI, built to train language models into stronger professional software-engineering assistants.
Dataset Summary
GOD_Coder_Complete_DataSet is a large synthetic supervised fine-tuning dataset designed to help turn a general language model into a professional complete-project AI coder.
The dataset focuses on teaching models how to:
diagnose… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/GOD_Coder_Complete_DataSet.gemini_3.5_flash_distilled_25k
Gemini 3.5 Flash Distilled Dataset (25k)
A 25,000-sample synthetic distilled dataset designed to replicate the core capabilities of Gemini 3.5 Flash: frontier-level agentic execution, rapid multi-step reasoning, dense context analysis, and advanced autonomous coding — all optimized for low-latency inference.
Dataset Summary
This dataset was created via template-based evolutionary synthesis with content-normalized SHA-256 deduplication. Every sample features… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/gemini_3.5_flash_distilled_25k.DeepSeek_V4_Flash_distilled_dataset_5k
DeepSeek V4 Flash — Distilled Reasoning Dataset
A synthetic dataset of 5,099 unique reasoning traces designed to mirror the step-by-step thinking style of DeepSeek V4 Flash. Generated entirely with template-based parameterized generation (no LLM API calls).
Format
JSONL (one JSON object per line):
{
"id": "ds4f_math_000042",
"domain": "mathematics",
"subdomain": "algebra",
"difficulty": "easy",
"prompt": "Solve 3x + 7 = 22.",
"reasoning_trace":… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/DeepSeek_V4_Flash_distilled_dataset_5k.P_P_GHOST_1Million
P_P_GHOST — 1,000,000 record pack (sharded)
This repository contains 1,000,000 JSONL records (gzipped shards) to train/evaluate prompt orchestration that creates the illusion of instant fine-tuning via:
prompt packets / compilation
routing + retrieval grounding (RAG)
tool-use loops (ReAct)
schema enforcement + reasking/repair loops
end-of-pass meta-optimization (self-refine style)
Truth policy
Each record references at least one evidence capsule with a public… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/P_P_GHOST_1Million.
