compiwer-ai
MUD-Code1MUD-Code3🤖 MUD-Code3 – 2M Real‑World Code Tokens!
This dataset contains 20483 code snippets (mixture of real and high-quality synthetic) totaling 2,000,018 tokens.
Languages: Python (and some others)
Sources: Real open‑source code (when available) plus enhanced synthetic
Use cases: Fine‑tuning code models, program synthesis, code search
📖 How to Use
from datasets import load_dataset
dataset = load_dataset("CompiwerAI/MUD-Code3")
print(dataset["train"][0])
📊 Stats
Metric Value
Total Documents 20483… See the full description on the dataset page: https://huggingface.co/datasets/CompiwerAI/MUD-Code3.MUD-Code2MUD-Code4🚀 MUD-Code3-Mega – 10M Tokens of High‑Quality Code
This dataset contains 59244 synthetic code snippets designed to mimic real‑world production‑grade code across multiple domains (ML, web, async, data processing, deep learning, etc.). All samples are carefully crafted to be realistic, well‑structured, and high‑quality.
Total tokens: 10,000,112
Languages: Python (with some snippets including other languages like SQL, Dockerfile)
Quality: High – generated from expert‑level templates with… See the full description on the dataset page: https://huggingface.co/datasets/CompiwerAI/MUD-Code4.MUD-Public-v5
🌌 ULM-v5 Singularity (MUD-Public-v5)
ULM-v5 is the ultimate evolution of the Mtrini Used Dataset series, containing 3,000,000 high-entropy samples.
💎 Key Features
Reflective Thought Traces: Every sample includes an internal thought_trace and a reflective_audit for logical verification.
95+ Languages: Deep reasoning saturating global scripts and dialects.
120+ Specialized Domains: From Quantum Cryptography to Kernel-level exploit development.
Max Saturation:… See the full description on the dataset page: https://huggingface.co/datasets/CompiwerAI/MUD-Public-v5.
