CompiwerAI/MUD-Code4
🚀 MUD-Code3-Mega – 10M Tokens of High‑Quality Code This dataset contains 59244 synthetic code snippets designed to mimic real‑world production‑grade code across multiple domains (ML, web, async, data processing, deep learning, etc.). All samples are carefully crafted to be realistic, well‑structured, and high‑quality. Total tokens: 10,000,112 Languages: Python (with some snippets including other languages like SQL, Dockerfile) Quality: High – generated from expert‑level templates with… See the full description on the dataset page: https://huggingface.co/datasets/CompiwerAI/MUD-Code4.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face