CompiwerAI/MUD-Code3
🤖 MUD-Code3 – 2M Real‑World Code Tokens! This dataset contains 20483 code snippets (mixture of real and high-quality synthetic) totaling 2,000,018 tokens. Languages: Python (and some others) Sources: Real open‑source code (when available) plus enhanced synthetic Use cases: Fine‑tuning code models, program synthesis, code search 📖 How to Use from datasets import load_dataset dataset = load_dataset("CompiwerAI/MUD-Code3") print(dataset["train"][0]) 📊 Stats Metric Value Total Documents 20483… See the full description on the dataset page: https://huggingface.co/datasets/CompiwerAI/MUD-Code3.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face