Gugu8/Code-Syntax-Expanded
Code-Syntax-Expanded A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations β no artificial padding, no duplicate rows. π Dataset Overview Property Value Total rows 5,000,000+ File size ~1.1 GB (uncompressed CSV) Languages 33 Unique templates 160+ error patterns Format CSV (4 columns) Licenseβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Gugu8/Code-Syntax-Expanded.
Code-Syntax-Expanded
A massive, high-quality synthetic dataset for training LLMs to identify and correct syntax errors across 33 programming languages. Contains 5+ million unique examples (~1.1 GB) with English explanations β no artificial padding, no duplicate rows.
π Dataset Overview
π Dataset Structure
Each row contains:
π§ Languages Covered
General Purpose
Python, JavaScript, TypeScript, Java, C#, C++, Rust, Go, Ruby, PHP, Perl, Swift, Kotlin, R, MATLAB, Scala, Lua, Dart
Web Technologies
HTML, CSS
Database
SQL (MySQL/PostgreSQL-style)
Shell & Scripting
Bash, PowerShell
Markup & Configuration
YAML, JSON, XML, Markdown
Backend & Frameworks
Node.js (Express, fs, JWT, bcrypt, Mongoose)
π Dataset Statistics
π― Use Cases
- Fine-tuning LLMs for code correction and syntax repair
- Building code review assistants that detect common mistakes
- Teaching programming with a massive bank of error/fix examples
- Benchmarking model understanding of language-specific syntax
- Creating educational tools for learning programming languages
