CoolFace
Datasetpublic

kaushik-harsh-99/Code-Language-Classification

Programming Language Classification Dataset A large-scale, balanced dataset for programming language identification from source code snippets. Overview This dataset contains 1.664 million cleaned and labeled source code samples across 16 programming languages, specifically designed for programming language classification and identification tasks. Unlike many code datasets that are primarily built for code generation or retrieval, this dataset was curated… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-harsh-99/Code-Language-Classification.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
3likes53downloads
5 commits on main
a5d77224mo ago

minor fix

kaushik-harsh-99
591ac614mo ago

Update README.md

kaushik-harsh-99
c57c7da4mo ago

Update README.md

kaushik-harsh-99
0c1c0944mo ago

Upload 3 files

kaushik-harsh-99
74f9a6d4mo ago

initial commit

kaushik-harsh-99