CoolFace
20 results

github

codeparrot /github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.text-generation420 likes40k downloads4y agoHugging Facehasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes26k downloads10mo agoHugging Facecodeparrot /github-code-cleanThe GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.text10M<n<100M143 likes25k downloads4y agoHugging FaceAdhyanshVerma /open-github-major-repos🌐 AdhyanshVerma's Open GitHub Major Repos An elite, curated collection of GitHub commit metadata from the world's most influential technology companies: Microsoft, Google, Meta, and Intel. 📖 Introduction Welcome to AdhyanshVerma's Open GitHub Major Repos dataset. This dataset focuses exclusively on high-impact, industry-standard repositories maintained by the world's leading technology giants. It utilizes the Lazy Pointer Pattern: instead of bloating your storage… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/open-github-major-repos.text-generation100K<n<1M1 likes13k downloads15d agoHugging FaceCodedotAI /code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.text1M<n<10M20 likes6.4k downloads4y agoHugging Facenick007x /github-code-2025text100M<n<1B121 likes6k downloads6mo agoHugging Face