CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01codeparrot /github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.text-generation420 likes40k downloads4y agoHugging Face02hasankursun /github-code-2025-language-split 📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models. Processing Steps To create this dataset, we performed the following processing on the source data: Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.text100M<n<1B13 likes25k downloads10mo agoHugging Face03codeparrot /github-code-cleanThe GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.text10M<n<100M143 likes25k downloads4y agoHugging Face04AdhyanshVerma /open-github-major-repos🌐 AdhyanshVerma's Open GitHub Major Repos An elite, curated collection of GitHub commit metadata from the world's most influential technology companies: Microsoft, Google, Meta, and Intel. 📖 Introduction Welcome to AdhyanshVerma's Open GitHub Major Repos dataset. This dataset focuses exclusively on high-impact, industry-standard repositories maintained by the world's leading technology giants. It utilizes the Lazy Pointer Pattern: instead of bloating your storage… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/open-github-major-repos.text-generation100K<n<1M1 likes13k downloads16d agoHugging Face05CodedotAI /code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.text1M<n<10M20 likes6k downloads4y agoHugging Face06nick007x /github-code-2025text100M<n<1B121 likes5.9k downloads6mo agoHugging Face07Yanbin99 /GITQA-Aug-Legacyimage100K<n<1M2 likes5.6k downloads3y agoHugging Face08Yanbin99 /GITQA-Aug-Pruned-Legacyimage1 likes5.5k downloads3y agoHugging Face09Yanbin99 /GITQA-Base-Pruned-Legacyimage2 likes5.3k downloads3y agoHugging Face10Yanbin99 /GITQA-Base-Legacyimage3 likes5.2k downloads3y agoHugging Face11SKT-NRS /GIT-SCRAPED 🚀 SKT-NRS / GIT-SCRAPED This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets. Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures. 📂 Repository Structure All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.1K<n<10K1 likes4.4k downloads3mo agoHugging Face12common-pile /github_archive GitHub Archive Description According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository. To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads. The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.texttext-generation10M<n<100M2 likes4.2k downloads1y agoHugging Face13ronantakizawa /github-top-code GitHub Top Developer Source Code A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025). This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers Dataset Summary 1.3M+ source code files from repositories across ~4,700 unique developers 80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more) Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.texttext-generation1M<n<10M125 likes4k downloads7mo agoHugging Face14mvaccargiu /gitskills GitSkills: A Dataset of Agent Skills on GitHub Paper (arXiv:2608.10906) · Sample repository · Zenodo DOI: 10.5281/zenodo.21875637 An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and reference files. The agent loads the skill when it judges that a task matches the skill description. Anthropic introduced the format in October 2025 as an open specification. Nine months later, skill files in the… See the full description on the dataset page: https://huggingface.co/datasets/mvaccargiu/gitskills.tabularother10M<n<100M38 likes3.8k downloads11d agoHugging Face15project-themis /git-commits Themis-Git-Commits Overview Themis-Git-Commits is a large-scale dataset of single-file code commits mined from permissively licensed GitHub repositories via the BigQuery GitHub public dataset. The SQL query restricts to repositories under permissive open-source licenses only (MIT, Apache-2.0, BSD-2/3-Clause, ISC, CC0-1.0, EPL-1.0, MPL-2.0, Unlicense, AGPL-3.0, LGPL-2.1, Artistic-2.0). The BigQuery snapshot used contains commits up to early 2022 —… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits.text-generation10M<n<100M1 likes3.7k downloads1mo agoHugging Face16jobs-git /StockChina-Minute A-Share Minute-Level Historical Data Dataset Description This dataset contains minute-level trading data for Chinese A-share stocks from 2005 to 2023, covering 5267 stocks with complete historical trading records. Data Format Each CSV file corresponds to one stock and contains the following fields: Field Description open Opening price close Closing price high Highest price low Lowest price volume Trading volume money Trading amount avg… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/StockChina-Minute.timeseries8 likes3.6k downloads1y agoHugging Face17code-rag-bench /github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos) text100K<n<1M2 likes3.3k downloads2y agoHugging Face18kevinS4455 /trellis500k-github-archives-70 likes3.3k downloads6mo agoHugging Face19hybridfree /github-code-2025 🚀 GitHub Code 2025: The Clean Code Manifesto A meticulously curated dataset of 1.5M+ repositories representing both quality and innovation in 2025's code ecosystem 🌟 The Philosophy Quality Over Quantity, Purpose Over Volume In an era of data abundance, we present a dataset built on radical curation. Every file, every repository, every byte has been carefully selected to represent the signal in the noise of open-source development. 🎯 What This Dataset Is… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/github-code-2025.text100M<n<1B1 likes3.3k downloads9mo agoHugging Face20nyuuzyou /gitee-code Gitee Code Dataset Dataset Description This dataset was compiled from code repositories hosted on Gitee, China's largest code hosting platform and a leading alternative to GitHub in the Chinese developer community. Gitee is widely used by Chinese developers, enterprises, and open-source projects, making this dataset particularly valuable for training code models with strong Chinese language understanding and Chinese coding conventions. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/gitee-code.texttext-generation100M<n<1B14 likes3.2k downloads9mo agoHugging Face21angie-chen55 /python-github-codetext1M<n<10M50 likes3k downloads4y agoHugging Face22OpenSynth /D-GITT-RTE7000-2023 README See the common Readme textn<1K1 likes2.8k downloads11mo agoHugging Face23wakamex /githubtext100K<n<1M0 likes2.7k downloads3y agoHugging Face24Brunobkr /llama.cpp_AlgMor24_github ΩFFFΣLLIa • llama.cpp • AlgMor24 ██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗ ██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗ ██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║ ██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║ ╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║ ╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝ High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.0 likes2.6k downloads1mo agoHugging Face25ruediste /codeparrot-github-code-10GThis is data is derived from the Codeparrot Dataset by taking the first 10GB of text from each language, and splitting it into individual configs. This results in a download size of about 3GB per language. Sample usage: from datasets import load_dataset dataset = load_dataset("ruediste/codeparrot-github-code-10G", "java") List of Languages: languages = { 'HTML': 'html', 'Java': 'java', 'JavaScript': 'js', 'CSS': 'css', 'C#': 'cs', 'TypeScript': 'ts', "Batchfile":… See the full description on the dataset page: https://huggingface.co/datasets/ruediste/codeparrot-github-code-10G.text10M<n<100M2 likes2.6k downloads2y agoHugging Face26hrishizone /Java-GitHub-Codestext1M<n<10M1 likes2.4k downloads1y agoHugging Face27angie-chen55 /javascript-github-codetext10M<n<100M18 likes2.3k downloads4y agoHugging Face28OpenSynth /D-GITT-RTE7000-2021 D-GITT_RTE 7000 Nodes Dataset We, at OpenSynth/D-GITT, are proud to announce that RTE released open-source the complete French electrical network data. D-GITT stands for Detailed Grid Inner Topology Time-series This first datasets provides a series of snapshots of the French transmission electricity network in node-breaker topology, with a temporal granularity of 5 minutes, covering the three-year period from January 2021 to December 2023. You'll find three datasets per year due to… See the full description on the dataset page: https://huggingface.co/datasets/OpenSynth/D-GITT-RTE7000-2021.text1K<n<10K10 likes2.3k downloads11mo agoHugging Face29bigcode /github-commits-diff-dedup-pjjs-april Deduplicated Commits Deduplicated based on diff: content = '\n'.join(difflib.unified_diff( old_content.splitlines(keepends=True), new_content.splitlines(keepends=True), n=5 )) Parameters: Minimum ngram size: 5 MinHash ngram size: 5 MinHash threshold: 0.8 text100K<n<1M4 likes1.7k downloads3y agoHugging Face30OpenSynth /D-GITT-RTE7000-2022 README See the common Readme text1K<n<10K1 likes1.7k downloads11mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.