datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
github-codeThe GitHub Code dataest consists of 115M code files from GitHub in 32 programming languages with 60 extensions totalling in 1TB of text data. The dataset was created from the GitHub dataset on BiqQuery.github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.github-code-cleanThe GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in almost 1TB of text data.open-github-major-repos🌐 AdhyanshVerma's Open GitHub Major Repos
An elite, curated collection of GitHub commit metadata from the world's most influential technology companies: Microsoft, Google, Meta, and Intel.
📖 Introduction
Welcome to AdhyanshVerma's Open GitHub Major Repos dataset. This dataset focuses exclusively on high-impact, industry-standard repositories maintained by the world's leading technology giants.
It utilizes the Lazy Pointer Pattern: instead of bloating your storage… See the full description on the dataset page: https://huggingface.co/datasets/AdhyanshVerma/open-github-major-repos.code_clippy_githubThe Code Clippy dataset consists of various public codebases from GitHub in 22 programming languages with 23 extensions totalling about 16 TB of data when uncompressed. The dataset was created from the public GitHub dataset on Google BiqQuery.github-code-2025GITQA-Aug-LegacyGITQA-Aug-Pruned-LegacyGITQA-Base-Pruned-LegacyGITQA-Base-LegacyGIT-SCRAPED
🚀 SKT-NRS / GIT-SCRAPED
This repository is dedicated to hosting structural, curated, and diverse datasets—including GitHub roadmaps,and Roadmaps.sh Sites system architectures, technical diagrams, and mass scraped assets.
Our ultimate mission is to fuel the development of next-generation Sovereign Indian Intelligence base models with high-fidelity, production-grade text-image structures.
📂 Repository Structure
All the raw and structured crawled data is… See the full description on the dataset page: https://huggingface.co/datasets/SKT-NRS/GIT-SCRAPED.git-commits
Themis-Git-Commits
Overview
Themis-Git-Commits is a large-scale dataset of single-file code commits mined from permissively licensed GitHub repositories via the BigQuery GitHub public dataset. The SQL query restricts to repositories under permissive open-source licenses only (MIT, Apache-2.0, BSD-2/3-Clause, ISC, CC0-1.0, EPL-1.0, MPL-2.0, Unlicense, AGPL-3.0, LGPL-2.1, Artistic-2.0). The BigQuery snapshot used contains commits up to early 2022 —… See the full description on the dataset page: https://huggingface.co/datasets/project-themis/git-commits.github-top-code
GitHub Top Developer Source Code
A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025).
This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers
Dataset Summary
1.3M+ source code files from repositories across ~4,700 unique developers
80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)
Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.gitskills
GitSkills: A Dataset of Agent Skills on GitHub
Paper (arXiv:2608.10906) ·
Sample repository ·
Zenodo DOI: 10.5281/zenodo.21875637
An agent skill is a folder containing a SKILL.md file with instructions for
a language-model agent, optionally accompanied by scripts and reference
files. The agent loads the skill when it judges that a task matches the
skill description. Anthropic introduced the format in October 2025 as an
open specification. Nine months later, skill files in the… See the full description on the dataset page: https://huggingface.co/datasets/mvaccargiu/gitskills.StockChina-Minute
A-Share Minute-Level Historical Data
Dataset Description
This dataset contains minute-level trading data for Chinese A-share stocks from 2005 to 2023, covering 5267 stocks with complete historical trading records.
Data Format
Each CSV file corresponds to one stock and contains the following fields:
Field
Description
open
Opening price
close
Closing price
high
Highest price
low
Lowest price
volume
Trading volume
money
Trading amount
avg… See the full description on the dataset page: https://huggingface.co/datasets/jobs-git/StockChina-Minute.github-repos-pythonThe Github repository retrieval source for [code-rag-bench], containing all Python files from the entire GitHub dump (in github-repos)
trellis500k-github-archives-7github-code-2025
🚀 GitHub Code 2025: The Clean Code Manifesto
A meticulously curated dataset of 1.5M+ repositories representing both quality and innovation in 2025's code ecosystem
🌟 The Philosophy
Quality Over Quantity, Purpose Over Volume
In an era of data abundance, we present a dataset built on radical curation. Every file, every repository, every byte has been carefully selected to represent the signal in the noise of open-source development.
🎯 What This Dataset Is… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/github-code-2025.python-github-codegitee-code
Gitee Code Dataset
Dataset Description
This dataset was compiled from code repositories hosted on Gitee, China's largest code hosting platform and a leading alternative to GitHub in the Chinese developer community. Gitee is widely used by Chinese developers, enterprises, and open-source projects, making this dataset particularly valuable for training code models with strong Chinese language understanding and Chinese coding conventions.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/gitee-code.Java-GitHub-Codesgithub_archive
GitHub Archive
Description
According to GitHub’s terms of service, issues and pull request descriptions—along with the their comments—inherit the license of their associated repository.
To collect this data, we used the GitHub Archive’s public BigQuery table of events to extracted all issue, pull request, and comment events since 2011 and aggregated them into threads.
The table appeared to be missing “edit” events so the text from each comment is the original from when… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/github_archive.llama.cpp_AlgMor24_github
ΩFFFΣLLIa • llama.cpp • AlgMor24
██████╗ ███████╗███████╗███████╗██╗ ██╗ ██╗ █████╗
██╔═══██╗██╔════╝██╔════╝██╔════╝██║ ██║ ██║██╔══██╗
██║ ██║█████╗ █████╗ █████╗ ██║ ██║ ██║███████║
██║ ██║██╔══╝ ██╔══╝ ██╔══╝ ██║ ██║ ██║██╔══██║
╚██████╔╝██║ ██║ ███████╗███████╗███████╗██║██║ ██║
╚═════╝ ╚═╝ ╚═╝ ╚══════╝╚══════╝╚══════╝╚═╝╚═╝ ╚═╝
High-Performance LLM / VLM Inference & Autonomous Agentic Ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/Brunobkr/llama.cpp_AlgMor24_github.D-GITT-RTE7000-2023
README
See the common Readme
githubcodeparrot-github-code-10GThis is data is derived from the Codeparrot Dataset by taking the first 10GB of text from each language, and splitting it into individual configs. This results in a download size of about 3GB per language.
Sample usage:
from datasets import load_dataset
dataset = load_dataset("ruediste/codeparrot-github-code-10G", "java")
List of Languages:
languages = {
'HTML': 'html',
'Java': 'java',
'JavaScript': 'js',
'CSS': 'css',
'C#': 'cs',
'TypeScript': 'ts',
"Batchfile":… See the full description on the dataset page: https://huggingface.co/datasets/ruediste/codeparrot-github-code-10G.D-GITT-RTE7000-2021
D-GITT_RTE 7000 Nodes Dataset
We, at OpenSynth/D-GITT, are proud to announce that RTE released open-source the complete French electrical network data. D-GITT stands for Detailed Grid Inner Topology Time-series
This first datasets provides a series of snapshots of the French transmission electricity network in node-breaker topology, with a temporal granularity of 5 minutes, covering the three-year period from January 2021 to December 2023.
You'll find three datasets per year due to… See the full description on the dataset page: https://huggingface.co/datasets/OpenSynth/D-GITT-RTE7000-2021.javascript-github-codegithub-commits-diff-dedup-pjjs-april
Deduplicated Commits
Deduplicated based on diff:
content = '\n'.join(difflib.unified_diff(
old_content.splitlines(keepends=True),
new_content.splitlines(keepends=True),
n=5
))
Parameters:
Minimum ngram size: 5
MinHash ngram size: 5
MinHash threshold: 0.8
Git-10MThe Git-10M dataset is a global-scale remote sensing image-text pair dataset, consisting of over 10 million image-text pairs with geographical locations and resolution information.
CC-BY-NC-ND-4.0 License: This dataset is not allowed to be modified or distributed without authorization!
Project Page: https://chen-yang-liu.github.io/Text2Earth/
View samples from the dataset
from datasets import load_dataset
import math
def XYZToLonLat(x,y,z):
# Transform… See the full description on the dataset page: https://huggingface.co/datasets/lcybuaa/Git-10M.
