CoolFace
Datasetpublic

nyuuzyou/gitee-code

Gitee Code Dataset Dataset Description This dataset was compiled from code repositories hosted on Gitee, China's largest code hosting platform and a leading alternative to GitHub in the Chinese developer community. Gitee is widely used by Chinese developers, enterprises, and open-source projects, making this dataset particularly valuable for training code models with strong Chinese language understanding and Chinese coding conventions. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/gitee-code.

sourceHugging Faceotherupdated 9mo agoView on Hugging Face
14likes2.9kdownloads
Dataset Card

Gitee Code Dataset

Dataset Description

This dataset was compiled from code repositories hosted on Gitee, China's largest code hosting platform and a leading alternative to GitHub in the Chinese developer community. Gitee is widely used by Chinese developers, enterprises, and open-source projects, making this dataset particularly valuable for training code models with strong Chinese language understanding and Chinese coding conventions.

Dataset Summary

StatisticValue
Total Files819,472,785
Total Repositories3,105,923
Total Size536 GB (compressed Parquet)
Programming Languages554
File FormatParquet with Zstd compression (468 files)

Key Features

  • Large-scale Chinese code corpus: Contains code from over 3 million repositories, many featuring Chinese comments, documentation, and variable names
  • Diverse language coverage: Spans 554 programming languages identified by go-enry (based on GitHub Linguist rules)
  • Rich metadata: Includes repository name, file path, detected language, license information, and file size
  • Enterprise and open-source projects: Includes code from both individual developers and Chinese enterprises
  • Quality filtered: Extensive filtering to remove vendor code, build artifacts, generated files, and low-quality content

Languages

The dataset includes 554 programming languages. The top 30 languages by file count:

RankLanguageFile Count
1Java293,439,777
2JavaScript77,715,425
3C62,836,721
4C++49,134,251
5HTML46,191,063
6Vue40,468,646
7PHP37,132,954
8C#33,842,369
9Python25,192,704
10CSS20,802,464
11TypeScript20,122,528
12Go16,176,561
13Shell8,371,429
14Makefile6,341,964
15Java Server Pages6,224,523
16TSX5,768,542
17CMake5,581,774
18SCSS5,291,031
19Objective-C4,922,736
20Less4,669,672
21Ruby3,027,385
22Kotlin2,986,211
23Scala2,869,640
24Rust2,466,122
25Starlark2,027,514
26Dart2,010,079
27Unix Assembly1,900,320
28Fluent1,882,380
29HTML+Razor1,863,914
30Swift1,607,477

Licenses

The dataset includes files from repositories with various licenses. Repositories with restrictive licenses (CC-BY-ND variants, Commons Clause, SSPL) were excluded:

LicenseFile Count
apache-2.0273,706,950
mit201,880,040
unknown195,868,240
agpl-3.060,181,320
bsd30,013,190
gpl-2.027,831,530
lgpl-3.011,746,750
lgpl-2.14,807,600
bsd-3-clause4,442,480
cc0-1.03,144,920
gpl-3.01,631,590
unlicense1,181,930
bsd-2-clause1,154,300
epl-1.01,045,470
Other licenses~5,800,000

Dataset Structure

Data Fields

FieldTypeDescription
codestringContent of the source file (UTF-8 encoded)
repo_namestringName of the Gitee repository (format: username/repo)
pathstringPath of the file within the repository (relative to repo root)
languagestringProgramming language as identified by go-enry
licensestringLicense of the repository (SPDX identifier or "unknown")
sizeint64Size of the source file in bytes

Data Format

  • Format: Apache Parquet with Zstd compression
  • File Structure: 468 files (gitee_0000.parquet to gitee_0467.parquet)

Data Splits

All examples are in the train split. There is no validation or test split.

Example Data Point

{
    'code': 'package com.example.demo;\n\nimport org.springframework.boot.SpringApplication;\n...',
    'repo_name': 'username/spring-demo',
    'path': 'src/main/java/com/example/demo/Application.java',
    'language': 'Java',
    'license': 'apache-2.0',
    'size': 1234
}

Dataset Creation

Pipeline Overview

The dataset was created through a multi-stage pipeline:

  1. 1.Repository Discovery
  2. 2.Branch Selection: Selecting the main branch for each repository (priority: master > main > develop > dev > first branch)
  3. 3.Repository Downloading
  4. 4.Content Extraction: Extracting and filtering source code files
  5. 5.Parquet Generation: Writing filtered records to Parquet shards with Zstd compression

Language Detection

Programming languages are detected using go-enry, a Go port of GitHub's Linguist library. Only files classified as Programming or Markup language types are included (Data and Prose types are excluded).

License Detection

Licenses are detected by:

  1. 1.Scanning for license files (LICENSE, LICENSE.txt, LICENSE.md, COPYING, etc.)
  2. 2.Matching license text against known patterns (MIT, Apache 2.0, GPL variants, BSD, Creative Commons, etc.)
  3. 3.Defaulting to "unknown" if no license can be detected

Blocked Licenses: The following restrictive licenses are excluded from the dataset:

  • cc-by-nd, cc-by-nd-2.0, cc-by-nd-3.0, cc-by-nd-4.0 (Creative Commons No-Derivatives)
  • commons-clause
  • sspl, sspl-1.0 (Server Side Public License)

File Filtering

Extensive filtering is applied to ensure data quality:

Size Limits
LimitValue
Max repository ZIP size48 MB
Max single file size1 MB
Max line length1,000 characters
Excluded Directories
  • Configuration: .git/, .github/, .gitlab/, .vscode/, .idea/, .vs/, .settings/, .eclipse/, .project/, .metadata/
  • Vendor/Dependencies: node_modules/, bower_components/, jspm_packages/, vendor/, third_party/, 3rdparty/, external/, packages/, deps/, lib/vendor/, target/dependency/, Pods/
  • Build Output: build/, dist/, out/, bin/, target/, release/, debug/, .next/, .nuxt/, _site/, _build/, __pycache__/, .pytest_cache/, cmake-build-*, .gradle/, .maven/
Excluded Files
  • Lock Files: package-lock.json, yarn.lock, pnpm-lock.yaml, Gemfile.lock, Cargo.lock, poetry.lock, Pipfile.lock, composer.lock, go.sum, mix.lock
  • Minified Files: Any file containing .min. in the name
  • Binary Files: .exe, .dll, .so, .dylib, .a, .lib, .o, .obj, .jar, .war, .ear, .class, .pyc, .pyo, .wasm, .bin, .dat, .pdf, .doc, .docx, .xls, .xlsx, .ppt, .pptx, .zip, .tar, .gz, .bz2, .7z, .rar, .jpg, .jpeg, .png, .gif, .bmp, .ico, .svg, .mp3, .mp4, .avi, .mov, .wav, .flac, .ttf, .otf, .woff, .woff2, .eot
  • System Files: .DS_Store, thumbs.db
Content Filtering
  • UTF-8 Validation: Files must be valid UTF-8 encoded text
  • Binary Detection: Files detected as binary by go-enry are excluded
  • Generated Files: Files with generation markers in the first 500 bytes are excluded:
  • generated by, do not edit, auto-generated, autogenerated, automatically generated, code generator, generated code, this file is generated, @generated, <auto-generated
  • Empty Files: Files that are empty or contain only whitespace are excluded
  • Long Lines: Files with any line exceeding 1,000 characters are excluded
  • go-enry Filters: Additional filtering using go-enry's IsVendor(), IsImage(), IsDotFile(), IsTest(), and IsGenerated() functions
  • Documentation-only Repos: Repositories containing only documentation files (no actual code) are skipped

Source Data

All data originates from public repositories hosted on Gitee.

Considerations for Using the Data

Personal and Sensitive Information

The dataset may contain:

  • Email addresses in code comments or configuration files
  • API keys or credentials that were accidentally committed
  • Personal information in comments or documentation

Users should exercise caution and implement appropriate filtering when using this data.

Licensing Information

This dataset is a collection of source code from repositories with various licenses. Any use of all or part of the code gathered in this dataset must abide by the terms of the original licenses, including attribution clauses when relevant. The license field in each data point indicates the license of the source repository.