CoolFace
Datasetpublic

mvasiliniuc/iva-swift-codeint-clean

IVA Swift GitHub Code Dataset Dataset Description This is the curated IVA Swift dataset extracted from GitHub. It contains curated Swift files gathered with the purpose to train a code generation model. The dataset consists of 383380 swift code files from GitHub totaling ~542MB of data. The uncurated dataset was created from the public GitHub dataset on Google BiqQuery. How to use it To download the full dataset: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean.

sourceHugging Faceotherupdated 3y agoView on Hugging Face
7likes55downloads
Dataset Card

IVA Swift GitHub Code Dataset

Dataset Description

This is the curated IVA Swift dataset extracted from GitHub. It contains curated Swift files gathered with the purpose to train a code generation model.

The dataset consists of 383380 swift code files from GitHub totaling ~542MB of data. The uncurated dataset was created from the public GitHub dataset on Google BiqQuery.

How to use it

To download the full dataset:

python
from datasets import load_dataset
dataset = load_dataset('mvasiliniuc/iva-swift-codeint-clean', split='train')
python

from datasets import load_dataset
dataset = load_dataset('mvasiliniuc/iva-swift-codeint-clean', split='train')
print(dataset[723])

#OUTPUT:
{
   "repo_name":"jdkelley/Udacity-OnTheMap-ExampleApps",
   "path":"TheMovieManager-v2/TheMovieManager/BorderedButton.swift",
   "copies":"2",
   "size":"2649",
   "content":"...let phoneBorderedButtonExtraPadding: CGFloat = 14.0\n    \n    var backingColor: UIColor? = nil\n    var highlightedBackingColor: UIColor? = nil\n    \n    // MARK: Initialization\n}",
   "license":"mit",
   "hash":"db1587fd117e9a835f58cf8203d8bf05",
   "line_mean":29.1136363636,
   "line_max":87,
   "alpha_frac":0.6700641752,
   "ratio":5.298,
   "autogenerated":false,
   "config_or_test":false,
   "has_no_keywords":false,
   "has_few_assignments":false
}

Data Structure

Data Fields

FieldTypeDescription
repo_namestringname of the GitHub repository
pathstringpath of the file in GitHub repository
copiesstringnumber of occurrences in dataset
contentstringcontent of source file
sizestringsize of the source file in bytes
licensestringlicense of GitHub repository
hashstringHash of content field.
line_meannumberMean line length of the content.
line_maxnumberMax line length of the content.
alpha_fracnumberFraction between mean and max line length of content.
rationumberCharacter/token ratio of the file with tokenizer.
autogeneratedbooleanTrue if the content is autogenerated by looking for keywords in the first few lines of the file.
configortestbooleanTrue if the content is a configuration file or a unit test.
hasnokeywordsbooleanTrue if a file has none of the keywords for Swift Programming Language.
hasfewassignmentsbooleanTrue if file uses symbol '=' less than minimum times.

Instance

json
{
   "repo_name":"...",
   "path":".../BorderedButton.swift",
   "copies":"2",
   "size":"2649",
   "content":"...",
   "license":"mit",
   "hash":"db1587fd117e9a835f58cf8203d8bf05",
   "line_mean":29.1136363636,
   "line_max":87,
   "alpha_frac":0.6700641752,
   "ratio":5.298,
   "autogenerated":false,
   "config_or_test":false,
   "has_no_keywords":false,
   "has_few_assignments":false
}

Languages

The dataset contains only Swift files.

json
{
    "Swift": [".swift"]
}

Licenses

Each entry in the dataset contains the associated license. The following is a list of licenses involved and their occurrences.

json
{
   "agpl-3.0":1695,
   "apache-2.0":85514,
   "artistic-2.0":207,
   "bsd-2-clause":3132,
   "bsd-3-clause":6600,
   "cc0-1.0":1409,
   "epl-1.0":605,
   "gpl-2.0":9374,
   "gpl-3.0":18920,
   "isc":808,
   "lgpl-2.1":1122,
   "lgpl-3.0":3103,
   "mit":240929,
   "mpl-2.0":8181,
   "unlicense":1781
}

Dataset Statistics

json
{
    "Total size": "~542 MB",
    "Number of files": 383380,
    "Number of files under 500 bytes": 3680,
    "Average file size in bytes": 5942,
}

Curation Process

  • Removal of duplication files based on file hash.
  • Removal of file templates. File containing the following: ___FILENAME___, ___PACKAGENAME___, ___FILEBASENAME___, ___FILEHEADER___, ___VARIABLE
  • Removal of the files containing the following words in the first 10 lines: generated, auto-generated", "autogenerated", "automatically generated
  • Removal of the files containing the following words in the first 10 lines with a probability of 0.7: test", "unit test", "config", "XCTest", "JUnit
  • Removal of file with the rate of alphanumeric characters below 0.3 of the file.
  • Removal of near duplication based MinHash and Jaccard similarity.
  • Removal of files with mean line length above 100.
  • Removal of files without mention of keywords with a probability of 0.7: struct ", "class ", "for ", "while ", "enum ", "func ", "typealias ", "var ", "let ", "protocol ", "public ", "private ", "internal ", "import "
  • Removal of files that use the assignment operator = less than 3 times.
  • Removal of files with the ratio between the number of characters and number of tokens after tokenization lower than 1.5.

Curation process is a derivation of the one used in CodeParrot project: https://huggingface.co/codeparrot

Data Splits

The dataset only contains a train split which is separated into train and valid which can be found here:

  • Clean Version Train: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean-train
  • Clean Version Valid: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean-valid

Considerations for Using the Data

The dataset comprises source code from various repositories, potentially containing harmful or biased code, along with sensitive information such as passwords or usernames.