CoolFace
Datasetpublic

mvasiliniuc/iva-swift-codeint-clean-train

IVA Swift GitHub Code Dataset Dataset Description This is the curated train split of IVA Swift dataset extracted from GitHub. It contains curated Swift files gathered with the purpose to train a code generation model. The dataset consists of 320000 Swift code files from GitHub. Here is the unsliced curated dataset and here is the raw dataset. How to use it To download the full dataset: from datasets import load_dataset dataset =… See the full description on the dataset page: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean-train.

sourceHugging Faceotherupdated 3y agoView on Hugging Face
3likes84downloads
Dataset Card

IVA Swift GitHub Code Dataset

Dataset Description

This is the curated train split of IVA Swift dataset extracted from GitHub. It contains curated Swift files gathered with the purpose to train a code generation model.

The dataset consists of 320000 Swift code files from GitHub. Here is the unsliced curated dataset and here is the raw dataset.

How to use it

To download the full dataset:

python
from datasets import load_dataset
dataset = load_dataset('mvasiliniuc/iva-swift-codeint-clean', split='train')

Data Structure

Data Fields

FieldTypeDescription
repo_namestringname of the GitHub repository
pathstringpath of the file in GitHub repository
copiesstringnumber of occurrences in dataset
contentstringcontent of source file
sizestringsize of the source file in bytes
licensestringlicense of GitHub repository
hashstringHash of content field.
line_meannumberMean line length of the content.
line_maxnumberMax line length of the content.
alpha_fracnumberFraction between mean and max line length of content.
rationumberCharacter/token ratio of the file with tokenizer.
autogeneratedbooleanTrue if the content is autogenerated by looking for keywords in the first few lines of the file.
configortestbooleanTrue if the content is a configuration file or a unit test.
hasnokeywordsbooleanTrue if a file has none of the keywords for Swift Programming Language.
hasfewassignmentsbooleanTrue if file uses symbol '=' less than minimum times.

Instance

json
{
   "repo_name":"...",
   "path":".../BorderedButton.swift",
   "copies":"2",
   "size":"2649",
   "content":"...",
   "license":"mit",
   "hash":"db1587fd117e9a835f58cf8203d8bf05",
   "line_mean":29.1136363636,
   "line_max":87,
   "alpha_frac":0.6700641752,
   "ratio":5.298,
   "autogenerated":false,
   "config_or_test":false,
   "has_no_keywords":false,
   "has_few_assignments":false
}

Languages

The dataset contains only Swift files.

json
{
    "Swift": [".swift"]
}

Licenses

Each entry in the dataset contains the associated license. The following is a list of licenses involved and their occurrences.

json
{
   "agpl-3.0":1415,
   "apache-2.0":71451,
   "artistic-2.0":169,
   "bsd-2-clause":2628,
   "bsd-3-clause":5492,
   "cc0-1.0":1176,
   "epl-1.0":498,
   "gpl-2.0":7846,
   "gpl-3.0":15716,
   "isc":676,
   "lgpl-2.1":932,
   "lgpl-3.0":2553,
   "mit":201134,
   "mpl-2.0":6846,
   "unlicense":1468
}

Dataset Statistics

json
{
    "Total size": "~453 MB",
    "Number of files": 320000,
    "Number of files under 500 bytes": 3116,
    "Average file size in bytes": 5940,
}

Curation Process

See the unsliced curated dataset for mode details.

Data Splits

The dataset only contains a train split focused only on training data. For validation and unspliced versions, please check the following links:

  • Clean Version Unsliced: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean
  • Clean Version Valid: https://huggingface.co/datasets/mvasiliniuc/iva-swift-codeint-clean-valid

Considerations for Using the Data

The dataset comprises source code from various repositories, potentially containing harmful or biased code, along with sensitive information such as passwords or usernames.