CoolFace
Datasetpublic

codeparrot/codeparrot-clean

CodeParrot 🦜 Dataset Cleaned What is it? A dataset of Python files from Github. This is the deduplicated version of the codeparrot. Processing The original dataset contains a lot of duplicated and noisy data. Therefore, the dataset was cleaned with the following steps: Deduplication Remove exact matches Filtering Average line length < 100 Maximum line length < 1000 Alpha numeric characters fraction > 0.25 Remove auto-generated files (keyword… See the full description on the dataset page: https://huggingface.co/datasets/codeparrot/codeparrot-clean.

sourceHugging Faceupdated 4y agoView on Hugging Face
89likes41kdownloads
Dataset Card

CodeParrot 🦜 Dataset Cleaned

What is it?

A dataset of Python files from Github. This is the deduplicated version of the codeparrot.

Processing

The original dataset contains a lot of duplicated and noisy data. Therefore, the dataset was cleaned with the following steps:

  • Deduplication
  • Remove exact matches
  • Filtering
  • Average line length < 100
  • Maximum line length < 1000
  • Alpha numeric characters fraction > 0.25
  • Remove auto-generated files (keyword search)

For more details see the preprocessing script in the transformers repository here.

Splits

The dataset is split in a train and validation split used for training and evaluation.

Structure

This dataset has ~50GB of code and 5361373 files.

python
DatasetDict({
    train: Dataset({
        features: ['repo_name', 'path', 'copies', 'size', 'content', 'license', 'hash', 'line_mean', 'line_max', 'alpha_frac', 'autogenerated'],
        num_rows: 5361373
    })
})