CoolFace
Datasetpublic

ammarnasr/the-stack-swift-clean

Dataset 1: TheStack - Swift - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Swift, a popular statically typed language. Target Language: Swift Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Swift as the target language due to… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-swift-clean.

sourceHugging Faceopenrailupdated 3y agoView on Hugging Face
7likes295downloads
settings

This repository belongs to ammarnasr on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namethe-stack-swift-clean
visibilitypublic
licenceopenrail
gatedno
ownerammarnasr
Account settings
ammarnasr/the-stack-swift-clean · CoolFace