CoolFace
Datasetpublic

bigcode/the-stack-inspection-data

Dataset Description A subset of the-stack dataset, from 87 programming languages, and 295 extensions. Each language is in a separate folder under data/ and contains folders of its extensions. We select samples from 20,000 random files of the original dataset, and keep a maximum of 1,000 files per extension. Check this space for inspecting this dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-inspection-data.

sourceHugging Faceupdated 4y agoView on Hugging Face
3likes778downloads
data.json4 linesDownload Raw Back to adp
1version https://git-lfs.github.com/spec/v12oid sha256:a0becfa50f1046fbf3c3c560a2fa2674cef9dc659dcd8514f408705f011e7fcb3size 11257864