CoolFace
Datasetpublic

bigcode/the-stack-inspection-data

Dataset Description A subset of the-stack dataset, from 87 programming languages, and 295 extensions. Each language is in a separate folder under data/ and contains folders of its extensions. We select samples from 20,000 random files of the original dataset, and keep a maximum of 1,000 files per extension. Check this space for inspecting this dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-inspection-data.

sourceHugging Faceupdated 4y agoView on Hugging Face
3likes778downloads
8 commits on main
145c9f04y ago

Create dataset_creation.py

loubnabnl
174ff164y ago

Create README.md

loubnabnl
ef10eb54y ago

add data

loubnabnl
a4f91d14y ago

add new data

loubnabnl
b63c5b24y ago

add data

loubnabnl
a53a8354y ago

add data

loubnabnl
177b8b64y ago

gitattr

loubnabnl
62088bf4y ago

initial commit

loubnabnl