CoolFace
Datasetpublicgated

bigcode/the-stack-smol

Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the original dataset. The dataset has 2.6GB of text (code). Languages The dataset contains 30 programming languages: "assembly", "batchfile", "c++", "c", "c-sharp", "cmake", "css", "dockerfile", "fortran", "go", "haskell", "html", "java", "javascript", "julia", "lua", "makefile", "markdown", "perl", "php", "powershell", "python", "ruby"… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol.

sourceHugging Faceupdated 3y agoView on Hugging Face
93likes27kdownloads

bigcode/the-stack-smol · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.