CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cyrile /dataset-the-stack-v2-dedup-sub The Stack v2 Subset with File Contents (Python, Java, JavaScript, C, C++) TempestTeam/dataset-the-stack-v2-dedup-sub Dataset Summary This dataset is a language-filtered and self-contained subset of bigcode/the-stack-v2-dedup, part of the BigCode Project. It contains only files written in the following programming languages: Python 🐍 Java ☕ JavaScript 📜 C ⚙️ C++ ⚙️ Unlike the original dataset, which only includes metadata and Software Heritage IDs, this subset includes… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/dataset-the-stack-v2-dedup-sub.tabulartext-generation10M<n<100M6 likes3.2k downloads1y agoHugging Face02bigcode /the-stack-v2-dedupgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated <-- you are here bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-dedup.tabulartext-generation1B<n<10B139 likes2.3k downloads2mo agoHugging Face03Reset23 /the-stack-v2tabular100M<n<1B1 likes1.9k downloads2y agoHugging Face04Reset23 /the-stack-v2-pythontabular1M<n<10M0 likes1.8k downloads2y agoHugging Face05Reset23 /the-stack-v2-ctabular1M<n<10M0 likes1.4k downloads2y agoHugging Face06Reset23 /the-stack-v2-new-pythontabular1M<n<10M0 likes1.2k downloads2y agoHugging Face07Reset23 /the-stack-v2-new-ctabular1M<n<10M0 likes1.1k downloads2y agoHugging Face08Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes1.1k downloads1y agoHugging Face09thepowerfuldeez /the-stack-v2-train-smol-ids-updatedUpdate on The Stack V2 dataset: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids All repos from original dataset are parsed with Github API and re-downloaded, so respective updates are kept, metadata is updated. This took 10+ days to process due to GraphQL limits. Filtering rules Removed repos with no update in the last 6 years (no updates since September 2019) Removed files with a single line Removed repos with a single file Removed repos with more than 99%… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated.tabular100K<n<1M0 likes967 downloads1y agoHugging Face10bigcode /the-stack-v2-train-smol-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.tabulartext-generation10M<n<100M60 likes841 downloads2mo agoHugging Face11bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M65 likes812 downloads2mo agoHugging Face12Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes806 downloads2y agoHugging Face13Reset23 /the-stack-v2-processedtabular1M<n<10M0 likes660 downloads1y agoHugging Face14M1keR /the-stack-v2-dedup-filtered-500-stars-100-forks-contentstabular1M<n<10M1 likes447 downloads1y agoHugging Face15Reset23 /the-stack-v2-blamedtabular1M<n<10M0 likes272 downloads1y agoHugging Face16Reset23 /the-stack-v2-filteredtabular1M<n<10M0 likes261 downloads1y agoHugging Face17handwoven8588 /the-stack-v2-train-xsmol-contentgated The Stack v2 — 12-Language Resolved Content This dataset provides resolved file content for twelve programming languages, derived from the repository/file identifiers published in bigcode/the-stack-v2-train-full-ids. The upstream dataset ships identifiers only — each file is a pointer into the Software Heritage archive. Here, those identifiers have been resolved to their actual source text so the content is directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.tabulartext-generation100M<n<1B1 likes214 downloads2mo agoHugging Face18Reset23 /the-stack-v2-filtered-ctabular100K<n<1M0 likes176 downloads1y agoHugging Face19dboo1334 /the-stack-v2-pythontabular1M<n<10M0 likes172 downloads5mo agoHugging Face20Reset23 /the-stack-v2-filtered-cpptabular100K<n<1M0 likes142 downloads1y agoHugging Face21Reset23 /the-stack-v2-filtered-pythontabular100K<n<1M0 likes134 downloads1y agoHugging Face22Reset23 /the-stack-v2-javatabular1M<n<10M0 likes103 downloads2y agoHugging Face23vinsblack /The_Stack_Processed-v2 🔥 The Stack Processed V2 A curated, balanced, and ML-optimized multi-language programming dataset 🎯 Why Choose This Dataset? A meticulously curated version of "The Stack" optimized for training robust multi-language code models. Perfect balance between quality, diversity, and usability. ✨ Key Advantages: 🎯 Perfect Balance: ~10,000 files per major programming language ⚡ Training-Ready: Parquet format optimized for ML workflows 🏆 Superior Quality: 91.3% syntax… See the full description on the dataset page: https://huggingface.co/datasets/vinsblack/The_Stack_Processed-v2.tabulartext-generation100K<n<1M4 likes76 downloads1y agoHugging Face24Reset23 /the-stack-v2-blamed2tabular100K<n<1M0 likes70 downloads2y agoHugging Face25onekq-ai /the-stack-v2-dedup-sql-annotateThis dataset focuses the SQL corpus of the Stack v2 dedup dataset and provides annotations to each SQL file therein. sqlparse is used to parse the SQL code, then count keywords and symbols. Below are the annotation columns. Column Name Column Description Keyword.DML Data Manipulation Language commands for modifying database records, e.g. SELECT, INSERT, UPDATE, DELETE, COMMIT, MERGE, ROLLBACK. Keyword.DDL Data Definition Language commands for defining or altering database… See the full description on the dataset page: https://huggingface.co/datasets/onekq-ai/the-stack-v2-dedup-sql-annotate.tabulartext-generation1M<n<10M0 likes50 downloads2y agoHugging Face26Reset23 /the-stack-v2-blamed-pythontabular10K<n<100K0 likes48 downloads2y agoHugging Face27Reset23 /the-stack-v2-new-javatabular1M<n<10M0 likes42 downloads2y agoHugging Face28Reset23 /the-stack-v2-blamed-cpptabular10K<n<100K0 likes41 downloads2y agoHugging Face29NanoMatriX /the-stack-v2-dedup285ktabular100K<n<1M0 likes35 downloads8mo agoHugging Face30Reset23 /the-stack-v2-filtered-javatabular100K<n<1M0 likes28 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.