CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jordangong /the-stack-v2-prtext100M<n<1B0 likes6.8k downloads3d agoHugging Face02Cyrile /dataset-the-stack-v2-dedup-sub The Stack v2 Subset with File Contents (Python, Java, JavaScript, C, C++) TempestTeam/dataset-the-stack-v2-dedup-sub Dataset Summary This dataset is a language-filtered and self-contained subset of bigcode/the-stack-v2-dedup, part of the BigCode Project. It contains only files written in the following programming languages: Python 🐍 Java ☕ JavaScript 📜 C ⚙️ C++ ⚙️ Unlike the original dataset, which only includes metadata and Software Heritage IDs, this subset includes… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/dataset-the-stack-v2-dedup-sub.tabulartext-generation10M<n<100M6 likes3k downloads1y agoHugging Face03tingtang2 /the_stack_v2_python_repos_pretraining_dataset_imported_context-datasettext1M<n<10M0 likes2.2k downloads1y agoHugging Face04bigcode /the-stack-v2-dedupgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated <-- you are here bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-dedup.tabulartext-generation1B<n<10B139 likes2.2k downloads2mo agoHugging Face05Reset23 /the-stack-v2tabular100M<n<1B1 likes1.8k downloads2y agoHugging Face06Reset23 /the-stack-v2-pythontabular1M<n<10M0 likes1.8k downloads2y agoHugging Face07Reset23 /the-stack-v2-ctabular1M<n<10M0 likes1.4k downloads2y agoHugging Face08Reset23 /the-stack-v2-new-pythontabular1M<n<10M0 likes1.2k downloads2y agoHugging Face09Reset23 /the-stack-v2-new-ctabular1M<n<10M0 likes1.1k downloads2y agoHugging Face10Reset23 /the-stack-v2-new-cpptabular1M<n<10M1 likes1.1k downloads1y agoHugging Face11CohenQu /the-stack-v2-dedup-Python_10ktext1M<n<10M0 likes1k downloads1y agoHugging Face12thepowerfuldeez /the-stack-v2-train-smol-ids-updatedUpdate on The Stack V2 dataset: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids All repos from original dataset are parsed with Github API and re-downloaded, so respective updates are kept, metadata is updated. This took 10+ days to process due to GraphQL limits. Filtering rules Removed repos with no update in the last 6 years (no updates since September 2019) Removed files with a single line Removed repos with a single file Removed repos with more than 99%… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated.tabular100K<n<1M0 likes969 downloads1y agoHugging Face13Wholesomeisland /rust-the-stack-v2text1M<n<10M0 likes957 downloads5mo agoHugging Face14smohammadi /the-stack-v2-python-shuffletext1M<n<10M0 likes898 downloads7mo agoHugging Face15bigcode /the-stack-v2-train-smol-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.tabulartext-generation10M<n<100M60 likes857 downloads2mo agoHugging Face16bigcode /the-stack-v2-train-full-idsgated The Stack v2 The dataset consists of 4 versions: bigcode/the-stack-v2: the full "The Stack v2" dataset bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.tabulartext-generation10M<n<100M65 likes805 downloads2mo agoHugging Face17Reset23 /the-stack-v2-cpptabular1M<n<10M1 likes802 downloads2y agoHugging Face18Reset23 /the-stack-v2-processedtabular1M<n<10M0 likes634 downloads1y agoHugging Face19luowenyang /the-stack-v2-20B-sampletext1M<n<10M0 likes562 downloads1y agoHugging Face20tingtang2 /the_stack_v2_2M_repos_pretraining_dataset_imported_context-datasettext100K<n<1M1 likes553 downloads1y agoHugging Face21CohenQu /the-stack-v2-dedup-Python_10k_00text100K<n<1M0 likes495 downloads1y agoHugging Face22tingtang2 /the_stack_v2_2M_repos_filtered_for_python-datasettext1M<n<10M1 likes474 downloads1y agoHugging Face23M1keR /the-stack-v2-dedup-filtered-500-stars-100-forks-contentstabular1M<n<10M1 likes459 downloads1y agoHugging Face24CohenQu /the-stack-v2-dedup-Pythontext100K<n<1M0 likes417 downloads1y agoHugging Face25CohenQu /the-stack-v2-dedup-Python_10k_01text1M<n<10M0 likes385 downloads1y agoHugging Face26thepowerfuldeez /the-stack-v2-train-smol-ids-updated-contentIncrementally uploaded Parquet shards under data/ with columns: repo_name: str text: str Downloaded all repos from https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated and run formatting / linting / import sort on all files Using ruff + black + ty stack for python and biomejs for js / ts / html / css / json / graphql Total amount of tokens: ~100B Example: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated-content.text100K<n<1M0 likes367 downloads1y agoHugging Face27Reset23 /the-stack-v2-blamedtabular1M<n<10M0 likes276 downloads1y agoHugging Face28handwoven8588 /the-stack-v2-train-xsmol-contentgated The Stack v2 — 12-Language Resolved Content This dataset provides resolved file content for twelve programming languages, derived from the repository/file identifiers published in bigcode/the-stack-v2-train-full-ids. The upstream dataset ships identifiers only — each file is a pointer into the Software Heritage archive. Here, those identifiers have been resolved to their actual source text so the content is directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.tabulartext-generation100M<n<1B1 likes263 downloads2mo agoHugging Face29Reset23 /the-stack-v2-filteredtabular1M<n<10M0 likes259 downloads1y agoHugging Face30tingtang2 /the_stack_v2_first_300k_python_repos-datasettext1M<n<10M0 likes240 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.