datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-v2-prdataset-the-stack-v2-dedup-sub
The Stack v2 Subset with File Contents (Python, Java, JavaScript, C, C++)
TempestTeam/dataset-the-stack-v2-dedup-sub
Dataset Summary
This dataset is a language-filtered and self-contained subset of bigcode/the-stack-v2-dedup, part
of the BigCode Project.
It contains only files written in the following programming languages:
Python 🐍
Java ☕
JavaScript 📜
C ⚙️
C++ ⚙️
Unlike the original dataset, which only includes metadata and Software Heritage IDs, this subset includes… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/dataset-the-stack-v2-dedup-sub.the-stack-v2-dedup
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated <-- you are here
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories.
bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-dedup.the_stack_v2_python_repos_pretraining_dataset_imported_context-datasetthe-stack-v2the-stack-v2-pythonthe-stack-v2-python-shufflethe-stack-v2-cthe-stack-v2-new-pythonthe-stack-v2-new-cthe-stack-v2-new-cppthe-stack-v2-dedup-Python_10krust-the-stack-v2the-stack-v2-train-smol-ids-updatedUpdate on The Stack V2 dataset: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids
All repos from original dataset are parsed with Github API and re-downloaded,
so respective updates are kept, metadata is updated. This took 10+ days to process due to GraphQL limits.
Filtering rules
Removed repos with no update in the last 6 years (no updates since September 2019)
Removed files with a single line
Removed repos with a single file
Removed repos with more than 99%… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated.the-stack-v2-train-smol-ids
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories.
bigcode/the-stack-v2-train-smol-ids: based on the bigcode/the-stack-v2-dedup… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids.the-stack-v2-train-full-ids
The Stack v2
The dataset consists of 4 versions:
bigcode/the-stack-v2: the full "The Stack v2" dataset
bigcode/the-stack-v2-dedup: based on the bigcode/the-stack-v2 but further near-deduplicated
bigcode/the-stack-v2-train-full-ids: based on the bigcode/the-stack-v2-dedup dataset but further filtered with heuristics and spanning 600+ programming languages. The data is grouped into repositories. <-- you are here
bigcode/the-stack-v2-train-smol-ids: based on the… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-v2-train-full-ids.the-stack-v2-cppthe-stack-v2-dedup-Python_10k_01the-stack-v2-processedthe_stack_v2_2M_repos_filtered_for_python-datasetthe_stack_v2_2M_repos_pretraining_dataset_imported_context-datasetthe-stack-v2-20B-samplethe-stack-v2-dedup-Python_10k_00the-stack-v2-dedup-filtered-500-stars-100-forks-contentsthe-stack-v2-dedup-Pythonthe-stack-v2-train-smol-ids-updated-contentIncrementally uploaded Parquet shards under data/ with columns:
repo_name: str
text: str
Downloaded all repos from https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated and run
formatting / linting / import sort on all files
Using ruff + black + ty stack for python and biomejs for js / ts / html / css / json / graphql
Total amount of tokens: ~100B
Example:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/thepowerfuldeez/the-stack-v2-train-smol-ids-updated-content.the-stack-v2-blamedthe-stack-v2-filteredthe_stack_v2_first_300k_python_repos-datasetthe-stack-v2-train-xsmol-content
The Stack v2 — 12-Language Resolved Content
This dataset provides resolved file content for twelve programming languages,
derived from the repository/file identifiers published in
bigcode/the-stack-v2-train-full-ids.
The upstream dataset ships identifiers only — each file is a pointer into the
Software Heritage archive. Here, those
identifiers have been resolved to their actual source text so the content is
directly usable, with the upstream metadata carried through unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/handwoven8588/the-stack-v2-train-xsmol-content.
