datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
starcoderdata
StarCoder Training Dataset
Dataset description
This is the dataset used for training StarCoder and StarCoderBase. It contains 783GB of code in 86 programming languages, and includes 54GB GitHub Issues + 13GB Jupyter notebooks in scripts and text-code pairs,
and 32GB of GitHub commits, which is approximately 250 Billion tokens.
Dataset creation
The creation and filtering of The Stack is explained in the original dataset, we additionally decontaminate and… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/starcoderdata.starcoder-python5b5b gpt2 tokens
starcoder-curated
StarCoderData Curated
A curated subset of StarCoderData
optimised for training a 500M parameter model focused on structured data output
(JSON generation, function calling, schema compliance).
Dataset Summary
Total code files: 5,203,508
Total tokens: 3.9B (target: 3.5B)
Classifier-scored files: 1,553,596 (1.7B tokens)
Non-classified files: 3,649,912 (2.2B tokens) — filtered by heuristics, not the classifier
Source: bigcode/starcoderdata
Classifier:… See the full description on the dataset page: https://huggingface.co/datasets/mdonigian/starcoder-curated.local-code-arena-mbpp-starcoder_1b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder 1B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the legacy StarCoder 1B base foundational model.
This specific partition documents the absolute upper-bound inference speed achievable on our local consumer hardware while highlighting the functional limitations of un-aligned raw base… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-starcoder_1b.local-code-arena-starcoder_15b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder 15B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the heavyweight StarCoder 15B base foundational model.
This specific partition documents the absolute scaling limits of unaligned foundational weights inside conversational benchmarking loops, establishing a massive baseline for… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder_15b.local-code-arena-mbpp-starcoder_7b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder 7B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the legacy StarCoder 7B base foundational model.
This specific partition documents the behavioral dynamics of larger-scale raw foundational weights inside automated benchmarking pipelines, establishing an anchor point to analyze… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-starcoder_7b.local-code-arena-mbpp-starcoder_3b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder 3B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the legacy StarCoder 3B base foundational model.
This specific partition documents the behavioral dynamics of mid-tier raw foundational weights inside automated benchmarking environments, serving as a clean baseline for… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-mbpp-starcoder_3b.local-code-arena-starcoder2_7b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder2 7B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the next-generation StarCoder2 7B base foundational model.
This specific partition documents the behavioral dynamics of modern, mid-tier raw foundational weights inside automated conversational evaluation workflows, defining the… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder2_7b.local-code-arena-starcoder2_3b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder2 3B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the next-generation StarCoder2 3B base foundational model.
This specific partition documents the behavioral dynamics of modern raw foundational weights inside automated conversational pipelines, highlighting the persistent… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder2_3b.local-code-arena-starcoder2_15b
Local Code Arena Telemetry: MBPP Benchmark on StarCoder2 15B (Base)
This repository hosts the raw evaluation metrics, execution telemetry logs, and structural syntax outputs captured from running the Mostly Basic Python Problems (MBPP) benchmark against the flagship StarCoder2 15B base foundational model.
This specific partition documents the final limits of scaling raw, unaligned foundational weights inside conversational evaluation loops, establishing an absolute baseline for… See the full description on the dataset page: https://huggingface.co/datasets/ShahzebKhoso/local-code-arena-starcoder2_15b.
