datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mosaic-starcoder-filtered
Mosaic format for filtered starcoder dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-starcoder.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered
load it,
from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered.starcoder-python-instruct
StarCoder-Python-Qwen-Instruct
Dataset Description
This dataset contains Python code samples paired with synthetically generated natural language instructions. It is designed for supervised fine-tuning of language models for code generation tasks. The dataset is derived from the Python subset of the bigcode/starcoderdata corpus, and the instructional text for each code sample was generated using the Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 model.
Creation… See the full description on the dataset page: https://huggingface.co/datasets/OLMo-Coding/starcoder-python-instruct.starcoderdatastarcoderdata-samplebigcode__starcoder2-3b-details
Dataset Card for Evaluation run of bigcode/starcoder2-3b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-3b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-3b-details.bigcode__starcoder2-15b-details
Dataset Card for Evaluation run of bigcode/starcoder2-15b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-15b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-15b-details.bigcode__starcoder2-7b-details
Dataset Card for Evaluation run of bigcode/starcoder2-7b
Dataset automatically created during the evaluation run of model bigcode/starcoder2-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-7b-details.starcoder_solidity_finetunestarcoder_test_contextstarcoderdatasetnewstarcoder_java_refinestarcoder_java_finalstarcoder_3b_baseline_solidityontocord__starcoder2-29b-ls-details
Dataset Card for Evaluation run of ontocord/starcoder2-29b-ls
Dataset automatically created during the evaluation run of model ontocord/starcoder2-29b-ls
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2-29b-ls-details.starcoder-agent-eval
starcoder-agent-eval
Eval records for Colby/starcoder-7b-agent checkpoints (seed=999).
Source
Count
System prompt
Tool calls
Crownelius/Opus-4.6-Reasoning-3300x
20
none
none
Roman1111111/claude-opus-4.6-10000x
20
from record
none
All expected answers are short, auto-verifiable values (numbers, bools, short lists).
No synthetic ANSWER: prompt. Do not add these records to training data.
gemma-starcoder-baseline-soliditystarcoder_java_baselineontocord__starcoder2_3b-AutoRedteam-details
Dataset Card for Evaluation run of ontocord/starcoder2_3b-AutoRedteam
Dataset automatically created during the evaluation run of model ontocord/starcoder2_3b-AutoRedteam
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2_3b-AutoRedteam-details.starcoder10starcoder_java_finetuneinstruction_response_starcoder2-15bbigcode-starcoderdata
bigcode/starcoderdata
This repository documents a dataset used by the Mothership project. By default data is not mirrored here.
Primary source: https://huggingface.co/datasets/bigcode/starcoderdata
Local cache (if present during publishing): C:\Users\Sean Smith\Documents\Scraps\Knowledge\Mothership\library\datasets\bigcode\starcoderdata
Revision pin: none
To reproduce locally, use the project's downloader:
python scripts/download_datasets.py --include bigcode/starcoderdata
