CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01malaysia-ai /mosaic-starcoder-filtered Mosaic format for filtered starcoder dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-starcoder.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered load it, from streaming import… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-starcoder-filtered.textn<1K0 likes3.9k downloads3y agoHugging Face02OLMo-Coding /starcoder-python-instruct StarCoder-Python-Qwen-Instruct Dataset Description This dataset contains Python code samples paired with synthetically generated natural language instructions. It is designed for supervised fine-tuning of language models for code generation tasks. The dataset is derived from the Python subset of the bigcode/starcoderdata corpus, and the instructional text for each code sample was generated using the Qwen/Qwen3-Coder-30B-A3B-Instruct-FP8 model. Creation… See the full description on the dataset page: https://huggingface.co/datasets/OLMo-Coding/starcoder-python-instruct.text1M<n<10M14 likes1.4k downloads1y agoHugging Face03secmlr /starcoderdatatext10M<n<100M0 likes604 downloads2mo agoHugging Face04malaysia-ai /starcoderdata-sampletabular100K<n<1M0 likes406 downloads3y agoHugging Face05open-llm-leaderboard /bigcode__starcoder2-3b-detailsgated Dataset Card for Evaluation run of bigcode/starcoder2-3b Dataset automatically created during the evaluation run of model bigcode/starcoder2-3b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-3b-details.tabular10K<n<100K0 likes60 downloads2y agoHugging Face06open-llm-leaderboard /bigcode__starcoder2-15b-detailsgated Dataset Card for Evaluation run of bigcode/starcoder2-15b Dataset automatically created during the evaluation run of model bigcode/starcoder2-15b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-15b-details.tabular10K<n<100K0 likes42 downloads2y agoHugging Face07open-llm-leaderboard /bigcode__starcoder2-7b-detailsgated Dataset Card for Evaluation run of bigcode/starcoder2-7b Dataset automatically created during the evaluation run of model bigcode/starcoder2-7b The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/bigcode__starcoder2-7b-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face08zhaospei /starcoder_solidity_finetunetabular1K<n<10K1 likes24 downloads2y agoHugging Face09zhaospei /starcoder_test_contexttabular1K<n<10K0 likes19 downloads2y agoHugging Face10haris001 /starcoderdatasetnewtextquestion-answeringn<1K0 likes9 downloads3y agoHugging Face11zhaospei /starcoder_java_refinetabular10K<n<100K0 likes9 downloads2y agoHugging Face12zhaospei /starcoder_java_finaltabular10K<n<100K0 likes8 downloads2y agoHugging Face13zhaospei /starcoder_3b_baseline_soliditytabular1K<n<10K1 likes7 downloads2y agoHugging Face14open-llm-leaderboard /ontocord__starcoder2-29b-ls-detailsgated Dataset Card for Evaluation run of ontocord/starcoder2-29b-ls Dataset automatically created during the evaluation run of model ontocord/starcoder2-29b-ls The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2-29b-ls-details.tabular10K<n<100K0 likes7 downloads2y agoHugging Face15Colby /starcoder-agent-eval starcoder-agent-eval Eval records for Colby/starcoder-7b-agent checkpoints (seed=999). Source Count System prompt Tool calls Crownelius/Opus-4.6-Reasoning-3300x 20 none none Roman1111111/claude-opus-4.6-10000x 20 from record none All expected answers are short, auto-verifiable values (numbers, bools, short lists). No synthetic ANSWER: prompt. Do not add these records to training data. textn<1K0 likes6 downloads4mo agoHugging Face16zhaospei /gemma-starcoder-baseline-soliditytabular1K<n<10K1 likes5 downloads2y agoHugging Face17zhaospei /starcoder_java_baselinetabular1K<n<10K0 likes5 downloads2y agoHugging Face18open-llm-leaderboard /ontocord__starcoder2_3b-AutoRedteam-detailsgated Dataset Card for Evaluation run of ontocord/starcoder2_3b-AutoRedteam Dataset automatically created during the evaluation run of model ontocord/starcoder2_3b-AutoRedteam The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/ontocord__starcoder2_3b-AutoRedteam-details.tabular10K<n<100K0 likes5 downloads2y agoHugging Face19thanhnew2001 /starcoder10textn<1K0 likes4 downloads3y agoHugging Face20zhaospei /starcoder_java_finetunetabular1K<n<10K0 likes4 downloads2y agoHugging Face21yashwanthreddy7178 /instruction_response_starcoder2-15btext100K<n<1M0 likes3 downloads1y agoHugging Face22Makoveli89 /bigcode-starcoderdata bigcode/starcoderdata This repository documents a dataset used by the Mothership project. By default data is not mirrored here. Primary source: https://huggingface.co/datasets/bigcode/starcoderdata Local cache (if present during publishing): C:\Users\Sean Smith\Documents\Scraps\Knowledge\Mothership\library\datasets\bigcode\starcoderdata Revision pin: none To reproduce locally, use the project's downloader: python scripts/download_datasets.py --include bigcode/starcoderdata textn<1K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.