small-code
testing_codealpaca_small
Dataset Card for "testing_codealpaca_small"
More Information needed
code-comments-small
Comment Dataset
Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit.
Files are grouped as <dataset>/<language>/part-*.parquet.
The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language.
Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata.
For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.hardware_code_and_sec_smallHigh-Coder-SFT-Small
High-Coder-SFT-Small
A high-quality synthetic code dataset containing 54,950 long-form code samples across 8 programming languages. Generated using Hunter Alpha (1T+ parameter frontier model). Every single sample contains at least 200 lines of actual code — most contain 500+.
This is not a snippet dataset. Every file is a complete, production-quality source file with imports, error handling, design patterns, and modern language idioms. The average sample is 630 lines of code… See the full description on the dataset page: https://huggingface.co/datasets/Crownelius/High-Coder-SFT-Small.agent-trace-privacy-scrubber-codex-tracessmall_repos_multi_file_chatgpt_5_qas_part5_code_qa-dataset
