code-comments
code-comments-small
Comment Dataset
Opening comments extracted from code datasets with CommentMiner and ML4SE-toolkit.
Files are grouped as <dataset>/<language>/part-*.parquet.
The Hugging Face dataset card declares one config per source dataset and one split-safe language name per language.
Each row contains dataset, record_id, opening_comment, language, path, repo, extracted_at, and metadata.
For Parquet exports, metadata is stored as a JSON string so every source dataset shares one stable… See the full description on the dataset page: https://huggingface.co/datasets/Jkatzy/code-comments-small.smart_contract_code_commentsmultilingual-code-comments-fixed-8
Fixed-8
Based on fixed-7 revision 14e85fe00a8b284cd226c58281ddd8e6b990b190. Replaces six Greek rows with missing expert labels with six newly labelled samples. All five language configurations retain 500 training rows (2,500 total). All other rows are unchanged.
Removed ID
Replacement ID
8000_5
1056_0
8000_14
4357_8
8000_15
4848_9
8000_16
29069_13
8000_17
1385_4
8000_18
5142_0
All 500 Greek rows now have all five expert accuracy labels. Original… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments-fixed-8.multilingual-code-comments
A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics
This dataset helps us understand how Large Language Models (LLMs) can create code comments in different languages. While LLMs are good at coding tasks in English, we don't know much about how well they work in other languages. This dataset, along with our research, studies how LLMs generate code comments in English, Chinese, Dutch, Polish, and Greek. In our case, we have… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/multilingual-code-comments.multilingual-code-comments-fixed-7arch-code-transfer-lpi-260903T0846-w2-code_no_comments-dataset
