datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commitbench
CommitBench: A Benchmark for Commit Message Generation
EXECUTIVE SUMMARY
We provide CommitBench as an open-source, reproducible and privacy- and license-aware benchmark for commit message generation. The dataset is gathered from GitHub repositories with licenses that permit redistribution. We provide six programming languages, Java, Python, Go, JavaScript, PHP, and Ruby. The commit messages in natural language are restricted to English, as it is the working language in… See the full description on the dataset page: https://huggingface.co/datasets/Maxscha/commitbench.commit-classification-dataset
Commit Classification Dataset
This dataset is designed for multi-label classification of Git commit messages into predefined categories.
Dataset Summary
This dataset contains:
Training data: Commit messages and their corresponding labels for training the model.
Validation data: A separate set of messages for tuning and evaluation.
Testing data: Unlabeled commit messages for testing the model’s performance.
The goal of the dataset is to classify each commit message into… See the full description on the dataset page: https://huggingface.co/datasets/meriemm6/commit-classification-dataset.commit-classification-17kserena-synthetic-it-28h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.commitbench_longCommitBench: A Benchmark for Commit Message Generation
We provide CommitBench as an open-source, reproducible and privacy- and license-aware benchmark for commit message generation. The dataset is gathered from GitHub repositories with licenses that permit redistribution. We provide six programming languages, Java, Python, Go, JavaScript, PHP, and Ruby. The commit messages in natural language are restricted to English, as it is the working language in many software development projects. The… See the full description on the dataset page: https://huggingface.co/datasets/Maxscha/commitbench_long.tangled-ccs-commits
Detecting Multiple Semantic Concerns in Tangled Code Commits using Small Language Models
This dataset contains commit data for training and evaluating models on software engineering tasks, specifically focusing on identifying and separating concerns in multi-concern commits.
Every tangled (multi-concern) commit in this dataset is composed exclusively of atomic commits from a single repository — resolving a structural weakness in earlier cross-repo tangles (which were trivially… See the full description on the dataset page: https://huggingface.co/datasets/Berom0227/tangled-ccs-commits.commit_messagesccs_commits_with_bodygit-diff_to_commit_msg
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
There are 2 version of this dataset:
git-diff_to_commit_msg - 1.5K rows
huggingface link
kaggle link
git-diff_to_commit_msg_large - 1.75M rows
huggingface link
kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg.serena-synthetic-it-27h
Qwen3-TTS Italian Synthetic Speech (27h)
Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with
Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV,
Piper-ready metadata.
Dataset summary
Property
Value
Clips (train / val)
26,523 / 2,947
Total duration
~27.3 h (98,099 s)
Sample rate
22,050 Hz mono, 16-bit WAV
Loudness
Normalized to -23 LUFS, silence-trimmed
Language
Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-27h.cockroach-commitsgit-diff_to_commit_msg_large
Hi, I’m Seniru Epasinghe 👋
I’m an AI undergraduate and an AI enthusiast, working on machine learning projects and open-source contributions.I enjoy exploring AI pipelines, natural language processing, and building tools that make development easier.
🌐 Connect with me
There are 2 version of this dataset:
git-diff_to_commit_msg - 1.5K rows
huggingface link
kaggle link
git-diff_to_commit_msg_large - 1.75M rows
huggingface link
kaggle link… See the full description on the dataset page: https://huggingface.co/datasets/seniruk/git-diff_to_commit_msg_large.llm-commit-message-evaluation
Dataset Card for LLM Commit Message Evaluation
The LLM Commit Message Evaluation dataset is designed to evaluate and compare the performance of Large Language Models (LLMs) in generating high-quality git commit messages. It contains real-world code diffs, issue descriptions, and issue titles extracted from open-source repositories (such as OWASP/Nest).
For each code change, the dataset provides the original human-written commit message alongside commit messages generated by… See the full description on the dataset page: https://huggingface.co/datasets/g-for-gour/llm-commit-message-evaluation.loki-commitsgit_commitsleetcode-commitsmemos-commitsangular-commitsvuetify-commitsgreptimedb-commitsuni-app-commitsant-design-commitshyperswitch-commitsobs-studio-commitsCommitPack_CLtaro-commitssentry-commitsappsmith-commitsCommitPack_long_CLangular-cli-commits
