datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SmellBench
SmellBench: Towards Fine-Grained Evaluation of Code Agents on Refactoring Tasks
Dataset Summary
SmellBench is a benchmark designed to evaluate whether code agents can detect and refactor bad code (code smells). Each instance represents a validated code smell injection case constructed from real-world open-source repositories, enabling fine-grained assessment of code agents' refactoring capabilities.
Supported Tasks
Code Refactoring: Given code with… See the full description on the dataset page: https://huggingface.co/datasets/critical88/SmellBench.SmellLearning
