CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JetBrains-Research /lca-bug-localization 🏟️ Long Code Arena (Bug localization) This is the benchmark for the Bug localization task as part of the 🏟️ Long Code Arena benchmark. The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug. The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.imagetext-generation10K<n<100K4 likes1.2k downloads2y agoHugging Face02JetBrains-Research /lca-module-summarization 🏟️ Long Code Arena (Module summarization) This is the benchmark for Module summarization task as part of the 🏟️ Long Code Arena benchmark. The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects. The model is required to generate such description, given the relevant context code and the intent behind the documentation. All the repositories are published under permissive licenses (MIT, Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-module-summarization.imagetext-generationn<1K1 likes104 downloads2y agoHugging Face03DanCip /lca-StartingPoints-expanded🧠 LCA-Starting Points A benchmark for evaluating project-local code completion ranking. Curated to validate TreeRanker (ASE2025). 📖 Dataset Description Starting Points is a specialized dataset designed to evaluate code completion ranking, with a specific focus on locally defined identifiers (project-specific APIs) in Python. Most LLM benchmarks focus on global APIs (standard libraries). However, developers spend significant time using APIs defined within their own… See the full description on the dataset page: https://huggingface.co/datasets/DanCip/lca-StartingPoints-expanded.tabulartext-generation1K<n<10K0 likes20 downloads9mo agoHugging Face04alucent /mirror-lca-bug-localizationgated 🏟️ Long Code Arena (Bug localization) This is the benchmark for the Bug localization task as part of the 🏟️ Long Code Arena benchmark. The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug. The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-lca-bug-localization.tabulartext-generation10K<n<100K0 likes4 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.