anon-iclr-submission/benchname-module-summarization
🥷 BenchName (Module summarization) This is the benchmark for Module summarization task as part of the 🥷 BenchName benchmark. The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects. The model is required to generate such description, given the relevant context code and the intent behind the documentation. All the repositories are published under permissive licenses (MIT, Apache-2.0… See the full description on the dataset page: https://huggingface.co/datasets/anon-iclr-submission/benchname-module-summarization.
🥷 BenchName (Module summarization)
This is the benchmark for Module summarization task as part of the 🥷 BenchName benchmark. The current version includes 216 manually curated text files describing different documentation of open-source permissive Python projects. The model is required to generate such description, given the relevant context code and the intent behind the documentation. All the repositories are published under permissive licenses (MIT, Apache-2.0, BSD-3-Clause, and BSD-2-Clause). The datapoints can be removed upon request.
How-to
Load the data via `load_dataset`:
from datasets import load_dataset
dataset = load_dataset("anon-iclr-submission/benchname-module-summarization")Datapoint Structure
Each example has the following fields:
Note: you may collect and use your own relevant context. Our context may not be suitable. Zipped repositories can be found the repos directory.
Metric
To compare the predicted documentation and the ground truth documentation, we introduce the new metric based on LLM as an assessor. Our approach involves feeding the LLM with relevant code and two versions of documentation: the ground truth and the model-generated text. The LLM evaluates which documentation better explains and fits the code. To mitigate variance and potential ordering effects in model responses, we calculate the probability that the generated documentation is superior by averaging the results of two queries with the different order.
