litble/Multi-Docker-Eval
Dataset Summary Multi-Docker-Eval is a multi-language, multi-dimensional benchmark designed to rigorously evaluate the capability of Large Language Model (LLM)-based agents in automating a critical yet underexplored task: constructing executable Docker environments for real-world software repositories. How to Use from datasets import load_dataset ds = load_dataset('litble/Multi-Docker-Eval') Dataset Structure The data format of Multi-Docker-Eval is… See the full description on the dataset page: https://huggingface.co/datasets/litble/Multi-Docker-Eval.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face