litble/Multi-Docker-Eval
Dataset Summary Multi-Docker-Eval is a multi-language, multi-dimensional benchmark designed to rigorously evaluate the capability of Large Language Model (LLM)-based agents in automating a critical yet underexplored task: constructing executable Docker environments for real-world software repositories. How to Use from datasets import load_dataset ds = load_dataset('litble/Multi-Docker-Eval') Dataset Structure The data format of Multi-Docker-Eval is… See the full description on the dataset page: https://huggingface.co/datasets/litble/Multi-Docker-Eval.
This repository belongs to litble on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
