BananaMind/BananaMind-Instruct-Bench-1.1
BananaMind Instruct Bench 1.1 BananaMind Instruct Bench 1.1 is the second version of the open BananaMind Instruct Bench. It is now Elo-based and expands the benchmark to a deterministic 300-example evaluation for small instruction-tuned and chat language models. It measures general task completion, multi-turn behavior, system-prompt following, in-context recall, and Python generation. The dataset is intended to be gated on Hugging Face. Accept the repository access conditions… See the full description on the dataset page: https://huggingface.co/datasets/BananaMind/BananaMind-Instruct-Bench-1.1.
723
No card is published for this repository, or it could not be fetched from Hugging Face right now.
