ef-ai/wallet-eval-benchmark
Wallet tool-calling eval benchmark The eval side of the wallet fine-tuning work: what the models are scored on. The training rows are deliberately not published. These cases are held out from them by construction, and that is the only reason a score here means anything. If you train on this benchmark, say so — a number from a contaminated run is not comparable to the ones below. The model these cases were used to select is public: ef-dai-team/gemma-4-E4B-wallet-ft-v5, which… See the full description on the dataset page: https://huggingface.co/datasets/ef-ai/wallet-eval-benchmark.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face