jeorgexyz/lua-agent-evals
Lua Agent Evals The evidence behind lua-agent-lab, from lua-agent. The experiment separates structural tool-call validity from useful tool selection and complete task success. Contents Configuration Unit Method decoder_trials 60 first-turn trials TinyStories 15M; 30 constrained, 30 free; greedy sampling loop_ablations 64 agent runs Eight scripted tasks × eight loop configurations qwen_runs 16 agent runs Eight tasks × constrained/free Qwen 2.5 1.5B… See the full description on the dataset page: https://huggingface.co/datasets/jeorgexyz/lua-agent-evals.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face