thu-coai/MTAC-IFBench
MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding 🌟 Overview MTAC-IFBench benchmarks instruction following in multi-turn agentic coding. Existing agentic coding benchmarks (e.g., SWE-bench, Terminal-Bench) focus on final functional correctness, while current instruction-following benchmarks confine themselves to single-turn chat or code generation. Neither answers the question that matters in a real development session: does the agent… See the full description on the dataset page: https://huggingface.co/datasets/thu-coai/MTAC-IFBench.
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Add eval configs and dataset card
Upload folder using huggingface_hub
initial commit
