davanstrien/benchmark-diversity-atlas
0
Benchmark Diversity Atlas
Interactive embedding visualization comparing task diversity across three coding/agent benchmarks:
- SWE-bench Verified (500 tasks) - Python open-source bug fixing
- SWE-bench Pro (731 tasks) - Enterprise multi-language software engineering
- Terminal-bench 2.0 (89 tasks) - Diverse terminal agent tasks
Built with Embedding Atlas. MCP endpoint available at /mcp for tool integration.
