Shiftedx/shiftedx-bench
Shiftedx Bench Shiftedx Bench is a reproducible qualification suite for local language-model deployments. It measures model quality, effective long-context use, tool protocol reliability, multi-turn agent behavior, vision, and runtime performance without collapsing them into a single “intelligence” score. The project is designed for quantization and speculative-decoding decisions on Apple Silicon, but its API runner works with any OpenAI-compatible chat endpoint.… See the full description on the dataset page: https://huggingface.co/datasets/Shiftedx/shiftedx-bench.
0183
