Shiftedx/shiftedx-bench
Shiftedx Bench Shiftedx Bench is a reproducible qualification suite for local language-model deployments. It measures model quality, effective long-context use, tool protocol reliability, multi-turn agent behavior, vision, and runtime performance without collapsing them into a single “intelligence” score. The project is designed for quantization and speculative-decoding decisions on Apple Silicon, but its API runner works with any OpenAI-compatible chat endpoint.… See the full description on the dataset page: https://huggingface.co/datasets/Shiftedx/shiftedx-bench.
Publish AEON baseline vs Shiftedx Agent Harness results
Fix immutable provenance flags for Shiftedx Bench v0.5.1
Release Shiftedx Bench v0.5.0 with Shiftedx Agent Harness
fix: render unestablished effective context
Record immutable benchmark revision in run manifests
v0.3.0: add post-publish model-card qualification
v0.2.1: make 128K the default quant-gate ceiling
Add v0.2.0 lightweight full-window quant gate
Publish Shiftedx Bench v0.1.0
initial commit
