CoolFace
Datasetpublic

Shiftedx/shiftedx-bench

Shiftedx Bench Shiftedx Bench is a reproducible qualification suite for local language-model deployments. It measures model quality, effective long-context use, tool protocol reliability, multi-turn agent behavior, vision, and runtime performance without collapsing them into a single “intelligence” score. The project is designed for quantization and speculative-decoding decisions on Apple Silicon, but its API runner works with any OpenAI-compatible chat endpoint.… See the full description on the dataset page: https://huggingface.co/datasets/Shiftedx/shiftedx-bench.

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes183downloads
10 commits on main
9a430271mo ago

Publish AEON baseline vs Shiftedx Agent Harness results

Shiftedx
335e6691mo ago

Fix immutable provenance flags for Shiftedx Bench v0.5.1

Shiftedx
a6422421mo ago

Release Shiftedx Bench v0.5.0 with Shiftedx Agent Harness

Shiftedx
f60e0cb1mo ago

fix: render unestablished effective context

Shiftedx
3bbb0bf1mo ago

Record immutable benchmark revision in run manifests

Shiftedx
cc249f31mo ago

v0.3.0: add post-publish model-card qualification

Shiftedx
48e63681mo ago

v0.2.1: make 128K the default quant-gate ceiling

Shiftedx
9ddbe3c1mo ago

Add v0.2.0 lightweight full-window quant gate

Shiftedx
f039f411mo ago

Publish Shiftedx Bench v0.1.0

Shiftedx
9f6d0741mo ago

initial commit

Shiftedx