CoolFace
Datasetpublic

HanyueShen/YunXiaoHe-LongBench-v2-Comparison

YunXiaoHe on LongBench-v2 YunXiaoHe answered 337 of 503 LongBench-v2 questions correctly, or 67.0%. Four unfinished questions count as incorrect. This is a provisional, self-reported result rather than an official leaderboard submission. The left panel puts that result beside the LongBench-v2 site's first nine model rows, ranked by its overall chain-of-thought score, and all six LongBench-v2 configurations in the 2026 Prime Agent study. The official-site rows range from 56.0%… See the full description on the dataset page: https://huggingface.co/datasets/HanyueShen/YunXiaoHe-LongBench-v2-Comparison.

sourceHugging Faceupdated 1d agoView on Hugging Face
0likes
Dataset Card

YunXiaoHe on LongBench-v2

[image]

YunXiaoHe answered 337 of 503 LongBench-v2 questions correctly, or 67.0%. Four unfinished questions count as incorrect. This is a provisional, self-reported result rather than an official leaderboard submission.

The left panel puts that result beside the LongBench-v2 site's first nine model rows, ranked by its overall chain-of-thought score, and all six LongBench-v2 configurations in the 2026 Prime Agent study. The official-site rows range from 56.0% to 63.3%; the newer agentic study ranges from 68.0% to 74.6%. Each group used its own prompting and harness setup.

The right panel plots accuracy against the current public API price for one million uncached input tokens. Its horizontal scale runs from $0 to $6 in $1 steps. The YunXiaoHe point uses DeepSeek V4 Pro's $0.66 off-peak rate; the listed peak rate is $1.32. The open circles mark three lower-scoring configurations at the same model's input price: GLM-5.2 with Prime Agent, GPT-5.6 Sol with Codex, and Opus 5 with Prime Agent. The source study reports point estimates without confidence intervals.

Input price is one part of cost. A complete run also uses output tokens and may involve many model calls, caching, or long-prompt rates. The step line shows the best observed accuracy at each listed input price, not a measured run-cost frontier.

Files

  • —`figures/accuracy_vs_input_price.png`: display figure.
  • —`figures/accuracy_vs_input_price.svg`: vector figure.
  • —`data/points.csv`: the 16 plotted accuracy points and the seven input prices used in the right panel. A blank price means that the system appears only in the accuracy panel; it does not mean zero cost.
  • —`scripts/plot_comparison.py`: redraws both figures from the CSV with Matplotlib 3.11 or newer. Run python scripts/plot_comparison.py from a clone of this repository.

Sources

Accuracy: LongBench-v2 leaderboard, Prime Agent, Table 1. API input prices, checked 2026-09-25 UTC: DeepSeek, Z.AI, OpenAI, Anthropic.

Chart and comparison table: © 2026 YH Intelligence Technology, Co., Ltd. LongBench-v2 and the cited model results remain attributable to their original authors.