api-benchmark
llm-api-benchmark-matrix-2026
LLM Benchmark & Feature Matrix 2026
Which LLM is best at what? This dataset maps capabilities, performance, and limits of 22 major models.
Unlike pricing datasets, this focuses on what models can do — not just what they cost.
Files
File
Description
llm-benchmarks-2026.csv
MMLU, HumanEval, MATH, Arena ELO, coding/reasoning/multilingual rankings, tier (S+ to B)
llm-features-2026.csv
15 binary capabilities: vision, function calling, JSON mode, fine-tuning, tool… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/llm-api-benchmark-matrix-2026.arxiv_rlad_math_reasoning_benchmark_api_solsweb-access-api-benchmarks
NativePort Web-Access API Benchmarks
Measured quality, latency, cost and error-rate figures for 22 commercial web-access
APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots,
document parsing, browser actions and change watching — scored per capability on a
fixed task corpus. This is the 2026-08-05 run: 67 provider × capability
scorecards across 13 capabilities, flattened into 297 metric rows.
It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.arxiv_rlad_math_reasoning_benchmark_api_sol_cond_hint_newarxiv_rlad_math_reasoning_benchmark_api_sols_mergedarxiv_rlad_math_reasoning_benchmark_api_hint
