CoolFace
10 results

api-benchmark

ComparEdge /llm-api-benchmark-matrix-2026 LLM Benchmark & Feature Matrix 2026 Which LLM is best at what? This dataset maps capabilities, performance, and limits of 22 major models. Unlike pricing datasets, this focuses on what models can do — not just what they cost. Files File Description llm-benchmarks-2026.csv MMLU, HumanEval, MATH, Arena ELO, coding/reasoning/multilingual rankings, tier (S+ to B) llm-features-2026.csv 15 binary capabilities: vision, function calling, JSON mode, fine-tuning, tool… See the full description on the dataset page: https://huggingface.co/datasets/ComparEdge/llm-api-benchmark-matrix-2026.n<1K1 likes55 downloads5mo agoHugging FaceCohenQu /arxiv_rlad_math_reasoning_benchmark_api_solstextn<1K0 likes45 downloads1y agoHugging Facenativeport /web-access-api-benchmarks NativePort Web-Access API Benchmarks Measured quality, latency, cost and error-rate figures for 22 commercial web-access APIs — search, SERP, scraping, crawling, extraction, sourced answers, screenshots, document parsing, browser actions and change watching — scored per capability on a fixed task corpus. This is the 2026-08-05 run: 67 provider × capability scorecards across 13 capabilities, flattened into 297 metric rows. It exists for one practical decision: when an AI agent… See the full description on the dataset page: https://huggingface.co/datasets/nativeport/web-access-api-benchmarks.tabularn<1K0 likes40 downloads1mo agoHugging FaceCohenQu /arxiv_rlad_math_reasoning_benchmark_api_sol_cond_hint_newtextn<1K0 likes23 downloads1y agoHugging FaceCohenQu /arxiv_rlad_math_reasoning_benchmark_api_sols_mergedtextn<1K0 likes22 downloads1y agoHugging FaceCohenQu /arxiv_rlad_math_reasoning_benchmark_api_hinttextn<1K0 likes21 downloads1y agoHugging Face