datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
repro-abc-bench-an-agentic-bio-capabilities-benchmark-for-biosecurity-traces
Agent traces
Agent sessions published from a Trackio Logbook.
llm-benchmarks-capabilities-2020-2026
📊 LLM Benchmarks & Capabilities 2020–2026
The most comprehensive open dataset tracking the evolution of Large Language Models — from GPT-3 to GPT-5.5, Claude Opus 4.7, Gemini 3.5, and beyond.
🧭 Overview
This dataset captures the complete LLM landscape from 2020 to 2026 across five dimensions:
🤖 113 models from 25+ organizations
📈 17 benchmarks tracking capability growth over time
💰 Monthly API pricing showing 100x+ cost reductions
⚙️ Training compute… See the full description on the dataset page: https://huggingface.co/datasets/hmnshudhmn24/llm-benchmarks-capabilities-2020-2026.sandbagging_clarify_2k_capabilities_incentivesandbagging_conceal_6k_capabilities_incentivedefence-capabilitiescrosseval
