meta-cognitive
Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.metacognitive-profile-atlas
Metacognitive Profile Atlas
Domain-level metacognitive monitoring quality in 33 frontier LLMs.
47,151 (answer, confidence) observations from 33 frontier LLMs on 1,500 stratified MMLU items across six cognitive domains.
Dataset summary
The Metacognitive Profile Atlas provides item-level verbalized-confidence data for evaluating how well LLMs monitor their own accuracy, decomposed by cognitive domain. Each observation is one (model, item) pair containing the model's… See the full description on the dataset page: https://huggingface.co/datasets/synthiumjp/metacognitive-profile-atlas.rollouts-hotpotqallama_metacognitiveMetacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/aiqtech/Metacognitive.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/fantos/Metacognitive.
