CoolFace
20 results

meta-cognitive

FINAL-Bench /Metacognitive FINAL Bench: Functional Metacognitive Reasoning Benchmark "Not how much AI knows — but whether it knows what it doesn't know, and can fix it." --- Overview FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs). Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.documenttext-generationn<1K106 likes480 downloads7mo agoHugging Facesynthiumjp /metacognitive-profile-atlas Metacognitive Profile Atlas Domain-level metacognitive monitoring quality in 33 frontier LLMs. 47,151 (answer, confidence) observations from 33 frontier LLMs on 1,500 stratified MMLU items across six cognitive domains. Dataset summary The Metacognitive Profile Atlas provides item-level verbalized-confidence data for evaluating how well LLMs monitor their own accuracy, decomposed by cognitive domain. Each observation is one (model, item) pair containing the model's… See the full description on the dataset page: https://huggingface.co/datasets/synthiumjp/metacognitive-profile-atlas.text-classification10K<n<100K0 likes292 downloads2mo agoHugging Facemetacognitive-behavioral-tuning /rollouts-hotpotqatabular100K<n<1M0 likes195 downloads8mo agoHugging FacePixedar /llama_metacognitivetext10K<n<100K0 likes104 downloads2y agoHugging Faceaiqtech /Metacognitive FINAL Bench: Functional Metacognitive Reasoning Benchmark "Not how much AI knows — but whether it knows what it doesn't know, and can fix it." Overview FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs). Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/aiqtech/Metacognitive.texttext-generationn<1K0 likes83 downloads7mo agoHugging Facefantos /Metacognitive FINAL Bench: Functional Metacognitive Reasoning Benchmark "Not how much AI knows — but whether it knows what it doesn't know, and can fix it." Overview FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs). Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/fantos/Metacognitive.texttext-generationn<1K0 likes68 downloads7mo agoHugging Face