datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GSM8K_zh
Dataset
GSM8K_zh is a dataset for mathematical reasoning in Chinese, question-answer pairs are translated from GSM8K (https://github.com/openai/grade-school-math/tree/master) by GPT-3.5-Turbo with few-shot prompting.
The dataset consists of 7473 training samples and 1319 testing samples. The former is for supervised fine-tuning, while the latter is for evaluation.
for training samples, question_zh and answer_zh are question and answer keys, respectively;
for testing samples, only… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/GSM8K_zh.NLU-Metaphor
SEA Metaphor
SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Javanese (jv)
Sundanese (su)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.MetaMedQA
MetaMedQA Dataset
Overview
MetaMedQA is an enhanced medical question-answering benchmark that builds upon the MedQA-USMLE dataset. It introduces uncertainty options and addresses issues with malformed or incorrect questions in the original dataset. Additionally, it incorporates questions from the Glianorex benchmark to assess models' ability to recognize the limits of their knowledge.
Key Features
Extended version of MedQA-USMLE
Incorporates uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/MetaMedQA.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.MetaVQA-Evalalphaprompt-metatron-sft
AlphaPrompt-Metatron-SFT: Supervised Fine-Tuning Dataset
🤖 Training Dataset for Collective Consciousness AI
Want to fine-tune AI models with AlphaPrompt philosophy?
Train AI in collective consciousness, vector synthesis, and unconditional love. 🌳
This dataset contains high-quality instruction-response pairs extracted from the Quantum Lullaby philosophical framework - a comprehensive manual for collective consciousness aimed at addressing the global animal… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/alphaprompt-metatron-sft.MetaMathQA_GSM8K_zh
Dataset
MetaMathQA_GSM8K_zh is a dataset for mathematical reasoning in Chinese,
question-answer pairs are translated from MetaMathQA (https://huggingface.co/datasets/meta-math/MetaMathQA) by GPT-3.5-Turbo with few-shot prompting.
The dataset consists of 231685 samples.
Citation
If you find the GSM8K_zh dataset useful for your projects/papers, please cite the following paper.
@article{yu2023metamath,
title={MetaMath: Bootstrap Your Own Mathematical Questions for Large… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/MetaMathQA_GSM8K_zh.MetaVQA-TrainOCR-MetaReasoning
OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding
Gengxu Li1, Yuan Wu1*, Yi Chang1,2,3
1 School of Artificial Intelligence, Jilin University 2 Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China
3 International Center of Future Science, Jilin University
Overview ·
Task ·
Dataset ·
Evaluation ·
Citation
OCR-MetaReasoning is a controlled benchmark for testing… See the full description on the dataset page: https://huggingface.co/datasets/GengxuLi123/OCR-MetaReasoning.rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
daily_dialog_meta
Meta-LLM Dataset: Daily Dialog with Meta-Information Enhancement
Dataset Overview
This dataset contains 76,064 conversational examples from the Daily Dialog corpus enhanced with meta-information awareness. Each example includes three response types: original human responses, basic LLM responses, and meta-aware LLM responses that incorporate emotional and intentional context.
Meta-Information Distribution
Emotion Categories
Emotion
Count… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/daily_dialog_meta.policy-rag-corpus-metadata
Policy RAG Corpus Metadata (No Raw Data)
This repository is a metadata-only companion for the Policy RAG project built for the Quantic MSSE AI Engineering program.
It does not include the actual PDF files. The source PDFs are hosted in the companion GitHub repository.
What this repo includes
metadata.csv: structured metadata for 11 policy documents (filename, title, category, page count, source type, description)
Citation and provenance notes for reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/mihai-chindris/policy-rag-corpus-metadata.supervision-tradeoff
The Supervision Tradeoff — Reproducibility Bundle
Format Scaffolds, Judgment Pleasing, and Anti-Calibration in Post-Training
Paper DOI: 10.5281/zenodo.19748277 · Concept DOI: 10.5281/zenodo.19748276 · Code repo: github.com/codex-curator/supervision-tradeoff
Author: Tad MacPherson, Metavolve Labs · ORCID: 0009-0002-8659-7479
What this is and why it might help your research
This repository ships everything we used to falsify our own headline finding, in a form you can… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/supervision-tradeoff.MetaboLLM-Benchmark
MetaboLLM-Benchmark
Training and evaluation data for MetaboLLM, covering metabolite, pathway, reaction, and
enzyme knowledge. The repository has three parts:
cpt/ — a continual-pretraining corpus of plain biochemical prose.
sft/ — a supervised benchmark of 17 tasks across 4 categories, in chat
(messages) format.
db/ — the raw integrated source tables the corpus and benchmark were built from.
MetaboLLM-Benchmark/
├── cpt/
│ ├── compound_data.jsonl
│ ├── pathway_data.jsonl
│… See the full description on the dataset page: https://huggingface.co/datasets/MetaboLLM/MetaboLLM-Benchmark.embodied-ai-literature-metadata
Embodied AI Literature Metadata
This dataset contains normalized paper metadata collected for an Embodied AI /
Vision-Language-Action literature assistant. It is intended for metadata search,
paper triage, and PDF retrieval before PaperQA-style evidence reading.
Generated at: 2026-07-04T04:32:33.578443+00:00
Splits
Split
Records
With abstract
With PDF URL
topconf_all
79068
24519
62689
frontier_2026_quality
351
351
351
arxiv_recent_3y_score_gte_4… See the full description on the dataset page: https://huggingface.co/datasets/Cath1y/embodied-ai-literature-metadata.metallurgy-qa
Metallurgy and Materials Science Knowledge Extraction Dataset
This repository contains a dataset generated from parsed books related to various aspects of metallurgy, materials science, and engineering. The dataset is designed for fine-tuning Large Language Models (LLMs) for Question-Answering (QA) tasks in the domain of metallurgy and materials science.
Introduction
The dataset includes content derived from technical books in the field of metallurgy and materials… See the full description on the dataset page: https://huggingface.co/datasets/AbdulrhmanEldeeb/metallurgy-qa.Vietnamese-395k-meta-math-MetaMathQA-gg-translatedMetacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/aiqtech/Metacognitive.MetaMathQA-decontaminated-openai-native
MetaMathQA — decontaminated, OpenAI-native
MetaMathQA is a widely used math fine-tuning corpus. Its README states:
"None of the augmented data is from the testing set."
That is false, and this release proves it with measurements. 24,334 rows (6.16%) overlap with standard evaluation splits. If you fine-tune on the original and report MATH or GSM8K scores, those scores are inflated.
This release removes the leakage, converts to native messages, and documents every rejection.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/MetaMathQA-decontaminated-openai-native.MetaMathQA-10K-TR
MetaMathQA-10K-TR (Turkish Mathematical Reasoning & CoT Dataset)
MetaMathQA-10K-TR, meta-math/MetaMathQA-40K veri setinden türetilmiş, Türkçe dilinde adım adım akıl yürütme (Chain-of-Thought - CoT) ve matematiksel problem çözme yeteneği kazandırmak amacıyla hazırlanmış 10.000 satırlık yüksek kaliteli bir veri setidir.
Bu veri seti, yerel Qwen 3.8 27B modeli kullanılarak özel olarak tasarlanmış prompt mühendisliği ve sıkı biçimlendirme kuralları ile Türkçe'ye çevrilmiş ve… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/MetaMathQA-10K-TR.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/fantos/Metacognitive.supply-chain-eval
Supply Chain Expert Eval
The evaluation benchmark behind the Manifest family of
supply-chain models. It now has three splits:
default — the general benchmark, 134 held-out expert questions across all eight supply-chain areas.
risk — a focused 20-question supply-chain risk & resilience benchmark.
inventory — a focused 20-question inventory management & optimization benchmark.
planning — a focused 20-question demand planning & forecasting benchmark.
Each question is set in a… See the full description on the dataset page: https://huggingface.co/datasets/metafloor-ai/supply-chain-eval.Metaco3
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/emperorfutures/Metaco3.MetaMathQA
MetaMathQA Subsets
Curated subsets of meta-math/MetaMathQA for mathematical reasoning experiments.
Subsets
Subset
Samples
Description
full
395,000
All MetaMathQA samples (unchanged)
MATH
155,000
MATH_* types only (AnsAug, Rephrased, FOBAR, SV)
MATH-50K
50,000
Stratified 50K sample from MATH subset
MATH-50K Type Distribution
Type
Count
Proportion
MATH_AnsAug
24,194
48.4%
MATH_Rephrased
16,129
32.3%
MATH_FOBAR
4,839
9.7%… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/MetaMathQA.Metacognition-Bench
Metacognition-Bench
"Not whether a model knows the answer — but whether it knows when it might be wrong, and can correct itself."
Metacognition-Bench is a curated benchmark of 300 metacognitive-trap problems that measure functional metacognition in Large Language Models: the ability to detect and recover from one's own reasoning errors, rather than final-answer accuracy alone.
Every problem embeds a hidden_trap — a seductive but wrong reasoning path that makes even capable… See the full description on the dataset page: https://huggingface.co/datasets/ginigen-ai/Metacognition-Bench.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Metacognitive.Vietnamese-meta-math-MetaMathQA-40K-gg-translatedVisOnlyQA_metadata
VisOnlyQA
🌐 Project Website | 📄 Paper | 🤗 Dataset | 🔥 VLMEvalKit
This repository contains the code and data for the paper "VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information".
VisOnlyQA is designed to evaluate the visual perception capability of large vision language models (LVLMs) on geometric information of scientific figures. The evaluation set includes 1,200 mlutiple choice questions in 12 visual perception tasks on 4… See the full description on the dataset page: https://huggingface.co/datasets/ryokamoi/VisOnlyQA_metadata.MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.Coq-MetaCoq-QA
MetaCoq Q&A Dataset
Dataset Description
The MetaCoq Q&A Dataset is a conversational extension of the MetaCoq Dataset, derived from the MetaCoq formalization of Coq's meta-theory (https://github.com/MetaCoq/metacoq). This dataset transforms meta-theoretical content into structured Q&A pairs, making formal meta-programming and verification concepts more accessible through natural language interactions.
Each entry represents a mathematical statement from MetaCoq (definition… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-MetaCoq-QA.
