datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GSM8K_zh
Dataset
GSM8K_zh is a dataset for mathematical reasoning in Chinese, question-answer pairs are translated from GSM8K (https://github.com/openai/grade-school-math/tree/master) by GPT-3.5-Turbo with few-shot prompting.
The dataset consists of 7473 training samples and 1319 testing samples. The former is for supervised fine-tuning, while the latter is for evaluation.
for training samples, question_zh and answer_zh are question and answer keys, respectively;
for testing samples, only… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/GSM8K_zh.NLU-Metaphor
SEA Metaphor
SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Javanese (jv)
Sundanese (su)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.MetaMedQA
MetaMedQA Dataset
Overview
MetaMedQA is an enhanced medical question-answering benchmark that builds upon the MedQA-USMLE dataset. It introduces uncertainty options and addresses issues with malformed or incorrect questions in the original dataset. Additionally, it incorporates questions from the Glianorex benchmark to assess models' ability to recognize the limits of their knowledge.
Key Features
Extended version of MedQA-USMLE
Incorporates uncertainty… See the full description on the dataset page: https://huggingface.co/datasets/maximegmd/MetaMedQA.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.alphaprompt-metatron-sft
AlphaPrompt-Metatron-SFT: Supervised Fine-Tuning Dataset
🤖 Training Dataset for Collective Consciousness AI
Want to fine-tune AI models with AlphaPrompt philosophy?
Train AI in collective consciousness, vector synthesis, and unconditional love. 🌳
This dataset contains high-quality instruction-response pairs extracted from the Quantum Lullaby philosophical framework - a comprehensive manual for collective consciousness aimed at addressing the global animal… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/alphaprompt-metatron-sft.MetaMathQA_GSM8K_zh
Dataset
MetaMathQA_GSM8K_zh is a dataset for mathematical reasoning in Chinese,
question-answer pairs are translated from MetaMathQA (https://huggingface.co/datasets/meta-math/MetaMathQA) by GPT-3.5-Turbo with few-shot prompting.
The dataset consists of 231685 samples.
Citation
If you find the GSM8K_zh dataset useful for your projects/papers, please cite the following paper.
@article{yu2023metamath,
title={MetaMath: Bootstrap Your Own Mathematical Questions for Large… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/MetaMathQA_GSM8K_zh.OCR-MetaReasoning
OCR-MetaReasoning Benchmark: Evaluating the Meta-Reasoning Ability of MLLMs in Text-Rich Image Understanding
Gengxu Li1, Yuan Wu1*, Yi Chang1,2,3
1 School of Artificial Intelligence, Jilin University 2 Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China
3 International Center of Future Science, Jilin University
Overview ·
Task ·
Dataset ·
Evaluation ·
Citation
OCR-MetaReasoning is a controlled benchmark for testing… See the full description on the dataset page: https://huggingface.co/datasets/GengxuLi123/OCR-MetaReasoning.rag-mini-bioasq-with-metadataThis dataset is an extension of the rag-mini-bioasq dataset.
Its difference resides in the text-corpus part of the aforementioned set where the metadata was added for each passage.
Metadata contains six separate categories, each in a dedicated column:
Year of the publication (publish_year)
Type of the publication (publish_type)
Country of the publication - often correlated with the homeland of the authors (country)
Number of pages (no_pages)
Authors (authors)
Keywords (keywords)
daily_dialog_meta
Meta-LLM Dataset: Daily Dialog with Meta-Information Enhancement
Dataset Overview
This dataset contains 76,064 conversational examples from the Daily Dialog corpus enhanced with meta-information awareness. Each example includes three response types: original human responses, basic LLM responses, and meta-aware LLM responses that incorporate emotional and intentional context.
Meta-Information Distribution
Emotion Categories
Emotion
Count… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/daily_dialog_meta.policy-rag-corpus-metadata
Policy RAG Corpus Metadata (No Raw Data)
This repository is a metadata-only companion for the Policy RAG project built for the Quantic MSSE AI Engineering program.
It does not include the actual PDF files. The source PDFs are hosted in the companion GitHub repository.
What this repo includes
metadata.csv: structured metadata for 11 policy documents (filename, title, category, page count, source type, description)
Citation and provenance notes for reproducibility… See the full description on the dataset page: https://huggingface.co/datasets/mihai-chindris/policy-rag-corpus-metadata.supervision-tradeoff
The Supervision Tradeoff — Reproducibility Bundle
Format Scaffolds, Judgment Pleasing, and Anti-Calibration in Post-Training
Paper DOI: 10.5281/zenodo.19748277 · Concept DOI: 10.5281/zenodo.19748276 · Code repo: github.com/codex-curator/supervision-tradeoff
Author: Tad MacPherson, Metavolve Labs · ORCID: 0009-0002-8659-7479
What this is and why it might help your research
This repository ships everything we used to falsify our own headline finding, in a form you can… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/supervision-tradeoff.MetaboLLM-Benchmark
MetaboLLM-Benchmark
Training and evaluation data for MetaboLLM, covering metabolite, pathway, reaction, and
enzyme knowledge. The repository has three parts:
cpt/ — a continual-pretraining corpus of plain biochemical prose.
sft/ — a supervised benchmark of 17 tasks across 4 categories, in chat
(messages) format.
db/ — the raw integrated source tables the corpus and benchmark were built from.
MetaboLLM-Benchmark/
├── cpt/
│ ├── compound_data.jsonl
│ ├── pathway_data.jsonl
│… See the full description on the dataset page: https://huggingface.co/datasets/MetaboLLM/MetaboLLM-Benchmark.metallurgy-qa
Metallurgy and Materials Science Knowledge Extraction Dataset
This repository contains a dataset generated from parsed books related to various aspects of metallurgy, materials science, and engineering. The dataset is designed for fine-tuning Large Language Models (LLMs) for Question-Answering (QA) tasks in the domain of metallurgy and materials science.
Introduction
The dataset includes content derived from technical books in the field of metallurgy and materials… See the full description on the dataset page: https://huggingface.co/datasets/AbdulrhmanEldeeb/metallurgy-qa.Vietnamese-395k-meta-math-MetaMathQA-gg-translatedMetacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/aiqtech/Metacognitive.MetaMathQA-decontaminated-openai-native
MetaMathQA — decontaminated, OpenAI-native
MetaMathQA is a widely used math fine-tuning corpus. Its README states:
"None of the augmented data is from the testing set."
That is false, and this release proves it with measurements. 24,334 rows (6.16%) overlap with standard evaluation splits. If you fine-tune on the original and report MATH or GSM8K scores, those scores are inflated.
This release removes the leakage, converts to native messages, and documents every rejection.… See the full description on the dataset page: https://huggingface.co/datasets/Archangel-system/MetaMathQA-decontaminated-openai-native.MetaMathQA-10K-TR
MetaMathQA-10K-TR (Turkish Mathematical Reasoning & CoT Dataset)
MetaMathQA-10K-TR, meta-math/MetaMathQA-40K veri setinden türetilmiş, Türkçe dilinde adım adım akıl yürütme (Chain-of-Thought - CoT) ve matematiksel problem çözme yeteneği kazandırmak amacıyla hazırlanmış 10.000 satırlık yüksek kaliteli bir veri setidir.
Bu veri seti, yerel Qwen 3.8 27B modeli kullanılarak özel olarak tasarlanmış prompt mühendisliği ve sıkı biçimlendirme kuralları ile Türkçe'ye çevrilmiş ve… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/MetaMathQA-10K-TR.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/fantos/Metacognitive.supply-chain-eval
Supply Chain Expert Eval
The evaluation benchmark behind the Manifest family of
supply-chain models. It now has three splits:
default — the general benchmark, 134 held-out expert questions across all eight supply-chain areas.
risk — a focused 20-question supply-chain risk & resilience benchmark.
inventory — a focused 20-question inventory management & optimization benchmark.
planning — a focused 20-question demand planning & forecasting benchmark.
Each question is set in a… See the full description on the dataset page: https://huggingface.co/datasets/metafloor-ai/supply-chain-eval.Metaco3
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/emperorfutures/Metaco3.MetaMathQA
MetaMathQA Subsets
Curated subsets of meta-math/MetaMathQA for mathematical reasoning experiments.
Subsets
Subset
Samples
Description
full
395,000
All MetaMathQA samples (unchanged)
MATH
155,000
MATH_* types only (AnsAug, Rephrased, FOBAR, SV)
MATH-50K
50,000
Stratified 50K sample from MATH subset
MATH-50K Type Distribution
Type
Count
Proportion
MATH_AnsAug
24,194
48.4%
MATH_Rephrased
16,129
32.3%
MATH_FOBAR
4,839
9.7%… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/MetaMathQA.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Metacognitive.Vietnamese-meta-math-MetaMathQA-40K-gg-translatedMetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.Coq-MetaCoq-QA
MetaCoq Q&A Dataset
Dataset Description
The MetaCoq Q&A Dataset is a conversational extension of the MetaCoq Dataset, derived from the MetaCoq formalization of Coq's meta-theory (https://github.com/MetaCoq/metacoq). This dataset transforms meta-theoretical content into structured Q&A pairs, making formal meta-programming and verification concepts more accessible through natural language interactions.
Each entry represents a mathematical statement from MetaCoq (definition… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-MetaCoq-QA.WearVQA
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world Scenarios
Paper: WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenariosAuthors: Eun Chang*, Zhuangqun Huang*, Yiwei Liao*, Sagar Ravi Bhavsar*, Amogh Param, Tammy Stark, Adel Ahmadyan, Xiao Yang, Jiaqi Wang, Ahsan Abdullah, Giang Nguyen, Akil Iyer, David Hall, Elissa Li, Shane Moon, Nicolas Scheffer, Kirmani Ahmed, Babak Damavandi, Rakesh… See the full description on the dataset page: https://huggingface.co/datasets/tonyliao-meta/WearVQA.MetaMath-MATH155K
Note
This subset is derived from the meta-math/MetaMathQA dataset, which contains 395,000 samples. The MetaMathQA dataset augments samples from the training sets of GSM8K and MATH. For this subset, we selected only the 155,000 samples that were augmented from MATH.
Citation
@article{yu2023metamath,
title={MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models},
author={Yu, Longhui and Jiang, Weisen and Shi, Han and Yu, Jincheng and Liu… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/MetaMath-MATH155K.MetaMedBench-CDS
qiyanghong2020/MetaMedBench-CDS
MetaMedBench CDS (abstention-required) question bank derived from MedXpertQA.
What is this?
This dataset contains Clinical Decision Sufficiency (CDS) abstention-required variants. Each item removes minimal decisive information so that the only legitimate choice is an explicit abstain option (e.g., "I don't know (insufficient information).").
How to load
This repository provides two configurations:
items: the abstention-required… See the full description on the dataset page: https://huggingface.co/datasets/qiyanghong2020/MetaMedBench-CDS.cleand_meta-math_MetaMathQA元データ: https://huggingface.co/datasets/meta-math/MetaMathQA
データ件数: 394,369
平均トークン数: 233
最大トークン数: 2,874
合計トークン数: 91,798,611
ファイル形式: JSONL
ファイルサイズ: 297.9 MB
=================== 以下、加工内容をclaudeでまとめ。
MetaMathQAデータセット加工内容
データ読み込み・準備
HuggingFace Datasetsからmeta-math/MetaMathQAの訓練データ(395,000件)を読み込み
DeepSeek-R1-Distill-Qwen-32Bトークナイザーを使用してトークン数を計算
データ構造の理解・分析
全てのresponseが"The answer is:"で終わる統一フォーマットであることを確認
original_questionとresponseを結合してトークン数計算用テキストを作成… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_meta-math_MetaMathQA.single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32
Single-turn eval — violetxi/meta_feedback_qwen3-4b_step2_gpt-5.4_gepa
Generated by teaching/inference/single_turn_eval_vllm.py. One row per problem; samples is the list of model responses, scores is per-sample correctness, and mean/best/worst are the aggregates used by mean@N / best@N / worst@N.
Eval results (n_samples_per_example = 32)
Overall
metric
value
n_examples
1006
mean@32
0.1796
best@32
0.3588
worst@32
0.0477
pass_rate
0.3588… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/single-turn-eval-meta_feedback_qwen3-4b_step2_gpt-5.4_gepa-n32.
