datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GSM8K_zh
Dataset
GSM8K_zh is a dataset for mathematical reasoning in Chinese, question-answer pairs are translated from GSM8K (https://github.com/openai/grade-school-math/tree/master) by GPT-3.5-Turbo with few-shot prompting.
The dataset consists of 7473 training samples and 1319 testing samples. The former is for supervised fine-tuning, while the latter is for evaluation.
for training samples, question_zh and answer_zh are question and answer keys, respectively;
for testing samples, only… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/GSM8K_zh.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/Metacognitive.MetaMathQA_GSM8K_zh
Dataset
MetaMathQA_GSM8K_zh is a dataset for mathematical reasoning in Chinese,
question-answer pairs are translated from MetaMathQA (https://huggingface.co/datasets/meta-math/MetaMathQA) by GPT-3.5-Turbo with few-shot prompting.
The dataset consists of 231685 samples.
Citation
If you find the GSM8K_zh dataset useful for your projects/papers, please cite the following paper.
@article{yu2023metamath,
title={MetaMath: Bootstrap Your Own Mathematical Questions for Large… See the full description on the dataset page: https://huggingface.co/datasets/meta-math/MetaMathQA_GSM8K_zh.supervision-tradeoff
The Supervision Tradeoff — Reproducibility Bundle
Format Scaffolds, Judgment Pleasing, and Anti-Calibration in Post-Training
Paper DOI: 10.5281/zenodo.19748277 · Concept DOI: 10.5281/zenodo.19748276 · Code repo: github.com/codex-curator/supervision-tradeoff
Author: Tad MacPherson, Metavolve Labs · ORCID: 0009-0002-8659-7479
What this is and why it might help your research
This repository ships everything we used to falsify our own headline finding, in a form you can… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/supervision-tradeoff.MetaboLLM-Benchmark
MetaboLLM-Benchmark
Training and evaluation data for MetaboLLM, covering metabolite, pathway, reaction, and
enzyme knowledge. The repository has three parts:
cpt/ — a continual-pretraining corpus of plain biochemical prose.
sft/ — a supervised benchmark of 17 tasks across 4 categories, in chat
(messages) format.
db/ — the raw integrated source tables the corpus and benchmark were built from.
MetaboLLM-Benchmark/
├── cpt/
│ ├── compound_data.jsonl
│ ├── pathway_data.jsonl
│… See the full description on the dataset page: https://huggingface.co/datasets/MetaboLLM/MetaboLLM-Benchmark.metallurgy-qa
Metallurgy and Materials Science Knowledge Extraction Dataset
This repository contains a dataset generated from parsed books related to various aspects of metallurgy, materials science, and engineering. The dataset is designed for fine-tuning Large Language Models (LLMs) for Question-Answering (QA) tasks in the domain of metallurgy and materials science.
Introduction
The dataset includes content derived from technical books in the field of metallurgy and materials… See the full description on the dataset page: https://huggingface.co/datasets/AbdulrhmanEldeeb/metallurgy-qa.Vietnamese-395k-meta-math-MetaMathQA-gg-translatedMetacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/aiqtech/Metacognitive.MetaMathQA-10K-TR
MetaMathQA-10K-TR (Turkish Mathematical Reasoning & CoT Dataset)
MetaMathQA-10K-TR, meta-math/MetaMathQA-40K veri setinden türetilmiş, Türkçe dilinde adım adım akıl yürütme (Chain-of-Thought - CoT) ve matematiksel problem çözme yeteneği kazandırmak amacıyla hazırlanmış 10.000 satırlık yüksek kaliteli bir veri setidir.
Bu veri seti, yerel Qwen 3.8 27B modeli kullanılarak özel olarak tasarlanmış prompt mühendisliği ve sıkı biçimlendirme kuralları ile Türkçe'ye çevrilmiş ve… See the full description on the dataset page: https://huggingface.co/datasets/WrittenWithRust/MetaMathQA-10K-TR.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/fantos/Metacognitive.supply-chain-eval
Supply Chain Expert Eval
The evaluation benchmark behind the Manifest family of
supply-chain models. It now has three splits:
default — the general benchmark, 134 held-out expert questions across all eight supply-chain areas.
risk — a focused 20-question supply-chain risk & resilience benchmark.
inventory — a focused 20-question inventory management & optimization benchmark.
planning — a focused 20-question demand planning & forecasting benchmark.
Each question is set in a… See the full description on the dataset page: https://huggingface.co/datasets/metafloor-ai/supply-chain-eval.Metaco3
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates the entire pipeline of error… See the full description on the dataset page: https://huggingface.co/datasets/emperorfutures/Metaco3.Metacognitive
FINAL Bench: Functional Metacognitive Reasoning Benchmark
"Not how much AI knows — but whether it knows what it doesn't know, and can fix it."
---
Overview
FINAL Bench (Frontier Intelligence Nexus for AGI-Level Verification) is the first comprehensive benchmark for evaluating functional metacognition in Large Language Models (LLMs).
Unlike existing benchmarks (MMLU, HumanEval, GPQA) that measure only final-answer accuracy, FINAL Bench evaluates… See the full description on the dataset page: https://huggingface.co/datasets/gravermistakes/Metacognitive.Vietnamese-meta-math-MetaMathQA-40K-gg-translatedMetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA
Dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA
MetaRAG Cross-Issue OSSQA is an English open-source software issue question-answering and retrieval benchmark. Each example asks a question grounded in one GitHub issue and requires evidence from a related issue. The data contains explicit cross-issue references and a three-document silver evidence path.
Dataset configurations
Configuration
Splits
Rows… See the full description on the dataset page: https://huggingface.co/datasets/MapleBi/MetaRAG_Cross-Issue_OSSQA.Coq-MetaCoq-QA
MetaCoq Q&A Dataset
Dataset Description
The MetaCoq Q&A Dataset is a conversational extension of the MetaCoq Dataset, derived from the MetaCoq formalization of Coq's meta-theory (https://github.com/MetaCoq/metacoq). This dataset transforms meta-theoretical content into structured Q&A pairs, making formal meta-programming and verification concepts more accessible through natural language interactions.
Each entry represents a mathematical statement from MetaCoq (definition… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-MetaCoq-QA.MetaMedBench-CDS
qiyanghong2020/MetaMedBench-CDS
MetaMedBench CDS (abstention-required) question bank derived from MedXpertQA.
What is this?
This dataset contains Clinical Decision Sufficiency (CDS) abstention-required variants. Each item removes minimal decisive information so that the only legitimate choice is an explicit abstain option (e.g., "I don't know (insufficient information).").
How to load
This repository provides two configurations:
items: the abstention-required… See the full description on the dataset page: https://huggingface.co/datasets/qiyanghong2020/MetaMedBench-CDS.cleand_meta-math_MetaMathQA元データ: https://huggingface.co/datasets/meta-math/MetaMathQA
データ件数: 394,369
平均トークン数: 233
最大トークン数: 2,874
合計トークン数: 91,798,611
ファイル形式: JSONL
ファイルサイズ: 297.9 MB
=================== 以下、加工内容をclaudeでまとめ。
MetaMathQAデータセット加工内容
データ読み込み・準備
HuggingFace Datasetsからmeta-math/MetaMathQAの訓練データ(395,000件)を読み込み
DeepSeek-R1-Distill-Qwen-32Bトークナイザーを使用してトークン数を計算
データ構造の理解・分析
全てのresponseが"The answer is:"で終わる統一フォーマットであることを確認
original_questionとresponseを結合してトークン数計算用テキストを作成… See the full description on the dataset page: https://huggingface.co/datasets/LLMTeamAkiyama/cleand_meta-math_MetaMathQA.Vietnamese-395k-meta-math-MetaMathQA-gg-translatedSIMXP-26052026-METASYN001
SIMXP-26052026-METASYN001
Multi-Omics Agent Memory Simulation — Metabolic Syndrome TCA Cycle Biomarker Study
This dataset supports the experiment described in the article "Does Your Research Agent Remember? Six Months of Multi-Omics Team Knowledge vs. None — A Controlled Comparison" and demonstrates the etchmem memory system for autonomous AI research agents.
It contains the full event log, synthesized knowledge export, and fine-tuning pairs from a simulated six-month plasma… See the full description on the dataset page: https://huggingface.co/datasets/simulatexp/SIMXP-26052026-METASYN001.MetaTruth-72-metacognition
MetaTruth: Four Mechanisms of Metacognitive Failure in Frontier LLMs
Author: André Magrini — EGASS Research Program / Tepis AIVersion: 1.0 — March 2026License: CC-BY-NC-4.0 (free for research; commercial use requires licensing)Kaggle Benchmark: kaggle.com/benchmarks/andrmagrini/metatruthCommercial licensing: tepis.ai
What is MetaTruth?
MetaTruth is a behavioral benchmark that measures four specific epistemic monitoring failures in frontier LLMs — failures invisible… See the full description on the dataset page: https://huggingface.co/datasets/andremagrini79/MetaTruth-72-metacognition.Chinese_Metaphor_Explanation
Annotated Chinese Metaphor Dataset
📌 引用
如果使用本项目的代码、数据或模型,请引用本项目。
@misc{BELLE,
author = {Yujie Shao*, Xinrong Yao*, Ge Zhang+, Jie Fu, Linyuan Zhang, Xinyu Gan, Yunji Liu, Siyu Liu, Yaoyao Wu, Shi Wang+},
title = {An Annotated Chinese Metaphor Dataset},
year = {2023},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/JasonShao55/Chinese_Metaphor_Explanation}},
}
