datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-stack-metadata
Dataset Card for The Stack Metadata
Changelog
Release
Description
v1.1
This is the first release of the metadata. It is for The Stack v1.1
v1.2
Metadata dataset matching The Stack v1.2
Dataset Summary
This is a set of additional information for repositories used for The Stack. It contains file paths, detected licenes as well as some other information for the repositories.
Supported Tasks and Leaderboards
The main… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-metadata.arxiv-metadata-snapshot
Dataset Card for "arxiv-metadata-oai-snapshot"
More Information needed
This is a mirror of the metadata portion of the arXiv dataset.
The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset.
Metadata
This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing:
id: ArXiv ID (can be used to access the paper, see below)
submitter:… See the full description on the dataset page: https://huggingface.co/datasets/librarian-bots/arxiv-metadata-snapshot.Olympiads
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 32926
Filtered size: 32926
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads.NLU-Metaphor
SEA Metaphor
SEA Metaphor evaluates a model's ability to interpret paired figurative phrases with divergent meanings. It is sampled from Multilingual-Fig-QA for Indonesian, Javanese, and Sundanese.
Supported Tasks and Leaderboards
SEA Metaphor is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.
Languages
Indonesian (id)
Javanese (jv)
Sundanese (su)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Metaphor.MetaMathQA-R1
oumi-ai/MetaMathQA-R1
MetaMathQA-R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning.
Prompts were augmented from GSM8K and MATH training sets with responses directly from DeepSeek-R1.
MetaMathQA-R1 was used to train MiniMath-R1-1.5B, which achieves 44.4% accuracy on MMLU-Pro-Math, the highest of any model with <=1.5B parameters.
Curated by: Oumi AI using Oumi inference on Parasail
Language(s) (NLP): English
License:… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/MetaMathQA-R1.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.arxiv-metadata-snapshot
Dataset Card for "arxiv-metadata-oai-snapshot"
More Information needed
This is a mirror of the metadata portion of the arXiv dataset.
The sync will take place weekly so may fall behind the original datasets slightly if there are more regular updates to the source dataset.
Metadata
This dataset is a mirror of the original ArXiv data. This dataset contains an entry for each paper, containing:
id: ArXiv ID (can be used to access the paper, see below)
submitter: Who… See the full description on the dataset page: https://huggingface.co/datasets/Rurouni-II/arxiv-metadata-snapshot.MetaMathQA
Meta Math Filtered
This is a combined and filtered (removed all the redundant rows) version of meta-math/MetaMathQA and meta-math/MetaMathQA-40K
Usage
from datasets import load_dataset
dataset = load_dataset("Sharathhebbar24/MetaMathQA", split="train")
MetaMathQA-math-500DAPO-Math-17k with the MATH-500 test split converted to the same parquet schema and prompt format.
MetaSyn
MetaSyn
MetaSyn is a benchmark for protocol-driven scientific evidence synthesis. It
contains 422 Nature Portfolio source reviews and a shared corpus of 140,585
PubMed articles, with 336 training and 86 test instances.
Resources
Paper: arXiv:2606.17041
Code and evaluator: THUIR/MetaSyn
Trained retriever: BFTree/MA-Retriever
Configurations
reviews contains source-review records, PI/ECO fields, search and eligibility
information, synthesis… See the full description on the dataset page: https://huggingface.co/datasets/THUIR/MetaSyn.thahabiorg_metadata
📖 Thahabi Books Metadata Dataset
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book and includes bibliographic information such as title, author, category, and source details.
📦 Dataset Structure
This repository contains structured metadata for 28,896 Arabic books scraped from thahabi.org.
Each row represents one book with full bibliographic and structural information.
📚… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_metadata.daily_dialog_meta
Meta-LLM Dataset: Daily Dialog with Meta-Information Enhancement
Dataset Overview
This dataset contains 76,064 conversational examples from the Daily Dialog corpus enhanced with meta-information awareness. Each example includes three response types: original human responses, basic LLM responses, and meta-aware LLM responses that incorporate emotional and intentional context.
Meta-Information Distribution
Emotion Categories
Emotion
Count… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/daily_dialog_meta.cellarc_100k_meta
cellarc_100k_meta
CellARC 100k Meta is the metadata‑rich variant of the CellARC benchmark introduced in Lzicar, M. (2025). CellARC: Measuring Intelligence with Cellular Automata. It contains the exact same episodes and splits as cellarc_100k, with byte‑identical Parquet files; the JSONL files retain full per‑episode metadata (rule tables, coverage diagnostics, morphology descriptors, sampling parameters, etc.). Each episode exposes five support pairs plus a held‑out query/solution… See the full description on the dataset page: https://huggingface.co/datasets/mireklzicar/cellarc_100k_meta.Olympiads_hard
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 21525
Filtered size: 21408
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_hard.Olympiads_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 13284
Filtered size: 13240
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Olympiads_medium.DPL-meta
Difference-aware Personalized Learning (DPL) Dataset
This dataset is used in the paper:
Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization
Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yimeng Bai, Wenjie Wang, Hong Cheng, Fuli Feng, Tat-Seng Chua
Code: https://github.com/SnowCharmQ/DPL
This dataset is an adaptation of the Amazon Reviews'23 dataset. It contains item metadata for Books, CDs & Vinyl, and Movies & TV. Each item includes title… See the full description on the dataset page: https://huggingface.co/datasets/SnowCharmQ/DPL-meta.task1394_meta_woz_task_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1394_meta_woz_task_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1394_meta_woz_task_classification.Numina_medium
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 37133
Filtered size: 37133
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina_medium.danbooru2023-metadata-database
Metadata Database for Danbooru2023
Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023
The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset.
This dataset contains a sqlite db file which have all the tags and posts metadata in it.
The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together)
The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.MetaMathQA
MetaMathQA Subsets
Curated subsets of meta-math/MetaMathQA for mathematical reasoning experiments.
Subsets
Subset
Samples
Description
full
395,000
All MetaMathQA samples (unchanged)
MATH
155,000
MATH_* types only (AnsAug, Rephrased, FOBAR, SV)
MATH-50K
50,000
Stratified 50K sample from MATH subset
MATH-50K Type Distribution
Type
Count
Proportion
MATH_AnsAug
24,194
48.4%
MATH_Rephrased
16,129
32.3%
MATH_FOBAR
4,839
9.7%… See the full description on the dataset page: https://huggingface.co/datasets/mtybilly/MetaMathQA.oellm-dpo-metadataset
OpenEuroLLM DPO metadataset
A lightweight, versioned source of truth for building preference-training data for OpenEuroLLM. It contains metadata and planning decisions—not copies of upstream training examples.
The catalogue pins each upstream revision and records its license, size, language coverage, pair schema, overlap family, decision, risks, and required transformations. Upstream licenses and terms still apply. The Apache-2.0 license in this repository covers only the… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-dpo-metadataset.DPL-meta
Difference-aware Personalized Learning (DPL) Dataset
This dataset is used in the paper:
Measuring What Makes You Unique: Difference-Aware User Modeling for Enhancing LLM Personalization
Yilun Qiu, Xiaoyan Zhao, Yang Zhang, Yimeng Bai, Wenjie Wang, Hong Cheng, Fuli Feng, Tat-Seng Chua
Code: https://github.com/SnowCharmQ/DPL
This dataset is an adaptation of the Amazon Reviews'23 dataset. It contains item metadata for Books, CDs & Vinyl, and Movies & TV. Each item includes… See the full description on the dataset page: https://huggingface.co/datasets/uuuue/DPL-meta.Numina
Numina-Olympiads
Filtered NuminaMath-CoT dataset containing only olympiads problems with valid answers.
Dataset Information
Split: train
Original size: 210350
Filtered size: 210350
Source: olympiads
All examples contain valid boxed answers
Dataset Description
This dataset is a filtered version of the NuminaMath-CoT dataset, containing only problems from olympiad sources that have valid boxed answers. Each example includes:
A mathematical word problem
A… See the full description on the dataset page: https://huggingface.co/datasets/Metaskepsis/Numina.Metamath
Metamath
A structured dataset of formally verified theorems and axioms from Metamath, one of the largest collections of rigorously verified mathematics in the world.
Source
Repository: https://github.com/metamath/set.mm
Commit: 160dfc7e4ec5f201f5bae4ca5a5eeb67242902b5
Files: 5
License: cc0-1.0
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Metamath.Coq-MetaCoq
Coq-MetaCoq
Structured dataset of formalizations from MetaCoq (Coq meta-theory formalized in Coq).
Source
Repository: https://github.com/MetaCoq/metacoq
Commit: 971b2cc5c8bdc011068f67fa272245d8e5b54209
Files: 597
License: mit
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body, empty if the… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-MetaCoq.math_metaqa_dialogue_en
Description
The dataset is from unknown, formatted as dialogues for speed and ease of use. Many thanks to author for releasing it.
Importantly, this format is easy to use via the default chat template of transformers, meaning you can use huggingface/alignment-handbook immediately, unsloth.
Structure
View online through viewer.
Note
We advise you to reconsider before use, thank you. If you find it useful, please like and follow this account.… See the full description on the dataset page: https://huggingface.co/datasets/lamhieu/math_metaqa_dialogue_en.openhermes-preferences-metamath
Dataset Card for OpenHermes Preferences - MetaMath
This dataset is a subset from argilla/OpenHermesPreferences,
only keeping the preferences of metamath, and removing all the columns besides the chosen and rejected ones, that
come in OpenAI chat formatting, so that's easier to fine-tune a model using tools like: huggingface/alignment-handbook
or axolotl, among others.
Reference
argilla/OpenHermesPreferences dataset created as a collaborative
effort between Argilla and… See the full description on the dataset page: https://huggingface.co/datasets/alvarobartt/openhermes-preferences-metamath.meta-meme
Meta-Meme Consultation URLs Dataset
Description
This dataset contains 2177 consultation URLs generated from the Meta-Meme formally verified system. Each URL represents a consultation with one of 9 AI muses about a specific file in the repository.
Dataset Structure
file: Path to the file in the repository
muse: Assigned AI muse (Calliope, Clio, Erato, Euterpe, Melpomene, Polyhymnia, Terpsichore, Thalia, Urania)
tool: Consultation tool (llm, lean4, rustc… See the full description on the dataset page: https://huggingface.co/datasets/introspector/meta-meme.Coq-Metalib
Coq-Metalib
Structured declarations from Metalib - a Coq library for programming language metatheory using locally nameless representation. Source: github.com/plclub/metalib
Source
Repository: https://github.com/plclub/metalib
Commit: 144ddcd0fff6717229140314cf559d85fad6eae0
Files: 33
License: mit
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Metalib.ai-vs-human-meta-llama-Llama-3.2-1B-Instruct
AI vs Human dataset on the CNN Daily mails
Dataset Description
This dataset showcases pairs of truncated articles and their respective completions, crafted either by humans or an AI language model.
Each article was randomly truncated between 25% and 50% of its length.
The language model was then tasked with generating a completion that mirrored the characters count of the original human-written continuation.
Data Fields
'human': The original human-authored… See the full description on the dataset page: https://huggingface.co/datasets/zcamz/ai-vs-human-meta-llama-Llama-3.2-1B-Instruct.
