datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime_2025
AIME 2025
This dataset contains 30 problems from the 2025 AIME tests, including:
AIME I: 15 problems
AIME II: 15 problems
aime_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from AIME 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (int64): Gold final answer.
problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2025.github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.hmmt_feb_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from HMMT February 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (string): Gold final answer.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/hmmt_feb_2025.AIME2025
AIME 2025 Dataset
Dataset Description
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II.
hmmt_nov_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from HMMT November 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
problem (string): Problem statement, usually stored as LaTeX source.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/hmmt_nov_2025.crates-20250307CC-MAIN-2025-08
CC-MAIN-2025-08へようこそ
本データセットはCommonCrawlerと呼ばれるものから日本語のみを抽出したものです。
利用したものはcc-downloader-rsです。
なおIPAのICSCoEと呼ばれるところから資源を借りてやりましたゆえに、みなさんIPAに感謝しましょう。
※ IPAは独立行政法人 情報処理推進機構のことです。テストに出ますので覚えましょう。
利用について
本利用は研究目的のみとさせていただきます。
それ以外の利用につきましては途方のくれない数の著作権者に許可を求めてきてください。
CC-MAIN-2025-33SCPWiki-Archive-02-March-2025-Datasets2025-challenge-task-instancesCC-MAIN-2025-05
CC-MAIN-2025-05へようこそ
本データセットはCommonCrawlerと呼ばれるものから日本語のみを抽出したものです。
利用したものはcc-downloader-rsです。
なおIPAのICSCoEと呼ばれるところから資源を借りてやりましたゆえに、みなさんIPAに感謝しましょう。
※ IPAは独立行政法人 情報処理推進機構のことです。テストに出ますので覚えましょう。
利用について
本利用は研究目的のみとさせていただきます。
それ以外の利用につきましては途方のくれない数の著作権者に許可を求めてきてください。
TencentGR-1M
TencentGR-1M Dataset
Paper | Project Page | Code
TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Constructed from real, de-identified Tencent Ads logs, it aims to address the lack of realistic, public multi-modal datasets in the GR field.
Data Features: Contains rich collaborative IDs and multi-modal representations (text and… See the full description on the dataset page: https://huggingface.co/datasets/TAAC2025/TencentGR-1M.github-code-2025perl-20250811raku-20250811apex_2025
UPDATE: The dataset now contains the remaining samples from SMT 2025!
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from MathArena Apex 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/apex_2025.brumo_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from BRUMO 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (string): Gold final answer.
problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/brumo_2025.CC-MAIN-2025-26TencentGR-10M
TencentGR-10M Dataset
TAAC2025 Second Round Dataset(2025年腾讯广告算法大赛复赛数据集) TencentGR-10M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Similar to TencentGR-1M, it is constructed from real, de-identified Tencent Ads logs, and aims to address the lack of realistic, public multi-modal datasets in the GR field.
The main differences between TencentGR-10M and TencentGR-1M are:
Dataset Size: Provides 10… See the full description on the dataset page: https://huggingface.co/datasets/TAAC2025/TencentGR-10M.smt_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from SMT 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
problem_type (list[string]): Problem type/category labels.
problem (string): Problem… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/smt_2025.cmimc_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from CMIMC 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (string): Gold final answer.
problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/cmimc_2025.github-code-2025
🚀 GitHub Code 2025: The Clean Code Manifesto
A meticulously curated dataset of 1.5M+ repositories representing both quality and innovation in 2025's code ecosystem
🌟 The Philosophy
Quality Over Quantity, Purpose Over Volume
In an era of data abundance, we present a dataset built on radical curation. Every file, every repository, every byte has been carefully selected to represent the signal in the noise of open-source development.
🎯 What This Dataset Is… See the full description on the dataset page: https://huggingface.co/datasets/hybridfree/github-code-2025.kl3m-data-snapshot-20250324NayanaOCR_Corpus_2025
🪷 NayanaOCR Corpus 2025
A 1M-page, 22-language fully-parallel synthetic OCR + VQA corpus for document-centric vision-language models — every page rendered in every language.
NayanaOCR Corpus 2025 is one of the largest open-source multilingual, multi-task document datasets for training and evaluating OCR, layout detection, and visual question answering (VQA) in low-resource and underrepresented languages.
The headline property: it's a true parallel corpus. The same ~45,700 source… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/NayanaOCR_Corpus_2025.dclm-dedup_20250227-004105xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.usamo_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from USAMO 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (string): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
points (int64): Maximum score for a… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/usamo_2025.imc_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from IMC 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (string): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
points (int64): Maximum score for a… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/imc_2025.CC-MAIN-2025-18
