datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime_2025
AIME 2025
This dataset contains 30 problems from the 2025 AIME tests, including:
AIME I: 15 problems
AIME II: 15 problems
aime_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from AIME 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (int64): Gold final answer.
problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2025.2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/elonelonelon/2025-challenge-demos.behavior-1k_2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dario-Shit4/behavior-1k_2025-challenge-demos.TencentGR-1M
TencentGR-1M Dataset
Paper | Project Page | Code
TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Constructed from real, de-identified Tencent Ads logs, it aims to address the lack of realistic, public multi-modal datasets in the GR field.
Data Features: Contains rich collaborative IDs and multi-modal representations (text and… See the full description on the dataset page: https://huggingface.co/datasets/TAAC2025/TencentGR-1M.TencentGR-10M
TencentGR-10M Dataset
TAAC2025 Second Round Dataset(2025年腾讯广告算法大赛复赛数据集) TencentGR-10M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Similar to TencentGR-1M, it is constructed from real, de-identified Tencent Ads logs, and aims to address the lack of realistic, public multi-modal datasets in the GR field.
The main differences between TencentGR-10M and TencentGR-1M are:
Dataset Size: Provides 10… See the full description on the dataset page: https://huggingface.co/datasets/TAAC2025/TencentGR-10M.dclm-dedup_20250227-004105xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.btcusdt_spot_1m_03_2023_to_12_2025miklos_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Miklos Schweitzer 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
points (int64): Maximum score for a non-final-answer or proof-style problem.
grading_scheme (list): Rubric… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/miklos_2025.putnam_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from Putnam 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
points (int64): Maximum score for a non-final-answer or proof-style problem.
grading_scheme (list): Rubric or grading… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/putnam_2025.aime-2022-2025
Combined AIME 2021-2025 dataset
This dataset is a combination of:
AI-MO/aimo-validation-aime
Sunny8781/AIME2025_w_solution
with the latter being slightly changed to follow AI-MO's format, namely
changing id from str to int and updating ids to follow previous ones
changing answer from int to str and adding starting "0" to 2 digit answers
add urls
note that the AIME 2025 answers contain only a single answer (the first one)
whereas the AIME 2021-2024 answers contain all of them… See the full description on the dataset page: https://huggingface.co/datasets/allenai/aime-2022-2025.danbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.mahendrawada_2025
Mahendrawada 2025
This data is taken from the Supplement of
Mahendrawada, L., Warfield, L., Donczew, R. et al. Low overlap of transcription factor DNA binding and regulatory targets. Nature 642, 796–804 (2025). https://doi.org/10.1038/s41586-025-08916-0
and GSE236948
Accessing Data
The examples below require
labretriever
(pip install labretriever) and/or the
HuggingFace Hub client
(pip install huggingface_hub).
Accessing Data with labretriever
This… See the full description on the dataset page: https://huggingface.co/datasets/BrentLab/mahendrawada_2025.cve-and-cwe-dataset-1999-2025This collection brings together every Common Vulnerabilities & Exposures (CVE) entry published in the National Vulnerability Database (NVD) from the very first identifier — CVE-1999-0001 — through all records available on 30 May 2025.
It was built automatically with a Python script that calls the NVD REST API v2.0 page-by-page, handles rate-limits, and filters data.
After download each CVE object is pared down to the essentials and written to CVE_CWE_2025.csv with the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/stasvinokur/cve-and-cwe-dataset-1999-2025.fineweb-edu-highest-quality-2025
FineWeb-Edu Highest Quality Dataset (2025 Collection)
Dataset Summary
This dataset contains 4.17 billion tokens of the highest quality educational content, carefully filtered from the FineWeb-Edu dataset's 2025 Common Crawl snapshots. This represents the cream of the crop - only the top ~2% of documents that meet strict quality criteria.
Key Statistics
Total Tokens: 4,176,738,951
Total Documents: 1,477,151
Average Tokens per Document: 2,827
Storage Size: ~11… See the full description on the dataset page: https://huggingface.co/datasets/Yxanul/fineweb-edu-highest-quality-2025.behavior-1k-2025-challenge-demos-debugThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/savoji/behavior-1k-2025-challenge-demos-debug.e621_20250515_selected
内部暂存数据集 (Internal Temporary Dataset)
English
This is a temporary dataset for internal use.
It might contain:
Items being re-processed or corrected (e.g., some images requiring re-tagging using a distributed cluster, needing a convenient data source for it).
Data to supplement our internal systems (e.g., if a machine accidentally lost some images and we don't want to re-download everything).
Recent updates or experimental data not yet finalized (e.g., the image source… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/e621_20250515_selected.aime-1983-2025
AIME Datasets from 1983 to 2025
This dataset contains the AIME datasets from 1983 to 2025.
For AIME 1983 to 2026 use Pandores/aime-1983-2026
Features Description
Feature
Description
Example
year
The year this problem was released. From 1983 to 2025.
2022
index
The index of the problem for a year and part. From 1 to 15.
12
part
The dataset part if this dataset has multiple parts. Can be AIME, AIME I, AIME II or None. Datasets have multiple parts… See the full description on the dataset page: https://huggingface.co/datasets/Pandores/aime-1983-2025.aime_2025_I
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from AIME I 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (int64): Gold final answer.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2025_I.aime_2025_II
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from AIME II 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (int64): Gold final answer.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2025_II.FineWeb2025
FineWeb-Edu 2025 — Cleaned and Shuffled
This dataset is a year-specific, cleaned, shuffled, and sharded release derived from HuggingFaceFW/fineweb-edu. It contains English educational web text collected in Common Crawl snapshots whose dump identifier belongs to 2025.
This is not a news-only dataset. The year refers to the Common Crawl capture year, not necessarily the page's publication year.
Dataset summary
Item
Value
Year
2025
Rows
99,022,205… See the full description on the dataset page: https://huggingface.co/datasets/Yahoo-Finance-News/FineWeb2025.ncaa-college-athlete-rosters-2025-26
NCAA All Sports Rosters 2025-26
A near-census of a full NCAA athletic year — now named and enriched.
513,655 athlete roster records across all 28 championship and emerging
sports, 1,087 schools, all three divisions (D1/D2/D3), men's and women's
teams, one coherent year (2025-26, the first season under the House v. NCAA
settlement). Every field is an institution-published roster fact from official
school athletics sites, validated against the NCAA's official sponsor lists.
New in… See the full description on the dataset page: https://huggingface.co/datasets/dharits3/ncaa-college-athlete-rosters-2025-26.nfl-stats-2021-2025
NFL stats 2021-2025
Per-season parquet files sourced from the nflverse
data releases via nflreadpy. One folder
per dataset, one file per season, manifest.csv at the root with row counts.
Load a single dataset:
from datasets import load_dataset
pbp = load_dataset("dismalnow/nfl-stats-2021-2025", "pbp")
Or read directly with polars/pandas/duckdb from the parquet files.
Datasets
dataset
files
rows
depth_charts
5
149,906
injuries
5
29,151
ngs_passing
5… See the full description on the dataset page: https://huggingface.co/datasets/MikeySports/nfl-stats-2021-2025.hand_gesture_dataThe study is conducted on a total of 7 participants. The participants were instructed to perform three hand gestures (Hold, Single Tap and Double Tap) under different light conditions (low(100-200 lux, medium (600-750 lux) and high (1500-1600 lux)) and at different distances from the light sensor (low(2-4 cm) and high(8-10 cm))
agibot2025seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.Domestic_Futures_ticksdanbooru_2025_recaption
内部暂存数据集 (Internal Temporary Dataset)
English
This is a temporary dataset for internal use.
It might contain:
Items being re-processed or corrected (e.g., some images requiring re-tagging using a distributed cluster, needing a convenient data source for it).
Data to supplement our internal systems (e.g., if a machine accidentally lost some images and we don't want to re-download everything).
Recent updates or experimental data not yet finalized (e.g., the image source… See the full description on the dataset page: https://huggingface.co/datasets/NebulaeWis/danbooru_2025_recaption.
