datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.arxiv-cs-2020-2025-pdfsdom-pi-pdfs-2025
DOM-PI 2025 — PDFs-fonte (Diário Oficial dos Municípios do Piauí)
PDFs originais das publicações de 2025 do Diário Oficial dos Municípios do Piauí,
organizados por Território de Desenvolvimento. São a fonte da qual o corpus textual foi
extraído por OCR/parsing. 41.617 PDFs · ~70 GB.
Dataset de texto derivado (carregável, com limpeza e tiers de qualidade):
gutoportelaa/dom-pi-corpus-2025.
Cobertura de PDFs (parcial): presentes 7 territórios — tabuleiros_alto_parnaiba… See the full description on the dataset page: https://huggingface.co/datasets/gutoportelaa/dom-pi-pdfs-2025.Gen-HumanEgo
Gen-HumanEgo
1,800+ hours of egocentric human demonstrations with synchronized, structured supervision for embodied AI and robot learning.
Gen-HumanEgo contains real-world first-person demonstrations collected across diverse tasks, people, environments, and ways of performing activities using a unified six-camera DAS-Ego setup. GenRobot's Data Foundation Model (DFM) processes the recordings to provide complementary training signals for human motion, spatial understanding… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/Gen-HumanEgo.2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/behavior-1k/2025-challenge-demos.aime_2025
AIME 2025
This dataset contains 30 problems from the 2025 AIME tests, including:
AIME I: 15 problems
AIME II: 15 problems
aime_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from AIME 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (int64): Gold final answer.
problem_type… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/aime_2025.2025-challenge-orchestrator-v3_2github-code-2025-language-split
📜 Source Data & Attribution
This dataset is a processed derivative of nick007x/github-code-2025.
Origination
The original data was aggregated by nick007x from public GitHub repositories. We have retained the original content, file paths, and metadata while restructuring the format for easier consumption by language-specific models.
Processing Steps
To create this dataset, we performed the following processing on the source data:
Language… See the full description on the dataset page: https://huggingface.co/datasets/hasankursun/github-code-2025-language-split.hmmt_feb_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from HMMT February 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
problem (string): Problem statement, usually stored as LaTeX source.
answer (string): Gold final answer.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/hmmt_feb_2025.IROS-2025-Challenge-Manip
IROS-2025-Challenge-Manip
Dataset Summary 📖
This dataset contains the IROS Challenge - Manipulation Track benchmark, organized into pretrain, train, and validation splits.
Pretrain split: ~20,000 single pick-and-place trajectories, packaged into tar files (each containing ~1,000 trajectories).
Train split: task-specific demonstrations, with ~100 trajectories provided per task.
Validation split: includes the test-time scenes and object assets in USD format.
Each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/IROS-2025-Challenge-Manip.ICLR_2025_OCR
ICLR_2025_OCR
OCR Data.
AIME2025
AIME 2025 Dataset
Dataset Description
This dataset contains problems from the American Invitational Mathematics Examination (AIME) 2025-I & II.
Arxiv_2025_OCR
Arxiv_2025_OCR
OCR Data.
hmmt_nov_2025
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
This dataset contains the questions from HMMT November 2025 used for the MathArena Leaderboard
Data Fields
The dataset contains the following fields:
problem_idx (int64): Problem index within the corresponding MathArena benchmark.
answer (string): Gold final answer.
problem (string): Problem statement, usually stored as LaTeX source.… See the full description on the dataset page: https://huggingface.co/datasets/MathArena/hmmt_nov_2025.scref_ICLR_2025
scREF
This dataset contains human single cell RNA-sequencing (scRNA-seq) data collected from 46 studies and standardized
by Diaz-Mejia JJ et al. (2025) for the paper Benchmarking and optimizing organism wide single-cell RNA alignment methods presented at the LMRL Workshop at the International Conference on Learning Representations (2025).
Folder Phenomic-AI/scref_ICLR_2025/zarr contains standardized single-cell RNA data for each study in zarr format.
Sub-folder names show: {first… See the full description on the dataset page: https://huggingface.co/datasets/Phenomic-AI/scref_ICLR_2025.MARS-Hyperspectral-EnMAP-PRISMA-v2025
MARS-Hyperspectral dataset (EnMAP and PRISMA) - version v2025
Updated version of the dataset. More information will be added soon.
crates-20250307FACETCC-MAIN-2025-08
CC-MAIN-2025-08へようこそ
本データセットはCommonCrawlerと呼ばれるものから日本語のみを抽出したものです。
利用したものはcc-downloader-rsです。
なおIPAのICSCoEと呼ばれるところから資源を借りてやりましたゆえに、みなさんIPAに感謝しましょう。
※ IPAは独立行政法人 情報処理推進機構のことです。テストに出ますので覚えましょう。
利用について
本利用は研究目的のみとさせていただきます。
それ以外の利用につきましては途方のくれない数の著作権者に許可を求めてきてください。
2025-challenge-rawdataOpenalex-2005-20252025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/elonelonelon/2025-challenge-demos.behavior-1k_2025-challenge-demosThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "R1Pro",
"total_episodes": 10000,
"total_frames": 119094660,
"total_tasks": 50,
"total_videos": 90000,
"chunks_size": 10000,
"fps": 30,
"splits": {
"train": "0:10000"
},
"data_path": "data/task-{episode_chunk:04d}/episode_{episode_index:08d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Dario-Shit4/behavior-1k_2025-challenge-demos.CC-MAIN-2025-33SCPWiki-Archive-02-March-2025-Datasets2025-challenge-task-instancesCC-MAIN-2025-05
CC-MAIN-2025-05へようこそ
本データセットはCommonCrawlerと呼ばれるものから日本語のみを抽出したものです。
利用したものはcc-downloader-rsです。
なおIPAのICSCoEと呼ばれるところから資源を借りてやりましたゆえに、みなさんIPAに感謝しましょう。
※ IPAは独立行政法人 情報処理推進機構のことです。テストに出ますので覚えましょう。
利用について
本利用は研究目的のみとさせていただきます。
それ以外の利用につきましては途方のくれない数の著作権者に許可を求めてきてください。
TencentGR-1M
TencentGR-1M Dataset
Paper | Project Page | Code
TAAC2025 Preliminary Round Dataset (2025年腾讯广告算法大赛初赛数据集) TencentGR-1M Dataset is a large-scale, all-modality dataset designed specifically for generative recommendation (GR) in industrial advertising. Constructed from real, de-identified Tencent Ads logs, it aims to address the lack of realistic, public multi-modal datasets in the GR field.
Data Features: Contains rich collaborative IDs and multi-modal representations (text and… See the full description on the dataset page: https://huggingface.co/datasets/TAAC2025/TencentGR-1M.20250704_beifen
