datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fondant-cc-25m
Dataset Card for Fondant Creative Commons 25 million (fondant-cc-25m)
Changelog
Release
Description
v0.1
Release of the Fondant-cc-25m dataset
Dataset Summary
Fondant-cc-25m contains 25 million image URLs with their respective Creative Commons
license information collected from the Common Crawl web corpus.
The dataset was created using Fondant, an open source framework that aims to simplify and speed up
large-scale data processing by making… See the full description on the dataset page: https://huggingface.co/datasets/fondant-ai/fondant-cc-25m.25m-img-capssee https://huggingface.co/datasets/csarron/4m-img-caps for example usage
movielens-25m-thumb
🍿 Popcorn Thumbnails Embeddings
This dataset contains deep visual features obtained from +65000 movie thumbnails.
It contains extracted visual features using modern VLMs.
To simply load it, Popcorn framework has been developed that can be used in movie recommendation, information retrieval, classification, etc tasks.
📚 Citation
@article{popcorn,
title={Popcorn: A Configurable Benchmark for Visual Evidence in Multimodal Movie Recommendation},
author={Tourani… See the full description on the dataset page: https://huggingface.co/datasets/alitourani/movielens-25m-thumb.MedTrinity-25M
Tutorial of using Medtrinity-25M
MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities, with multigranular annotations for more than 65 diseases. These enriched annotations encompass both global textual information, such as disease/lesion type, modality, region-specific descriptions, and inter-regional relationships, as well as detailed local annotations for regions of interest (ROIs), including bounding… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/MedTrinity-25M.vlite7-mini-25m-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite7-mini-25m-dataset.maze2d-large-diverse-25mapsvlite7-mini-25m-eng-dataset
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite7-mini-25m-eng-dataset.dingo-t1-waveforms-25M
DINGO-T1 waveform dataset — 25M (IMRPhenomXPHM, multibanded FD, SVD-200)
25,000,000 frequency-domain precessing-BBH waveforms generated with the dingo pipeline,
matching the DINGO-T1 (arXiv:2512.02968) data setup.
Approximant: IMRPhenomXPHM (precession + higher modes), f_ref = 20 Hz
Domain: MultibandedFrequencyDomain, f ∈ [20, 1810] Hz, base δf = 0.125 → 1104 MFD bins
Compression: whitening (aLIGO_ZERO_DET_high_P_asd.txt) + SVD basis size 200 per polarization
Intrinsic prior:… See the full description on the dataset page: https://huggingface.co/datasets/Ashray-g/dingo-t1-waveforms-25M.vlite7-mini-25m-dataset-eng
ISAI - 이사이
I’m an independent developer building and maintaining AI projects on my own.
Everything from model development to server costs, datasets, and feature updates is managed personally.
Any support you can provide greatly helps keep this project running and allows for continuous improvements.
If you find this project helpful, please consider supporting my work. Thank you.
혼자서 AI 프로젝트를 개발하고 운영하고 있습니다.
모델 개발부터 데이터셋 준비, 서버 비용 감당, 기능 업데이트까지 모두 직접 진행하고 있습니다.
보내주시는 따뜻한 후원은 안정적인… See the full description on the dataset page: https://huggingface.co/datasets/aixk/vlite7-mini-25m-dataset-eng.swev-trm-trajectories-25models
SWE-Bench Verified TRM Trajectories (25 Models, Verified Labels)
Trajectories from 25 LLMs attempting SWE-Bench Verified tasks, formatted for
training a Trajectory Reward Model (TRM). Each record is one model's full
multi-turn attempt at one task, labeled with the real SWE-bench harness
verdict (scores.resolved).
Splits
Split
Records
Tasks
Pos
Neg
train
10,107
405
6,108
3,999
val
2,366
95
1,440
926
Train/val are task-disjoint (stable hash on task_id… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/swev-trm-trajectories-25models.ml-25mmaze-navigation-25x25-25mazes-250k-steps400-bias0.01
Dataset Card for "maze-navigation-25x25-25mazes-250k-steps400-bias0.01"
More Information needed
ML-25mamazon_split_25M_reviews_20_percent_condensedmaze-navigation-25x25-25mazes-250k-steps400-bias0.001
Dataset Card for "maze-navigation-25x25-25mazes-250k-steps400-bias0.001"
More Information needed
maze-navigation-25x25-25mazes-250k-steps250-bias0.3
Dataset Card for "maze-navigation-25x25-25mazes-250k-steps250-bias0.3"
More Information needed
25_merged_noresizeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 25,
"total_frames": 35765,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mpacek/25_merged_noresize.maze-navigation-50x50-25mazes-250k-steps250-bias0.3
Dataset Card for "maze-navigation-50x50-25mazes-250k-steps250-bias0.3"
More Information needed
25-MDSCI25_merged_resizeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 25,
"total_frames": 35765,
"total_tasks": 2,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mpacek/25_merged_resize.movielens-25m-ratingsThis is a dataset that streams user ratings from the MovieLens 25M dataset from the MovieLens servers.data_29_10_2024_25m_ver_8_check_2swev-trm-trajectory-embeddings-25models
SWE-Bench Verified TRM Trajectory Embeddings (25 Models, task-disjoint)
Precomputed trajectory embeddings for the 25-model SWE-Bench Verified TRM set,
the embedding companion to
tarsur385/swev-trm-trajectories-25models.
Same trajectories, same task-disjoint split (stable hash on task_id,
val_ratio=0.20, seed=42); trajectory_id matches the text dataset 1:1.
Each trajectory is embedded with Qwen/Qwen3.6-27B (mean-pooled last hidden
state), hidden size 5120.
Split… See the full description on the dataset page: https://huggingface.co/datasets/tarsur385/swev-trm-trajectory-embeddings-25models.MedTrinity-25M-demo-shuffle-45000-indonesianamazon_25M_10_000_condenseddata_29_10_2024_25mFreeChessGames25M
Free Synthetic Chess Games (25M)
A synthetic dataset of 25,003,202 individual chess moves across 92,733 fully legal games, generated move-by-move with real chess-rules validation (every move is legal, every game is genuinely playable from start to finish, and every move is recorded in standard algebraic notation).
Important note on play quality: these games are fully legal but NOT master-level play — each move is chosen uniformly at random from the set of legal moves available… See the full description on the dataset page: https://huggingface.co/datasets/ziadatalabs/FreeChessGames25M.25-Mar-set1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so100_follower",
"total_episodes": 0,
"total_frames": 0,
"total_tasks": 0,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dillonlyr04/25-Mar-set1.25-Mar-set2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so100_follower",
"total_episodes": 2,
"total_frames": 407,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/dillonlyr04/25-Mar-set2.bankless_ROLLUP_Memestock_Mania__ETH_Spot_ETF_Countdown__25M_Hackers_Charged__Cryptos_Election_I
