datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iPhone360
iPhone360 Dataset (Simple Version)
iPhone360 is a benchmark dataset for 360° reconstruction of dynamic objects from monocular video, introduced in the paper:
4DGS360: 360° Gaussian Reconstruction of Dynamic Objects from a Single Video
Jae Won Jang, Yeonjin Chang, Wonsik Shin, Juhwan Cho, Nojun Kwak
Project Page · arXiv
Dataset Description
iPhone360 features real-world dynamic scenes captured with an iPhone, where test cameras are positioned at significantly… See the full description on the dataset page: https://huggingface.co/datasets/mipal/iPhone360.vlmn_iphonecf100_cotrain_magicsoup_no_insta_sub5
vlmn_iphonecf100_cotrain_magicsoup_no_insta_sub5
Description
VLN Navigation dataset with 100% of counterfactual iphone data, 10% of magicsoup no insta subsampled to 5 points.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 0.1
mateoguaman/iphone_every1_sub5: 1.0
mateoguaman/scand_every1_50pct_sub5: 0.1
mateoguaman/spot_every1_sub5: 0.1
mateoguaman/tartandrive_every1_100pct_sub5: 0.1… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphonecf100_cotrain_magicsoup_no_insta_sub5.egocentric-iphone-rgb-imu
Hub Egocentric: iPhone RGB+IMU
12 egocentric human-manipulation clips captured on iPhone, with nominal 30 Hz CoreMotion + ARKit IMU/attitude on the shared media timeline, delivered as CSV/Parquet/JSON sidecars plus a Foxglove MCAP recording and per-clip camera intrinsics.
Across this 12-clip sample, IMU row count is 0–4 boundary rows lower than decoded video frame count (≤0.035%).
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on iPhone 13… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-iphone-rgb-imu.vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5
Description
VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0
mateoguaman/iphone_stairs_ramps: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_sub5.iPhone-WallpapersIPHO2026
IPhO 2026 Curated Problems
This repository packages the official English problem, solution, and marking
materials for the LVI International Physics Olympiad (Bucaramanga, Colombia,
2026) as machine-readable, subquestion-level records.
Contents
Configuration
Rows
Description
all
41
All curated subquestions
theory
23
Theory papers T1–T3
experiment
18
Experimental paper E1
formalization_ready
29
Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.egocentric-iphone-rgb
Hub Egocentric: iPhone RGB
40 unique egocentric human-manipulation clips captured on iPhone, delivered as uniform re-encoded RGB video (H.264, 8-bit, Rec.709). Video only (no IMU, no depth).
Part of the Hub Egocentric Human Demonstrations Sample Set collection. Captured on iPhone 15, iPhone 15 Pro, iPhone 15 Pro Max, and iPhone 16 Pro Max. Egocentric, human-demonstration data (passive; no robot action stream). July 2026.
Dataset structure
One re-encoded H.264 MP4… See the full description on the dataset page: https://huggingface.co/datasets/Hubdata/egocentric-iphone-rgb.vlmn_iphonecf100_cotrain_magicsoup_no_insta_rdp
vlmn_iphonecf100_cotrain_magicsoup_no_insta_rdp
Description
VLN Navigation dataset with 100% of counterfactual iphone data, 10% of magicsoup no insta subsampled with RDP.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_rdp: 0.1
mateoguaman/iphone_every1_100pct_rdp: 1.0
mateoguaman/scand_every1_50pct_rdp: 0.1
mateoguaman/spot_every1_100pct_rdp: 0.1
mateoguaman/tartandrive_every1_100pct_rdp: 0.1… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphonecf100_cotrain_magicsoup_no_insta_rdp.coi-rl-data
COI RL Dataset
A reinforcement learning dataset for training vision-language models to use visual tools (zoom, contrast adjust, sharpen, etc.) when answering questions about images. The dataset is designed for RL-based post-training where models learn when and which visual manipulation tools to invoke before generating an answer.
Dataset Summary
Train samples: 8,853
Test samples: 9
Images: 8,400 (1.64 GB)
Question types: multiple-choice (2,072) and open-ended (6… See the full description on the dataset page: https://huggingface.co/datasets/iPhone38/coi-rl-data.vlmn_iphone100_tartandrive100_scand50_coda25_spot100_insta360100_sub5
vlmn_iphone100_tartandrive100_scand50_coda25_spot100_insta360100_sub5
Description
VLN Navigation dataset with 100% of iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, 100% of in-domain spot data, and 100% of insta360 data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphone100_tartandrive100_scand50_coda25_spot100_insta360100_sub5.iphone_stairs_ramps
iphone_stairs_ramps
Description
Processed entire iphone_stairs_ramps with filter_every_nth=1, 100% of data, and num_subsampled_points=5
Processing Parameters
mateoguaman/iphone_chin:
exclude_outliers_pct: 0
filter_by_curvature: false
filter_every_nth: 1
horizon:
1000: 1.0
num_subsampled_points: 5
mateoguaman/iphone_hip:
exclude_outliers_pct: 0
filter_by_curvature: false
filter_every_nth: 1
horizon:1000: 1.0
num_subsampled_points: 5… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/iphone_stairs_ramps.iphone5_vlm_imgIChO-IPhO-RL-v2-formated
ICHO-IPH0 Dataset
The ICHO-IPH0 dataset was crawled from various websites containing challenges similar to the Olympiad for chemistry and physics.
Preprocessing
First, we crawl the problems in PDF format.
For each PDF file, we use gemini-flash-2.0 to extract the (problem, solution) pairs for each question.
To preserve context, we include all previous questions and solutions along with the current question.
We filter out problems that contain figures, images, URLs, etc.
iphone5_vlm_img_livesessioniphone5_vlm_imgvlmn_iphonecf100_tartandrive100_scand50_coda25_spot100_rdp
vlmn_iphonecf100_tartandrive100_scand50_coda25_spot100_rdp
Description
VLN Navigation dataset with 100% of counterfactual iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_rdp: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphonecf100_tartandrive100_scand50_coda25_spot100_rdp.iphone2dslr_flowerThis dataset is part of the CycleGAN datasets, originally hosted here: https://people.eecs.berkeley.edu/~taesung_park/CycleGAN/datasets/
Citation
@article{DBLP:journals/corr/ZhuPIE17,
author = {Jun{-}Yan Zhu and
Taesung Park and
Phillip Isola and
Alexei A. Efros},
title = {Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial
Networks},
journal = {CoRR},
volume = {abs/1703.10593},
year… See the full description on the dataset page: https://huggingface.co/datasets/huggan/iphone2dslr_flower.iphone5_vlm_livesessionradiology_audio_3_iphone_laptop_666_samplesPhy-RL-Non-IPhO-Non-HiPHOiPhone14TweetsApprox 144K tweets about iPhone 14
amazon-iphone8plus-review
Dataset Description
Source
This dataset is derived from the Amazon Reviews 2023 dataset released by the McAuley Lab.
Subset
The dataset contains Amazon Electronics reviews for the iPhone 8 Plus, identified by the parent ASIN B089SRK8VQ within the Amazon product family.
Size
Number of records: 4,085 reviews
Ratings
The dataset includes reviews with the following star ratings:
1, 2, 3, 4, and 5 stars
Task
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/manmanluo/amazon-iphone8plus-review.radiology_audio_3_iphoneiphone5_vlmiphone5_vlm_imgiphone5_vlm_imgiphone_text_datavlmn_iphonecf100_tartandrive100_scand50_coda25_spot100_sub5
vlmn_iphonecf100_tartandrive100_scand50_coda25_spot100_sub5
Description
VLN Navigation dataset with 100% of counterfactual iphone data, 100% of tartandrive data, 50% of scand data, 25% of coda data, and 100% of in-domain spot data. Whenever daatsets aren't 100%, they are ranked by curvature and output of length 5.
Processing Parameters
{}
Dataset Configuration
Train dataset:
mixer: mateoguaman/coda_every1_25pct_sub5: 1.0… See the full description on the dataset page: https://huggingface.co/datasets/mateoguaman/vlmn_iphonecf100_tartandrive100_scand50_coda25_spot100_sub5.iphone5_images_with_featuresiphone16-dataset
