datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IPHO2026
IPhO 2026 Curated Problems
This repository packages the official English problem, solution, and marking
materials for the LVI International Physics Olympiad (Bucaramanga, Colombia,
2026) as machine-readable, subquestion-level records.
Contents
Configuration
Rows
Description
all
41
All curated subquestions
theory
23
Theory papers T1–T3
experiment
18
Experimental paper E1
formalization_ready
29
Subset selected for theorem formalization… See the full description on the dataset page: https://huggingface.co/datasets/humanfia-lab/IPHO2026.coi-rl-data
COI RL Dataset
A reinforcement learning dataset for training vision-language models to use visual tools (zoom, contrast adjust, sharpen, etc.) when answering questions about images. The dataset is designed for RL-based post-training where models learn when and which visual manipulation tools to invoke before generating an answer.
Dataset Summary
Train samples: 8,853
Test samples: 9
Images: 8,400 (1.64 GB)
Question types: multiple-choice (2,072) and open-ended (6… See the full description on the dataset page: https://huggingface.co/datasets/iPhone38/coi-rl-data.
