datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OceanDepths
OceanDepths GeoTIFF Raster and Aligned ARGO Dataset
This dataset package contains the model-ready Ocean variables (ARGO submarine data, sea surface height, sea surface temperature and
salinity, as well as GLORYS reanalysis information for 50 depth levels. The ARGO data has been projected onto the GLORYS grid in order
to build a ML-ready dataset. The intention is that users can create tensors easily for CV-inspired ML approaches to ocean-variable
reconstruction. While… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/OceanDepths.retail-products-philippinesInpaintCOCO
InpaintCOCO - Fine-grained multimodal concept understanding (for color, size, and COCO objects)
Dataset Summary
A data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object.
Many multimodal tasks, such as Vision-Language Retrieval and Visual Question Answering, present results in terms of overall performance.
Unfortunately, this approach overlooks more nuanced concepts, leaving us unaware… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/InpaintCOCO.synthetic-ocr-en-det-rec-120k
Synthetic English OCR Detection and Recognition 240K
📌 Current dataset size: 240,000 paired OCR samples
The current v2.0 release contains exactly 240,000 detector images and
240,000 matching recognition crops.
Each sample ID corresponds to:
one full image for text detection;
one cropped text image for text recognition;
one detector JSONL record;
one recognizer JSONL record.
Therefore, the dataset contains 240,000 aligned OCR pairs and
480,000 JPEG files in… See the full description on the dataset page: https://huggingface.co/datasets/Phitran21/synthetic-ocr-en-det-rec-120k.sole_training_data
This is the training dataset for SOLE-R1-8B
SOLE-R1-8B is a video-language reward reasoning model for robotics. It is designed to estimate task progress from robot video frames and a natural-language task description, producing both per-timestep reasoning traces and scalar progress predictions that can be used as rewards for online robot reinforcement learning.
This dataset accompanies the paper “SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL” by Philip… See the full description on the dataset page: https://huggingface.co/datasets/Philip-MIT/sole_training_data.coco2017
coco2017
Image-text pairs from MS COCO2017.
Data origin
Data originates from cocodataset.org
While coco-karpathy uses a dense format (with several sentences and sendids per row), coco-karpathy-long uses a long format with one sentence (aka caption) and sendid per row. coco-karpathy-long uses the first five sentences and therefore is five times as long as coco-karpathy.
phiyodr/coco2017: One row corresponds one image with several sentences.
phiyodr/coco2017-long: One row… See the full description on the dataset page: https://huggingface.co/datasets/phiyodr/coco2017.phisat2-s2-lightglue-triplets
PhiSat-2 / Sentinel-2 LightGlue Triplets
This dataset contains finalized strict triplet outputs generated from the local
PhiSat-2/Sentinel-2 LightGlue pipeline. Each accepted patch includes real
PhiSat-2, Sentinel-2 L1C, simulated PhiSat-2, OmniCloudMask, ESA WorldCover, and
Koppen-Geiger metadata.
Quality policy: fail closed. Patches are accepted only when registration,
geometry, nodata, PhiSat-2 cloud, WorldCover, and Koppen gates pass.
Current upload:
finalized pairs: 1… See the full description on the dataset page: https://huggingface.co/datasets/ESA-philab/phisat2-s2-lightglue-triplets.phishing-website-screenshots
Phishing Website Screenshots
A dataset of 8,370 full-page website screenshots labelled as legitimate or phishing, intended for training and evaluating visual phishing-detection models.
Contents
Label
label
Images
legitimate
0
7,924
phishing
1
446
Total
8,370
Screenshots were captured at a desktop viewport (1920×1080) as PNG images.
Structure
legitimate/<brand>/<page>.png
phishing/<source>/<page>.png
metadata.csv
metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/shresthsamyak/phishing-website-screenshots.cilp_assessment_subset
Dataset Card for cilp_assessment_subset
Created with this Jupyter Notebook.
This is a FiftyOne dataset with 750 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("philippkolbe/cilp_assessment_subset")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/philippkolbe/cilp_assessment_subset.pdf-samplesgalahad-deconf-verb
Deconfounded verb set (RoboCasa)
Part of the Galahad release · Project page · Code + battery + generator
Lift-versus-slide with the same object under both verbs, role-balanced, so the verb is the only cue. Behind the verb-selectivity result (+44.3pp, told-slide→lift 2.0%).
LeRobot v3 format. Generated by the deconfounding generators in the release repo; the generator produces the exam (the battery) and the medicine (the training set) from the same code.
galahad-deconf-category
Deconfounded category set (RoboCasa)
Part of the Galahad release · Project page · Code + battery + generator
500 episodes, 100 per category, category ⟂ instance ⟂ co-appearance ⟂ position: per-episode mesh resampling breaks single-instance memorization, random distractor subsets break co-appearance.
LeRobot v2.1 format. Generated by the deconfounding generators in the release repo; the generator produces the exam (the battery) and the medicine (the training set) from the same… See the full description on the dataset page: https://huggingface.co/datasets/phi-monster/galahad-deconf-category.amazon-product-descriptions-vlm
Amazon Multimodal Product dataset
This is a modfied and slim verison of bprateek/amazon_product_description helpful to get started training multimodal LLMs.
The description field was generated used Gemini Flash.
modified-swiss-dwellings-enriched
Modified Swiss Dwellings (MSD), enriched
Floor plans of medium-to-large multi-apartment building complexes (ECCV 2024 benchmark
MSD), each linking three modalities of one floor plan:
image, geometry, and access graph.
1. Why this is here & what was done
Hosted on Hugging Face for reach and one-line loading by the ML community. The Swiss
Dwellings (SD) license (CC BY 4.0) permits redistribution with attribution — so this
also enriches the public MSD release, which… See the full description on the dataset page: https://huggingface.co/datasets/philippds/modified-swiss-dwellings-enriched.CLT-IML-Dataset
CLT-IML Tokamak MHD Simulation Database
Access and use. This database is source-available for academic
communication, inspection, and reproducibility assessment. It is not open
data. Copyright (c) 2026 Zhejiang University. All rights reserved. Any use
requires prior written permission from Zhejiang University or its duly
authorized representative. See Terms of access and use.
Dataset summary
This dataset contains tabular scalar responses, sampled two-dimensional… See the full description on the dataset page: https://huggingface.co/datasets/Philaus/CLT-IML-Dataset.PhillMagine120phisat2-ortho-referencere-recap
re-recap
re-recap is an open vision-language dataset pairing live web image URLs with dual captions: the original baseline recaption from Recap-DataComp-1B and an expanded, highly detailed visual description generated directly from image pixels using Qwen3.8-27B.
The dataset contains 9,717,471 verified image-text pairs partitioned into two distinct subsets based on caption length and prompt structure.
Overview and Dataset Subsets
To serve different modeling… See the full description on the dataset page: https://huggingface.co/datasets/phi-9/re-recap.esa_philab_embed2heights
THIS IS A COPY OF THE ORIGINAL DATASET TO MAKE IT MORE ACCESSIBLE
SOURCE: https://www.eotdl.com/datasets/embed2heights?ref=philabchallenges-cms.earthpulse.es
embed2heights Challenge - Reaching New Heights with GeoFM Embeddings
Overview
The objective of the embed2heights challenge is to develop a multi-task method that uses Geospatial Foundation Model embeddings to map land cover and estimate heights at scale. Participants are asked to combine multiple embedding… See the full description on the dataset page: https://huggingface.co/datasets/troni21/esa_philab_embed2heights.celeba-hq-1.5k
Dataset Card for "celeba-hq-1.5k"
More Information needed
phishing-website-screenshots
Phishing Website Screenshots
A dataset of 8,370 full-page website screenshots labelled as legitimate or phishing, intended for training and evaluating visual phishing-detection models.
Contents
Label
label
Images
legitimate
0
7,924
phishing
1
446
Total
8,370
Screenshots were captured at a desktop viewport (1920×1080) as PNG images.
Structure
legitimate/<brand>/<page>.png
phishing/<source>/<page>.png
metadata.csv
metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/gpm123/phishing-website-screenshots.hle_ja_phi4https://huggingface.co/datasets/cais/hle
Humanity's Last Examのquestionをphi4で日本語訳したものです。
どのような質問があるかの中身の確認用です。きれいに出力されているかはあまり確認していません。
中身の確認用ページ(初回ロード遅い)https://if001.github.io/hle_sample/
いくつか出力が途切れているものがあります。
Input IDs of length 7215 > the model's max sequence length of 4096.Input IDs of length 7348 > the model's max sequence length of 4096.Input IDs of length 9698 > the model's max sequence length of 4096.Input IDs of length 10144 > the model's max sequence length of 4096.Input IDs of… See the full description on the dataset page: https://huggingface.co/datasets/if001/hle_ja_phi4.flickr30k_longCoral_snake_mimicry_FUNED_photos
Dataset Card for Coral Snake Mimicry specimens from FUNED
This dataset comprises images of snake specimens belonging to species within the coral snake mimicry complex from the Fundacao Ezequiel Dias (FUNED) in Belo Horizonte, MG, Brasil. Images of museum specimens were taken with a Nikon camera and a standardized color palette, but white balance has not been corrected on these images. This dataset is a small representation of a larger dataset collected by Andressa Viol for her PhD… See the full description on the dataset page: https://huggingface.co/datasets/philodryas/Coral_snake_mimicry_FUNED_photos.mtg_cards-2025-04-04MTG cards labeled with similarity score.
Can be used for sentence and/or image similarity tasks.
I hearby declare that I don't own any rights to the data, I only added annotations.
Annotation
Thebdata was automatically annotated, the code will be added here soon.
Interesting samples:
2df7b947-bdb2-4204-8eb0-92fe66411613_8c5bd289-6069-4d99-83cd-d1dda24bc224
galahad-deconf-object
Deconfounded object-identity set (LIBERO)
Part of the Galahad release · Project page · Code + battery + generator
900 episodes, 90 per object across all 10 LIBERO objects, perfect role balance: every object appears as the target and as a distractor across randomized positions, so the instruction is the only predictor of the target. The training set behind the object-substitution cure (base 13.5% → 94.5%).
LeRobot v3 format. Generated by the deconfounding generators in the… See the full description on the dataset page: https://huggingface.co/datasets/phi-monster/galahad-deconf-object.PhilEOBench-road_density_regression
Simulated PhiSat Bench Dataset - Roads
This dataset comprises simulated PhiSat2 data derived from Sentinel-2, tailored for pixel-wise regression tasks aimed at estimating road coverage.
Dataset Overview
Each sample in the dataset includes a single-channel label.
The labels are stored as floating-point values that represent the estimated percentage of roads area within each pixel.
For a pixel with a 10-meter resolution (representing 100 square meters), the label… See the full description on the dataset page: https://huggingface.co/datasets/ESA-PhiLab-Edge/PhilEOBench-road_density_regression.celeba-hq-15k
Dataset Card for "celeba-hq-15k"
More Information needed
Web_page_Phishing
Web Page Phishing Detection — EDA
Student: Yonatane Ben-Aroch,Reichman UniversityDataset: Web Page Phishing Detection — KaggleDate: April 2026
Your browser does not support the video tag.
Overview
Web Page Phishing is when someone creates a fake website that looks real to steal your password or personal information. This dataset contains 11,430 URLs that were labeled as either legitimate or phishing, with 87 features extracted from each… See the full description on the dataset page: https://huggingface.co/datasets/yonatane22-bh/Web_page_Phishing.b300-lora-review
