datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eumetsat-cloudmask-iodceumetsat-cloudmask-rsseumetsat-cloudmask-0degcloudsen12
This dataset follows the TACO specification.
cloudsen12plus
Website: https://cloudsen12.github.io/
version: 1.1.2
The largest dataset of expert-labeled pixels for cloud and cloud shadow detection in Sentinel-2
CloudSEN12+ version 1.1.0 is a significant extension of the CloudSEN12 dataset, which doubles the number of
expert-reviewed labels, making it, by a large margin, the largest cloud detection dataset to
date for Sentinel-2. All labels from the previous version have… See the full description on the dataset page: https://huggingface.co/datasets/tacofoundation/cloudsen12.cloudCloudAVR
CloudAVR Point-Cloud Matrix Reasoning
CloudAVR is a procedurally generated benchmark for 3×3 matrix completion with
native colored 3D point clouds. Each puzzle hides the bottom-right cell and
asks the model to select the correct completion from four candidates.
CloudAVR comprises 108,000 puzzles in the full benchmark, evenly split
across six rule families, with a 104,400/1,800/1,800 train/validation/test
split. The Hub upload is still in progress; while files are being… See the full description on the dataset page: https://huggingface.co/datasets/ShowRison/CloudAVR.CloudSEN12-nolabel🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 NOLABEL
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-nolabel.OS-Omni-VM
OS-Omni VM
This dataset repository hosts prebuilt virtual machine images for OS-Omni
desktop-agent benchmark environments.
Contents
android/AndroidWorldAvd_baseline_20260503.zip: Android baseline AVD
artifact with the benchmark apps installed.
android/AndroidWorldAvd_baseline_20260503.zip.sha256: SHA256 checksum
for the Android AVD archive.
android/AndroidWorldAvd_baseline_20260503_package_manifest.txt: package
list captured from the exported Android emulator.… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/OS-Omni-VM.CloudSEN12-scribble🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 NOLABEL
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12 SCRIBBLE
A Benchmark Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-scribble.arabic-synthetic-scanned-booksCloudSEN12-high
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 HIGH-QUALITY
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-high.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.CloudAnoBench
CloudAnoBench
Paper: Towards Generalizable Context-aware Anomaly Detection: A Large-scale Benchmark in Cloud Environments
Project Page: https://jayzou3773.github.io/cloudanobench-agent/
CloudAnoBench is a large-scale benchmark for context-aware anomaly detection in cloud environments, jointly incorporating both metrics and logs to more faithfully reflect real-world conditions. It consists of 1,252 labeled cases spanning 28 anomalous scenarios and 16 deceptive normal scenarios… See the full description on the dataset page: https://huggingface.co/datasets/jayzou3773/CloudAnoBench.my-cloudpaste-storageVCBench
Overview
VCBench provides a standardized framework for evaluating vision-language models. This document outlines the procedures for both standard evaluation and GPT-assisted evaluation of your model's outputs.
1. Standard Evaluation
1.1 Output Format Requirements
Models must produce outputs in JSONL format with the following structure:
{"id": <int>, "pred_answer": "<answer_letter>"}
{"id": <int>, "pred_answer": "<answer_letter>"}
...
Example File… See the full description on the dataset page: https://huggingface.co/datasets/cloudcatcher2/VCBench.fengyun4A-cloud-detection-dataset
Fengyun-4A cloud detection dataset
The cloud detection dataset for Fengyun-4A (FY-4A) geostationary satellite supporting both physical and spatio-temporal data-driven algorithms.
Dataset description
Each folder contains cloud detection samples collected during the time period (UTC) indicated by the folder name, and includes the following three types of files:
yyyymmddHHMM_yyyymmddHHMM.hdf5 provides the index, latitude, longitude, land-surface type, 3×21×3×3… See the full description on the dataset page: https://huggingface.co/datasets/SimonSongHit/fengyun4A-cloud-detection-dataset.clouds-decoded-rain-check
Sentinel-2 Cloud Property Retrievals over the UK and India
Rendered per-scene quicklook layers of cloud optical and microphysical properties
retrieved from Sentinel-2 (MSI) Level-1C imagery over eight tiles (four in the UK
and four in India). These images power the interactive report at
asterisk-labs-clouds-decoded-rain-check.static.hf.space. Cloud properties are estimated with the clouds-decoded retrieval code, developed as part of the clouds decoded project funded by ARIA.… See the full description on the dataset page: https://huggingface.co/datasets/asterisk-labs/clouds-decoded-rain-check.512x1_ABI_CloudSatCloudSEN12Plus
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12+ is a significant extension of the CloudSEN12 dataset, which doubles the number of… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/CloudSEN12Plus.CloudAnoBench
CloudAnoBench
Paper: Towards Generalizable Context-aware Anomaly Detection: A Large-scale Benchmark in Cloud Environments
Project Page: https://jayzou3773.github.io/cloudanobench-agent/
CloudAnoBench is a large-scale benchmark for context-aware anomaly detection in cloud environments, jointly incorporating both metrics and logs to more faithfully reflect real-world conditions. It consists of 1,252 labeled cases spanning 28 anomalous scenarios and 16 deceptive normal scenarios… See the full description on the dataset page: https://huggingface.co/datasets/66aaa/CloudAnoBench.IRIS-CloudDeep
IRIS-CloudDeep
Ground-based long-wave infrared (LWIR) images of the night sky, with the binary ground-truth masks and clear/cloud labels behind Sommer, Kabalan and Brunet (2025), Atmos. Meas. Tech. 18, 2083–2101.
An uncooled FLIR Tau2 microbolometer (640×512, 17 μm pitch, 8–14 μm band, 9 Hz) recorded two night-time campaigns in early 2023 at Prades-le-Lez, France (43°41′51″ N, 3°51′53″ E). A 60 mm f/1.25 lens gives a narrow imaging area of 10.4° × 8.3°, about 58″ per pixel. The… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/IRIS-CloudDeep.PhyX
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Dataset for the paper "PhyX: Does Your Model Have the "Wits" for Physical Reasoning?".
For more details, please refer to the project page with dataset exploration and visualization tools: PhyX Project Page.
[🌐 Project Page] [📖 Paper] [🔧 Evaluation Code] [🌐 Blog (中文)]
🔔 News
[2026.02.15] 🎉 The Seed 2.0 technical report has been released and it outperforms GPT-5.2-High by 0.6% on PhyX, congratulations!… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/PhyX.data_wmcrossed_arm_point_clouds
Crossed Arm Point Clouds Dataset
This dataset contains 3D point cloud data captured from a LiDAR scanner for crossed arm classification in the context of robot magic trick performance.
Overview
This dataset was collected for training and evaluating the Crossed Arm Voxel Network (CAVN) architecture, a deep learning model designed for 3D point cloud classification in human-robot interaction magic performances. The data supports classification of human arm positions during… See the full description on the dataset page: https://huggingface.co/datasets/ahanjaya/crossed_arm_point_clouds.cloudops_tsf
Pushing the Limits of Pre-training for Time Series Forecasting in the CloudOps Domain
Paper | Code
Datasets accompanying the paper "Pushing the Limits of Pre-training for Time Series Forecasting in the CloudOps Domain".
Quick Start
pip install datasets==2.12.0 fsspec==2023.5.0
azure_vm_traces_2017
from datasets import load_dataset
dataset = load_dataset('Salesforce/cloudops_tsf', 'azure_vm_traces_2017')
print(dataset)
DatasetDict({
train_test: Dataset({… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/cloudops_tsf.ann-cc-3m
CC3M Image-Text Embeddings
images_part{1-3}.txt are text files with base64-encoded images.
texts.txt is a text file with captions for images.
images.{model_name}.fbin is a binary file with {model_name} image embeddings.
images.{model_name}.usearch is a binary file with a serialized USearch image index which contains images.{model_name}.fbin.
texts.{model_name}.fbin is a binary file with {model_name} text embeddings.
texts.{model_name}.usearch is a binary file with a serialized… See the full description on the dataset page: https://huggingface.co/datasets/unum-cloud/ann-cc-3m.USearchWiki
USearchWiki
Multi-model embedding dataset built on HuggingFace FineWiki, designed for approximate nearest neighbor (ANN) search benchmarking with USearch and other vector search engines.
The same Wikipedia corpus — chunked, cleaned, and enriched with graph metadata — is embedded by multiple models spanning dense BERT-like encoders, GPT-style decoder-based LLMs, and late-interaction ColBERT-style architectures.
Each model's embeddings ship with precomputed ground-truth k-nearest… See the full description on the dataset page: https://huggingface.co/datasets/unum-cloud/USearchWiki.nepali-corpus-compilecloudflare_imgBed_publicpsi0-g1-sneaker-205ep-v2-source
Psi0 G1 Sneaker-in-Box — 205 episodes (v2 canonical source)
⚠️ Do not use this dataset directly for training. This is the canonical immutable union of the v1 and v2 collections, kept as a source of truth for reproducibility. For v2 fine-tuning use psi0-g1-sneaker-199ep-v2; for held-out evaluation use psi0-g1-sneaker-6ep-v2-eval. Together these two derivatives reconstruct this canonical dataset exactly: 199 + 6 = 205.
205 teleoperated episodes of a Unitree G1 humanoid (with Inspire… See the full description on the dataset page: https://huggingface.co/datasets/cloudwalk-research/psi0-g1-sneaker-205ep-v2-source.
