datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OS-Omni-VM
OS-Omni VM
This dataset repository hosts prebuilt virtual machine images for OS-Omni
desktop-agent benchmark environments.
Contents
android/AndroidWorldAvd_baseline_20260503.zip: Android baseline AVD
artifact with the benchmark apps installed.
android/AndroidWorldAvd_baseline_20260503.zip.sha256: SHA256 checksum
for the Android AVD archive.
android/AndroidWorldAvd_baseline_20260503_package_manifest.txt: package
list captured from the exported Android emulator.… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/OS-Omni-VM.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.VCBench
Overview
VCBench provides a standardized framework for evaluating vision-language models. This document outlines the procedures for both standard evaluation and GPT-assisted evaluation of your model's outputs.
1. Standard Evaluation
1.1 Output Format Requirements
Models must produce outputs in JSONL format with the following structure:
{"id": <int>, "pred_answer": "<answer_letter>"}
{"id": <int>, "pred_answer": "<answer_letter>"}
...
Example File… See the full description on the dataset page: https://huggingface.co/datasets/cloudcatcher2/VCBench.CloudSEN12Plus
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12+ is a significant extension of the CloudSEN12 dataset, which doubles the number of… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/CloudSEN12Plus.IRIS-CloudDeep
IRIS-CloudDeep
Ground-based long-wave infrared (LWIR) images of the night sky, with the binary ground-truth masks and clear/cloud labels behind Sommer, Kabalan and Brunet (2025), Atmos. Meas. Tech. 18, 2083–2101.
An uncooled FLIR Tau2 microbolometer (640×512, 17 μm pitch, 8–14 μm band, 9 Hz) recorded two night-time campaigns in early 2023 at Prades-le-Lez, France (43°41′51″ N, 3°51′53″ E). A 60 mm f/1.25 lens gives a narrow imaging area of 10.4° × 8.3°, about 58″ per pixel. The… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/IRIS-CloudDeep.PhyX
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Dataset for the paper "PhyX: Does Your Model Have the "Wits" for Physical Reasoning?".
For more details, please refer to the project page with dataset exploration and visualization tools: PhyX Project Page.
[🌐 Project Page] [📖 Paper] [🔧 Evaluation Code] [🌐 Blog (中文)]
🔔 News
[2026.02.15] 🎉 The Seed 2.0 technical report has been released and it outperforms GPT-5.2-High by 0.6% on PhyX, congratulations!… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/PhyX.cloudflare_imgBed_publiccloudtunesann-unsplash-25k
25K Unsplash Images for Search
This is a derivative work based on two existing datasets.
images.csv metadata from Unsplash, sorted and converted to CSV.
images/ in 250x250 resolution by kaggle/@jettchentt.
images.fbin is a binary file with UForm image embeddings.
images.usearch is a binary file with a serialized USearch index.
The original images.tsv from Unsplash has been filtered to avoid missing images.
The embeddings and the index can be reconstructed with the main.py script.… See the full description on the dataset page: https://huggingface.co/datasets/unum-cloud/ann-unsplash-25k.cloud-stereoCloud-Stereo Dataset (BMVC 2025)
Project Page: https://cloud-stereo.jacob-lin.com/
Lora_Cloud_Dataset_Train
Safety Inspector V2 LoRA Training Dataset & Hyperparameter Specification
This dataset repository contains the offline warm-up SFT dataset and standardized LoRA training configuration for training the Vision-Language Model (VLM) Safety Inspector on the 50 tabletop manipulation scenes (Split into 45 Train + 5 Validation).
1. Dataset Overview
Source Scenes: 50 Tabletop Scenes (45 Train, 5 Val, 0 Test)
Task Levels: single_step (1 primitive), safety_two_step (2… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Train.synthetic-cloud-removal
Dataset Card for "synthetic-cloud-removal"
More Information needed
CCAiM-CloudsDataset
CCAiM CloudsDataset
Description
This dataset contains photographs of clouds collected for the CCAiM project, a model for cloud classification. It includes various types of clouds captured from the ground and can be used for training and testing computer vision models.
Dataset Structure
Cloud images in JPEG/PNG format
Optional metadata: cloud type, date, location
Dataset Statistics
Total number of images: 916
Cloud Type
Number… See the full description on the dataset page: https://huggingface.co/datasets/serbekun/CCAiM-CloudsDataset.ocr-document-processing-evalcloudflare-imgbedCloudFlare-ImgBedCloudBench
CloudBench: A Benchmark Dataset for Cloud Image Retrieval
Dataset Description
CloudBench is a benchmark dataset for evaluating image retrieval systems in the domain of Atmospheric Science specifically focused on clouds. The dataset consists of natural language queries paired with images, along with binary relevance labels indicating whether each image is relevant to the query. The dataset is designed to test retrieval systems' ability to find relevant images based on… See the full description on the dataset page: https://huggingface.co/datasets/sagecontinuum/CloudBench.cloudflare-imgbedCloudFlare-ImgBedcloudflare-imgbedCloudFlare-ImgBedCloudSEN12Plus
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12+ is a significant extension of the CloudSEN12 dataset, which doubles the number of… See the full description on the dataset page: https://huggingface.co/datasets/MohamedAyman456/CloudSEN12Plus.os-omni-benchmark
OS-Omni Benchmark
OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks.
Contents
data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation.
data/tasks.jsonl: JSON Lines copy of the same task index.
metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/os-omni-benchmark.CloudFlare-ImgBedComfyUI-Prompt-CloudDB
ComfyUI-Prompt-CloudDB
⚠️ 免责声明 (Disclaimer)
本仓库收录的所有 Prompt 预设、画师风格及其生成的预览图,仅供 AI 绘画技术学习、风格研究与插件功能测试使用,绝无任何不良引导或商业盈利目的。
由于部分开源 AI 绘画模型本身的训练集偏好,在使用通用或特定画师 Prompt 跑图时,极易生成包含性感元素(如泳装、内衣、紧身衣等)的预览图。本仓库虽已尽力人工筛选,但难以完全杜绝此类擦边内容。使用者须自行承担在公共场合阅览或使用本图库可能引发的风险。
1. 仓库简介
本仓库是 ComfyUI-Prompt-Manager 插件的官方云端公共词库与图库数据中心。
通过接入 jsDelivr CDN,本仓库为插件提供了“开箱即用”的在线画廊功能。所有安装了该插件的用户,可以在无需本地配置的情况下,直接在 ComfyUI 界面中浏览、检索并调用由社区共同维护的高质量 Prompt 预设、画师风格及其对应的生成预览图。
2. 如何参与内容贡献… See the full description on the dataset page: https://huggingface.co/datasets/FRuoL/ComfyUI-Prompt-CloudDB.38-cloud-dataset
Dataset Card for "38-cloud-train-only-v2"
More Information needed
south-africa-crop-type-clouds
Dataset Card for South Africa Crop Type Clouds
This dataset contains the cloud masks generated and used for the paper KAN You See It? KANs and Sentinel for Effective and Explainable Crop Field Segmentation.
Curated by: Daniele Rege Cambrin
License: OpenRAIL
Uses
The dataset will provide a quality assessment for Sentinel-2 images of the South Africa Crop Type dataset.
Since MSI is ineffective through clouds, it was used to exclude samples that contain a large portion… See the full description on the dataset page: https://huggingface.co/datasets/DarthReca/south-africa-crop-type-clouds.38-cloud-train-only-v3-with-NIR
Dataset Card for "38-cloud-train-only-v3-with-nir"
More Information needed
ThaiIDCardSynt
Dataset Details
Dataset Description
Curated by: Matichon Maneegard
Shared by [optional]: Matichon Maneegard
Language(s) (NLP): image-to-text
License: apache-2.0
Dataset Sources [optional]
The dataset was entirely synthetic. It does not contain real information or pertain to any specific person.
Uses
Direct Use
Using for tranning OCR or Multimodal.
Dataset Structure
This dataset contains 98 x 6 = 588 samples, and the… See the full description on the dataset page: https://huggingface.co/datasets/Float16-cloud/ThaiIDCardSynt.yandex-cloud-dataset-resize
