datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CloudAVR
CloudAVR Point-Cloud Matrix Reasoning
CloudAVR is a procedurally generated benchmark for 3×3 matrix completion with
native colored 3D point clouds. Each puzzle hides the bottom-right cell and
asks the model to select the correct completion from four candidates.
CloudAVR comprises 108,000 puzzles in the full benchmark, evenly split
across six rule families, with a 104,400/1,800/1,800 train/validation/test
split. The Hub upload is still in progress; while files are being… See the full description on the dataset page: https://huggingface.co/datasets/ShowRison/CloudAVR.CloudSEN12-nolabel🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 NOLABEL
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-nolabel.CloudSEN12-scribble🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 NOLABEL
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12 SCRIBBLE
A Benchmark Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-scribble.arabic-synthetic-scanned-booksCloudSEN12-high
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 HIGH-QUALITY
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-high.Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.fengyun4A-cloud-detection-dataset
Fengyun-4A cloud detection dataset
The cloud detection dataset for Fengyun-4A (FY-4A) geostationary satellite supporting both physical and spatio-temporal data-driven algorithms.
Dataset description
Each folder contains cloud detection samples collected during the time period (UTC) indicated by the folder name, and includes the following three types of files:
yyyymmddHHMM_yyyymmddHHMM.hdf5 provides the index, latitude, longitude, land-surface type, 3×21×3×3… See the full description on the dataset page: https://huggingface.co/datasets/SimonSongHit/fengyun4A-cloud-detection-dataset.CloudSEN12Plus
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12+ is a significant extension of the CloudSEN12 dataset, which doubles the number of… See the full description on the dataset page: https://huggingface.co/datasets/isp-uv-es/CloudSEN12Plus.PhyX
PhyX: Does Your Model Have the "Wits" for Physical Reasoning?
Dataset for the paper "PhyX: Does Your Model Have the "Wits" for Physical Reasoning?".
For more details, please refer to the project page with dataset exploration and visualization tools: PhyX Project Page.
[🌐 Project Page] [📖 Paper] [🔧 Evaluation Code] [🌐 Blog (中文)]
🔔 News
[2026.02.15] 🎉 The Seed 2.0 technical report has been released and it outperforms GPT-5.2-High by 0.6% on PhyX, congratulations!… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/PhyX.data_wmann-cc-3m
CC3M Image-Text Embeddings
images_part{1-3}.txt are text files with base64-encoded images.
texts.txt is a text file with captions for images.
images.{model_name}.fbin is a binary file with {model_name} image embeddings.
images.{model_name}.usearch is a binary file with a serialized USearch image index which contains images.{model_name}.fbin.
texts.{model_name}.fbin is a binary file with {model_name} text embeddings.
texts.{model_name}.usearch is a binary file with a serialized… See the full description on the dataset page: https://huggingface.co/datasets/unum-cloud/ann-cc-3m.nepali-corpus-compilelongrca-bench
LongRCA Bench
A benchmark for diagnosing responsible roles and root causes in long-horizon agent failures.
1,140 observed, non-injected failed trajectories from SWE-bench Pro, Terminal-Bench 2, TravelPlanner, VitaBench, and WebArena Verified.
Human annotations identifying the responsible role, earliest decisive root-cause step, and supporting rationale.
LongRCA-Mini: a fixed 200-trajectory subset, randomly sampled without replacement with 40 from each benchmark, preserving the… See the full description on the dataset page: https://huggingface.co/datasets/CLoud5-real/longrca-bench.sangraha_nepalisynthesized-cloud-optimization-recommendations
Synthesized Cloud-Optimization Recommendations
18 scenarios that pair cloud telemetry with a hand-crafted optimization
recommendation. Use them to train models or to evaluate AI agents.
Summary
Each scenario has multi-tier telemetry, a Terraform file describing the
deployed infrastructure, and a gold-standard recommendation.
The dataset is built around a simple input-output mapping. The input is
telemetry plus the infrastructure. The output is an optimization… See the full description on the dataset page: https://huggingface.co/datasets/ameau01/synthesized-cloud-optimization-recommendations.cloud-stereoCloud-Stereo Dataset (BMVC 2025)
Project Page: https://cloud-stereo.jacob-lin.com/
synthetic-cloud-removal
Dataset Card for "synthetic-cloud-removal"
More Information needed
security-paper-datasets
Dataset Card for "security-paper-datasets"
More Information needed
ann-arxiv-2m
2M Title-Abstract Arxiv Pairs
title_abstract.tsv data from Cornell University Arxiv Dataset, preprocessed and coverted to TSV.
title.e5-base-v2.fbin is a binary file with e5-base-v2 title embeddings.
abstract.e5-base-v2.fbin is a binary file with e5-base-v2 abstract embeddings.
opensre
OpenSRE / OpenRCA dataset
Root cause analysis benchmark data: queries, incident records, telemetry (metrics, logs, traces), and query_alerts (per-row JSON derived from each query.csv).
Telemetry CSVs use different schemas by file type; load them by path (they are not merged into the Hub subset configs above).
Original archives are also described here: Google Drive.
Regenerating query_alerts
python3 scripts/query_csv_to_alert_json.py
CloudBench
CloudBench: A Benchmark Dataset for Cloud Image Retrieval
Dataset Description
CloudBench is a benchmark dataset for evaluating image retrieval systems in the domain of Atmospheric Science specifically focused on clouds. The dataset consists of natural language queries paired with images, along with binary relevance labels indicating whether each image is relevant to the query. The dataset is designed to test retrieval systems' ability to find relevant images based on… See the full description on the dataset page: https://huggingface.co/datasets/sagecontinuum/CloudBench.devanagari_pretrainann-codesearch-4m
Cleaning
Unlike the original dataset, the func_code_string column was updated to remove any comments and keep just the code.
The original version can still be found in the whole_func_string.
import re
def remove_comments_docstrings(code, language):
if language == 'python':
# Remove docstrings
code = re.sub(r'"""(.*?)"""', '', code, flags=re.DOTALL)
code = re.sub(r"'''(.*?)'''", '', code, flags=re.DOTALL)
# Remove comments
code =… See the full description on the dataset page: https://huggingface.co/datasets/unum-cloud/ann-codesearch-4m.finewiki-en-1mCloudSEN12Plus
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12+ is a significant extension of the CloudSEN12 dataset, which doubles the number of… See the full description on the dataset page: https://huggingface.co/datasets/MohamedAyman456/CloudSEN12Plus.nepali-news-corpusmulti-domain-cloudflare-observability
Multi-Domain Cloudflare Web Traffic, Performance and Security Observability Dataset
This dataset contains multi-domain Cloudflare analytics exported into analysis-ready Parquet tables. It combines HTTP request aggregates, hourly traffic trends, path and referrer dimensions, country/device/browser breakdowns, DNS analytics, cache behavior, Web Vitals/RUM performance signals, redirect and ruleset metadata, and available firewall/security aggregates across 200 websites.
The… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/multi-domain-cloudflare-observability.africa-cloud-cover-bias
Cloud Cover and Structural Observation Gaps in African Agricultural EO (six-zone dataset)
Supporting data for the paper "Cloud Cover and Structural Observation Gaps in
African Agricultural Earth Observation: Evidence from Six Agroecological Zones"
(Olaoye Anthony Somide, CropSense AI Research / CipherSense AI; Zenodo, doi:10.5281/zenodo.22642336; also EarthArXiv, doi:10.31223/X5J503).
Weekly usable Sentinel-2 optical observation frequency over cropland for six administrative… See the full description on the dataset page: https://huggingface.co/datasets/CipherSenseAI/africa-cloud-cover-bias.os-omni-benchmark
OS-Omni Benchmark
OS-Omni is a cross-platform benchmark for evaluating agents that operate graphical operating-system environments. This dataset repository contains the static benchmark task definitions and supporting assets used to configure and evaluate OS-Omni tasks.
Contents
data/tasks.parquet: tabular task index for Hugging Face Dataset Viewer and Croissant generation.
data/tasks.jsonl: JSON Lines copy of the same task index.
metadata/tasks.parquet: duplicate task… See the full description on the dataset page: https://huggingface.co/datasets/Cloudriver/os-omni-benchmark.nepali-roman-pretrain
