datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CloudSEN12-nolabel🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 NOLABEL
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12 is a LARGE dataset (~1 TB) for cloud semantic… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-nolabel.CloudSEN12-scribble🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 NOLABEL
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12 SCRIBBLE
A Benchmark Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-scribble.CloudSEN12-high
🚨 New Dataset Version Released!
We are excited to announce the release of Version [1.1] of our dataset!
This update includes:
[L2A & L1C support].
[Temporal support].
[Check the data without downloading (Cloud-optimized properties)].
📥 Go to: https://huggingface.co/datasets/tacofoundation/cloudsen12 and follow the instructions in colab
CloudSEN12 HIGH-QUALITY
A Benchmark Dataset for Cloud Semantic Understanding
CloudSEN12… See the full description on the dataset page: https://huggingface.co/datasets/csaybar/CloudSEN12-high.opensre
OpenSRE / OpenRCA dataset
Root cause analysis benchmark data: queries, incident records, telemetry (metrics, logs, traces), and query_alerts (per-row JSON derived from each query.csv).
Telemetry CSVs use different schemas by file type; load them by path (they are not merged into the Hub subset configs above).
Original archives are also described here: Google Drive.
Regenerating query_alerts
python3 scripts/query_csv_to_alert_json.py
ann-arxiv-2m
2M Title-Abstract Arxiv Pairs
title_abstract.tsv data from Cornell University Arxiv Dataset, preprocessed and coverted to TSV.
title.e5-base-v2.fbin is a binary file with e5-base-v2 title embeddings.
abstract.e5-base-v2.fbin is a binary file with e5-base-v2 abstract embeddings.
gcp-cloud-billing-costThaiIDCardSynt
Dataset Details
Dataset Description
Curated by: Matichon Maneegard
Shared by [optional]: Matichon Maneegard
Language(s) (NLP): image-to-text
License: apache-2.0
Dataset Sources [optional]
The dataset was entirely synthetic. It does not contain real information or pertain to any specific person.
Uses
Direct Use
Using for tranning OCR or Multimodal.
Dataset Structure
This dataset contains 98 x 6 = 588 samples, and the… See the full description on the dataset page: https://huggingface.co/datasets/Float16-cloud/ThaiIDCardSynt.CloudTimeSeriesData
Intro
The data organization follows TFB format: https://github.com/decisionintelligence/TFB.
TFB data format
TFB stores time series in a format of three column long tables, which we will introduce below:
Format Introduction
First column: date (the exact column name is required, the same applies below.)
The columns stores the time information in the time series, which can be in either of the following formats:
Timestamps in string, datetime, or other types… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/CloudTimeSeriesData.cloud-gpu-price-index
Cloud GPU Price Index
Canonical source: https://gpueconomy.com/price-index. That page is
recomputed every hour; this record is a dated snapshot of it, version
2026-09-19, built from data updated 2026-09-19T09:26:26.511913+00:00. When you cite,
cite GPU Economy and link the page; the snapshot is here so that a number you
used keeps existing exactly as you used it.
The index is the weekly median publicly listed on-demand price of one NVIDIA
H100 SXM GPU-hour across the cloud GPU… See the full description on the dataset page: https://huggingface.co/datasets/gpueconomy/cloud-gpu-price-index.recentprofit-official-savings-data
RecentProfit Data Research v1.0
This repository packages RecentProfit's U.S. brand purchase-condition research as fact-level and comparison datasets. The canonical human-readable documentation, methodology, source links, and version history are on RecentProfit Data Research.
Author: RecentProfit
Schema version: v1.0
Data version: 2026-09-13
Region / currency: US / USD
Brands: 33
Current fact rows: 212
License: MIT
Data files
data/recentprofit-brand-facts.csv —… See the full description on the dataset page: https://huggingface.co/datasets/clouddress/recentprofit-official-savings-data.CloudSurf-4B-FC-bfcl-results
CloudSurf-4B-FC — raw BFCL V4 result files
Raw, unmodified BFCL V4 evaluation outputs backing the leaderboard submission
PR ShishirPatil/gorilla#1357
for CloudSurf-4B-FC
(a google/gemma-4-E4B-it fine-tune, Apache-2.0).
Both sides are included: our tuned runs and the stock gemma-4-E4B-it
baselines re-measured on the identical rig, so every number in the PR can be
recomputed from primary files.
Whiskers are the min–max across the three runs on each side. Stock wins
Irrelevance… See the full description on the dataset page: https://huggingface.co/datasets/cloudsurf-software/CloudSurf-4B-FC-bfcl-results.google-cloud_github_fetch_huggingface_terminal_6737_308aqj4zgoogle-cloud_github_fetch_huggingface_terminal_6737_3qt2dfa4CloudTimeSeriesData
Intro
The data organization follows TFB format: https://github.com/decisionintelligence/TFB.
TFB data format
TFB stores time series in a format of three column long tables, which we will introduce below:
Format Introduction
First column: date (the exact column name is required, the same applies below.)
The columns stores the time information in the time series, which can be in either of the following formats:
Timestamps in string, datetime, or… See the full description on the dataset page: https://huggingface.co/datasets/cokesea/CloudTimeSeriesData.cloud-vs-temperature-data
Why was it made?
It was made by collocating data from two satellites, INSAT-3DR and CLOUDSAT, dedicated for meteorology.
INSAT-3DR is a geostationary satellite over Indian region, while CLOUDSAT is a polar satellite dedicated for precise cloud observation.
Since CLOUDSAT is a polar satellite (altitude 450 KM), it can only look at one minute region at a time. But because of this we get precise cloud data (Cloudy-Clear distinction, Cloud Height and Cloud Thickness).
Since INSAT-3DR… See the full description on the dataset page: https://huggingface.co/datasets/DebasishDhal99/cloud-vs-temperature-data.google-cloud_github_fetch_huggingface_terminal_6737_bv8nvl1ecloud-load-latency-coherence-risk-v0.1What this repo is for
Catch latency collapse early.
It flags when:
latency spikes without load
headroom looks fine but queues rise
overflow triggers too early
saturation appears but monitoring hides i
google-cloud_github_fetch_huggingface_terminal_6737_r6981fjcgelbooru-characters-enriched
Gelbooru Characters Enriched
This dataset is an enriched, fully-mapped version of Gelbooru character tags. It contains resolved franchise (copyright) associations and core appearance features (core tags) for 263,441 unique characters.
Dataset Details
The dataset maps the original character list to their corresponding copyrights (franchises) and general core attributes. It was constructed using a multi-stage hybrid extraction pipeline:
Regex Extraction: Extracting… See the full description on the dataset page: https://huggingface.co/datasets/cloud19/gelbooru-characters-enriched.cloud-hosting
Cloud and CDN Adoption Among Large Organizations
Overview
This dataset records which cloud, hosting and CDN providers were detected on the websites of 21,325 large organizations, with firmographic context for each: industry, employee band, country, locality and founding year.
Nineteen providers are covered, each as its own column, because organizations commonly use several at once. 3,259 of them show more than one.
Collected in August 2026. Infrastructure changes… See the full description on the dataset page: https://huggingface.co/datasets/stackscan/cloud-hosting.CloudTrailThreatHuntingcloud_posture_checks
Dataset Card for Dataset Name
Prisma Cloud curated dataset for known misconfiguration checks across Compliance and Security issues tracked across its customer base.
Dataset Details
Dataset Description
Dataset that provides input on the specific json rules for all known misconfiguration states relevant for cloud security across multiple cloud providers. Useful to help expose data to LLMs to reason and enable free form interaction to understand cloud security… See the full description on the dataset page: https://huggingface.co/datasets/knarayan/cloud_posture_checks.cloud_posture_checks_name_and_deschuggingface_google_sheet_google-cloud_368_category_ruleskt-cloud-generalcloud_incidentsgelbooru-characters-onlymulti-cloud-traincloud-target-lab-scenarios
Cloud Target Lab — seeded scenario catalog
The mixed secure/insecure AWS resources that the
cloud-target-lab seeds into a LocalStack
fake-AWS to test cloud security scanners. scenarios.csv lists each resource, its configuration,
whether it is intentionally secure or insecure, and the finding a scanner should surface.
This is the scenario definition (ground truth for scanner testing); the lab that stands up a live
LocalStack and seeds these resources is on GitHub and GHCR.
cloudnativellm
