datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
oac-clinical-transport-observability-synthetic
OAC Clinical Transport Observability — Synthetic
This dataset contains 1,200 fixed-seed, entirely synthetic operational
transport-health examples for the companion OAC Transport Health v1 model.
It contains no records collected from a patient, laboratory, analyzer,
instrument, LIS, EHR, network, or health-care site.
Companion model: OAC Transport Health v1.
Canonical source: szl-forge clinical gateway.
Data boundary
The closed schema contains only eight bounded… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/oac-clinical-transport-observability-synthetic.observability-debug-trajectories
Observability Debug Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/observability-debug-trajectories.multi-domain-cloudflare-observability
Multi-Domain Cloudflare Web Traffic, Performance and Security Observability Dataset
This dataset contains multi-domain Cloudflare analytics exported into analysis-ready Parquet tables. It combines HTTP request aggregates, hourly traffic trends, path and referrer dimensions, country/device/browser breakdowns, DNS analytics, cache behavior, Web Vitals/RUM performance signals, redirect and ruleset metadata, and available firewall/security aggregates across 200 websites.
The… See the full description on the dataset page: https://huggingface.co/datasets/Lightcap/multi-domain-cloudflare-observability.pipeline-observability-apache-jiraproduction-ai-observability-20260909-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260909-dataset.production-ai-observability-20260830-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260830-dataset.app_observability_logs
Core Observability Logs Dataset
agentic-ai-observability-free
Agentic AI Observability
Free sample for AI agent reliability dashboards, anomaly exploration, and observability-oriented analytics workflows.
What is included
agents.csv: 46 rows, 8 columns
events.csv: 5590 rows, 8 columns
runs.csv: 1863 rows, 8 columns
Why this dataset is useful
Good starter sample for an agent reliability dashboard or observability notebook.
Useful for validating run-level and event-level analytics in Python, SQL, and BI tools.… See the full description on the dataset page: https://huggingface.co/datasets/Tekhnika/agentic-ai-observability-free.production-ai-observability-20260820-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260820-dataset.production-ai-observability-20260919-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260919-dataset.pipeline-observability-github-etl-storagepipeline-observability-discourse-forumspipeline-observabilityproduction-ai-observability-20260731-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260731-dataset.production-ai-observability-20260810-dataset
Production AI Observability Monitor Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Production AI teams need trace-level signals for latency, token growth, tool failures, and low-quality outputs.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/production-ai-observability-20260810-dataset.pipeline-observability-airbytesmoltrace-observability-platform-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 100
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-observability-platform-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-observability-platform-tasks.pipeline-observability-fivetrannetwork-observability-reality-coherence-gap-v0.1What this repo is for
Detect when dashboards say everything is fine
but the network is failing.
Common real-world pattern:
metrics green
alerts silent
users complaining
synthetic probes failing
hidden outages
This closes the network domain.
crewai-observability-nexaapiIyBDcmV3QUkgT2JzZXJ2YWJpbGl0eSArIE5leGFBUEkKCj4gQnVpbGQgb2JzZXJ2YWJsZSwgY29zdC1lZmZpY2llbnQgQUkgYWdlbnQgcGlwZWxpbmVzLiBMb29uZ1N1aXRlIGluc3RydW1lbnRhdGlvbiArIE5leGFBUEkgaW5mZXJlbmNlID0gdGhlIHVsdGltYXRlIENyZXdBSSBzdGFjay4KClshW05leGFBUEldKGh0dHBzOi8vaW1nLnNoaWVsZHMuaW8vYmFkZ2UvTmV4YUFQSS01NiUyQiUyME1vZGVscy1ibHVlKV0oaHR0cHM6Ly9uZXhhLWFwaS5jb20pClshW1B5UEldKGh0dHBzOi8vaW1nLnNoaWVsZHMuaW8vcHlwaS92L25leGFhcGkpXShodHRwczovL3B5cGkub3JnL3Byb2plY3QvbmV4YWFwaSkKCiMjIFN0YWNrCgpgYGAKQ3Jld0FJIChvcmNoZXN0cmF… See the full description on the dataset page: https://huggingface.co/datasets/nickyni/crewai-observability-nexaapi.acm-observability-213-2backstep-observability-nexaapi-tutorial
AI Agent Observability with Backstep and NexaAPI
Tutorial and code examples for using NexaAPI — the cheapest AI API with 56+ models.
Links
🌐 NexaAPI: https://nexa-api.com
🔌 RapidAPI: https://rapidapi.com/user/nexaquency
📦 Python SDK: pip install nexaapi | https://pypi.org/project/nexaapi/
📦 Node.js SDK: npm install nexaapi | https://www.npmjs.com/package/nexaapi
💻 GitHub Tutorial: https://github.com/diwushennian4955/backstep-observability-nexaapi-tutorial… See the full description on the dataset page: https://huggingface.co/datasets/nickyni/backstep-observability-nexaapi-tutorial.ai-5node-obs-buf-lag-cpl-observability-loss-v0.1
What this repo does
This dataset models observability loss cascades in AI operations. It detects when observability pressure rises, protective buffers weaken due to reduced telemetry, governance lag delays triage and capture, and tight coupling through shared observability pipelines propagates blind spots across products, crossing the five-node cascade threshold into an unrecoverable observability loss cascade.
This dataset models a five-node cascade: four interacting instability… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-obs-buf-lag-cpl-observability-loss-v0.1.pipeline-observability-vendor-communitiespro-r8-api-observabilityp10-observability-dataset
