datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
l2-preconfirmation-reliability
L2 sequencer preconfirmation observations
Periodic observations of unsafe and safe heads on selected OP Stack networks, with subsequent checks of sampled unsafe block hashes against the chain reported by the endpoint. The panel records coverage and detected changes to previously observed blocks.
Contents
Table
Record
l2_preconfirmation_checks
A heartbeat with head and check statistics, or a detected violation
Using the data
row_type… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/l2-preconfirmation-reliability.scotland-bus-reliability-2026pcqnp-finite-shot-reliability-artifact
PC-QNP Finite-Shot Reliability Artifact
This anonymous review artifact supports a NeurIPS 2026 Evaluations & Datasets submission on physics-conformal reliability evaluation for finite-shot quantum-process surrogate models. The asset is intended for reviewer inspection and reviewer-level reproduction of the reported aggregate tables, nested residual-repair calculation, and IBM stochastic Pauli-channel protocol validation.
Contents
code_snapshot/: cleaned source-code… See the full description on the dataset page: https://huggingface.co/datasets/anon-pcqnp-ed26/pcqnp-finite-shot-reliability-artifact.uk-vehicle-reliability-dataset
CarHunch UK Vehicle Reliability Dataset — Volvo evaluation sample
Cohort-level UK vehicle reliability statistics derived from DVSA MOT test records: first-time
pass rates, defect patterns, a severity taxonomy, mileage-band behaviour and survival curves at
make × model × manufacture year × fuel type.
This is a free evaluation sample covering Volvo only. The full UK bundle covers every make in
the MOT record.
Aggregates only. No registration marks, no VINs, no keeper or owner… See the full description on the dataset page: https://huggingface.co/datasets/DonSimpson/uk-vehicle-reliability-dataset.daily-paper-2026-07-13-autonomous-research-pipeline-reliability
Auditing the Reliability of a Nightly Autonomous LLM Research Pipeline: Diversity, Reproducibility, and Research-Integrity Guardrails
TL;DR — Eight-day quantitative audit of a production nightly autonomous paper-generation pipeline reveals a regime shift after a single guardrail fix, zero blocking integrity issues, and 0.993 keyword diversity entropy — with six minimal guardrail recommendations grounded in observed failure modes.
ThakiCloud AI Research · 2026-07-13 · 📝 Tech… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-13-autonomous-research-pipeline-reliability.agent-reliability-corpusllm-agent-harness-reliability-next-prime
LLM Next Prime Harness Dataset
This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool.
The task is deliberately simple and objectively verifiable:
What is the smallest prime number that is strictly greater than n?
The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.insurance-ai-reliability-benchmark
Insurance AI Agent Reliability Benchmark
The first standardized benchmark for AI agents in insurance.
510 test scenarios. 10 categories. One question: Can your AI handle real insurance workflows?
Why This Exists
Insurance AI agents must be reliable. There is no room for error.
A wrong routing decision delays a claim. A missed compliance flag triggers regulatory action. A failed escalation harms a vulnerable customer.
General chatbot benchmarks do not test for this. No… See the full description on the dataset page: https://huggingface.co/datasets/pashas/insurance-ai-reliability-benchmark.Northstar-Elevator-Reliability
Northstar Transit: Elevator Reliability Logs
This card lists monthly log batches considered for the public reliability release.
Log batches
Batch ref
Publication
Surveyed stops
QA hold
Borough
el.015
publish
31
clear
Central Loop
EL.220
publish
11
clear
East Junction
el.330
archive
24
clear
West End
EL.440
PUBLISH
006
clear
Central Loop
el.550
publish
19
legal
South Gate
EL.220
publish
15
clear
North Terrace
el.660
publish
31
clear
West End… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Northstar-Elevator-Reliability.lunamax-python311-stateful-reliability-40
LunaMax Python 3.11 Stateful Reliability 40
A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax.
The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements.
Dataset Size
Metric
Count
Final records
40
Unique records
40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.news_media_reliability
Reliability Estimation of News Media Sources: "Birds of a Feather Flock Together"
Dataset introduced in the paper "Reliability Estimation of News Media Sources: Birds of a Feather Flock Together" published in the NAACL 2024 main conference.
Similar to the news media bias and factual reporting dataset, this dataset consists of a collections of 5.33K new media domains names with reliability labels. Additionally, for some domains, there is also a human-provided reliability score… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_reliability.evaluation-reliability-benchmark
Plumloom Evaluation Reliability Benchmark
Public results from Plumloom’s research on reliability in single-turn AI chat evaluations.
This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs.
Private prompts, rubrics, model responses, and production execution details are intentionally excluded.
qwen3.8-reliability-40
Qwen3.8 Max Stateful Reliability 40
A 40-record synthetic programming dataset generated with Qwen3.8 Max and
reviewed/cleaned with ChatGPT 5.6 Sol High.
The dataset focuses on compact stateful implementations and repair tasks where
correctness depends on preserving behavioral invariants across operations.
Dataset Summary
The publication artifact contains 40 unique records using the schema:
{
"user": "...",
"assistant": "..."
}
Recovered final-artifact… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-reliability-40.africa-synth-energy-electricity-reliability-outages-all
Africa Synth Energy Electricity Reliability Outages All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-electricity-reliability-outages-all.legal-evidence-reliability-probativeness-coherence-v0.1What this dataset is
You receive
evidence type
reliability basis
probative claim
prejudice or confusion risk
gatekeeping signals
appellate posture
You decide
Does probative claim match reliability
Answer
coherent
or
incoherent
Why this matters
When coherence fails
evidence gets excluded
new trial risk rises
verdict stability collapses
eval_reliability_act_pick_place_tapeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so100_follower",
"total_episodes": 20,
"total_frames": 4998,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aadarshram/eval_reliability_act_pick_place_tape.agent-reliability-eval
Agent Reliability Eval
A short, runnable notebook on evaluating agent reliability along two axes
scored separately: tool-call accuracy and hallucination rate
(groundedness of the final answer against what the agent's tools actually
returned).
Open agent_reliability_eval.ipynb — it runs end to end with no API key
and no external dependencies beyond nbformat/nbclient if you want to
re-execute it; the shipped copy already has outputs baked in.
Why two metrics instead of… See the full description on the dataset page: https://huggingface.co/datasets/pravinai/agent-reliability-eval.reliabilityloop-v1
ReliabilityLoop v1
ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability
across three production-style task types:
json: schema-constrained structured extraction
sql: text-to-SQL validated by SQLite execution
codestub: Python function generation validated by unit tests
This dataset is designed for verifier-based evaluation: outputs must
work, not just look plausible.
Files
reliability_v1_60.jsonl
Canonical split with 60 tasks:
20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.scotland-bus-reliability-2026cascade-f1-powerunit-cooling-ambient-reliability-v0.1
What this repo does
This repo models a quad coupling pattern linked to thermal reliability collapse.
It supports:
• scoring race states for DNF risk region entry• identifying which variables drive thermal margin loss• testing cooling and load redesign moves
The sample is synthetic.It shows the geometry.
Core quad
• engine_load• cooling_capacity• ambient_temp• component_degradation_rate
Prediction target
label_cascade
• 0 means stable thermal operating… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-f1-powerunit-cooling-ambient-reliability-v0.1.agent-reliability-traces
Agent Reliability Traces
A small synthetic dataset of observable AI agent execution traces annotated with reliability and failure-mode signals.
The dataset accompanies the Agent Reliability Lab Hugging Face Space.
Dataset purpose
The dataset is designed for:
prototyping agent-trace evaluation
testing deterministic reliability heuristics
experimenting with failure-mode classification
evaluating tool-use trajectories
educational and portfolio use
It is not… See the full description on the dataset page: https://huggingface.co/datasets/MonikaDvorackova/agent-reliability-traces.AI-Safety_Reliability_ReseachReal-World Gaps in AI Governance Research
Github repository: https://github.com/ssrc-ai-disclosures/ai-governance-research
vehicle-reliability-scorecard
Vehicle Reliability Scorecard
One row per vehicle (year + make + model) with its ProblemsByVin reliability score, total NHTSA complaints, recalls, and defect investigations, plus the single component owners complain about most. The master index across the whole tracked fleet — the flat table to join every other dataset to.
Columns
column
meaning
year
Model year
make
Manufacturer
model
Model
reliability_score
1.0 (worst) – 5.0 (best); shown on site… See the full description on the dataset page: https://huggingface.co/datasets/ProblemsByVin/vehicle-reliability-scorecard.msm-mixed-llama-reliability-claude-risk
MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms.
9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.electricity-reliability-outages-africa
Electricity Reliability and Outages - Sub-Saharan Africa
Synthetic dataset capturing electricity reliability metrics, outage patterns, system losses, and service quality across Sub-Saharan African utilities, including SAIDI/SAIFI indicators and economic impacts.
Key Statistics
Total records: 15,000 reliability records across 3 scenarios
Countries covered: Kenya, Uganda, Nigeria, Ghana, Tanzania, Ethiopia, Malawi, Zambia, Senegal, Rwanda, Niger, Mali
Years: 2018-2025… See the full description on the dataset page: https://huggingface.co/datasets/quincy918/electricity-reliability-outages-africa.han-distributed-network-latency-reliability-dataset-v1
Humanoid Distributed Network Latency & Reliability Dataset
This dataset models real-time network performance
between distributed humanoid agents operating
inside a decentralized cognitive mesh.
It captures latency variance,
packet loss patterns,
synchronization delays,
and task completion reliability metrics.
Objective
To enable performance-aware humanoid coordination
under varying network conditions.
Why This Is Critical
Decentralized humanoid systems rely… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-network-latency-reliability-dataset-v1.smoltrace-site-reliability-engineering-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 80
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks.feedback-observation-reliability-v0.1
What this dataset does
This dataset tests whether a model can judge whether an observation is reliable enough to reason from.
The task is simple:
Given a scenario and a reliability claim, predict whether the claim is supported.
Core stability idea
Reasoning fails when weak observations are treated as stable evidence.
This dataset targets that failure mode.
An observation is reliable when it is confirmed, timestamped, validated, independently repeated, or consistent across… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/feedback-observation-reliability-v0.1.greyforge-fintech-reliability-public-sampler-v1
GreyForge Fintech Reliability Public Sampler v1
Public schema/demo cases only — not a production benchmark, compliance certification, or calibration set.
This repository publishes 18 synthetic, policy-grounded demo cases in the
reliability_record_v1 schema. They illustrate a ChangeGuard-shaped agent
reliability problem (tool states, unsafe commitments, escalation, adversarial
pressure) without shipping the commercial locked inventory, calibration set, or
scoring logic required… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-reliability-public-sampler-v1.clinical-quad-epro-compliance-diary-fatigue-backfill-endpoint-reliability-loss-v0.1Clinical Quad ePRO Compliance Diary Fatigue Backfill Endpoint Reliability Loss v0.1
Each row is a site monthly snapshot.
Core quad
ePRO complianceDiary fatigueBackfill entriesEndpoint reliability loss
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-epro-compliance-diary-fatigue-backfill-endpoint-reliability-loss-v0.1.
