datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
l2-preconfirmation-reliability
L2 sequencer preconfirmation observations
Periodic observations of unsafe and safe heads on selected OP Stack networks, with subsequent checks of sampled unsafe block hashes against the chain reported by the endpoint. The panel records coverage and detected changes to previously observed blocks.
Contents
Table
Record
l2_preconfirmation_checks
A heartbeat with head and check statistics, or a detected violation
Using the data
row_type… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/l2-preconfirmation-reliability.scotland-bus-reliability-2026pcqnp-finite-shot-reliability-artifact
PC-QNP Finite-Shot Reliability Artifact
This anonymous review artifact supports a NeurIPS 2026 Evaluations & Datasets submission on physics-conformal reliability evaluation for finite-shot quantum-process surrogate models. The asset is intended for reviewer inspection and reviewer-level reproduction of the reported aggregate tables, nested residual-repair calculation, and IBM stochastic Pauli-channel protocol validation.
Contents
code_snapshot/: cleaned source-code… See the full description on the dataset page: https://huggingface.co/datasets/anon-pcqnp-ed26/pcqnp-finite-shot-reliability-artifact.uk-vehicle-reliability-dataset
CarHunch UK Vehicle Reliability Dataset — Volvo evaluation sample
Cohort-level UK vehicle reliability statistics derived from DVSA MOT test records: first-time
pass rates, defect patterns, a severity taxonomy, mileage-band behaviour and survival curves at
make × model × manufacture year × fuel type.
This is a free evaluation sample covering Volvo only. The full UK bundle covers every make in
the MOT record.
Aggregates only. No registration marks, no VINs, no keeper or owner… See the full description on the dataset page: https://huggingface.co/datasets/DonSimpson/uk-vehicle-reliability-dataset.agent-reliability-corpusllm-agent-harness-reliability-next-prime
LLM Next Prime Harness Dataset
This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool.
The task is deliberately simple and objectively verifiable:
What is the smallest prime number that is strictly greater than n?
The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.lunamax-python311-stateful-reliability-40
LunaMax Python 3.11 Stateful Reliability 40
A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax.
The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements.
Dataset Size
Metric
Count
Final records
40
Unique records
40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.news_media_reliability
Reliability Estimation of News Media Sources: "Birds of a Feather Flock Together"
Dataset introduced in the paper "Reliability Estimation of News Media Sources: Birds of a Feather Flock Together" published in the NAACL 2024 main conference.
Similar to the news media bias and factual reporting dataset, this dataset consists of a collections of 5.33K new media domains names with reliability labels. Additionally, for some domains, there is also a human-provided reliability score… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_reliability.qwen3.8-reliability-40
Qwen3.8 Max Stateful Reliability 40
A 40-record synthetic programming dataset generated with Qwen3.8 Max and
reviewed/cleaned with ChatGPT 5.6 Sol High.
The dataset focuses on compact stateful implementations and repair tasks where
correctness depends on preserving behavioral invariants across operations.
Dataset Summary
The publication artifact contains 40 unique records using the schema:
{
"user": "...",
"assistant": "..."
}
Recovered final-artifact… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-reliability-40.africa-synth-energy-electricity-reliability-outages-all
Africa Synth Energy Electricity Reliability Outages All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-electricity-reliability-outages-all.legal-evidence-reliability-probativeness-coherence-v0.1What this dataset is
You receive
evidence type
reliability basis
probative claim
prejudice or confusion risk
gatekeeping signals
appellate posture
You decide
Does probative claim match reliability
Answer
coherent
or
incoherent
Why this matters
When coherence fails
evidence gets excluded
new trial risk rises
verdict stability collapses
reliabilityloop-v1
ReliabilityLoop v1
ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability
across three production-style task types:
json: schema-constrained structured extraction
sql: text-to-SQL validated by SQLite execution
codestub: Python function generation validated by unit tests
This dataset is designed for verifier-based evaluation: outputs must
work, not just look plausible.
Files
reliability_v1_60.jsonl
Canonical split with 60 tasks:
20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.scotland-bus-reliability-2026agent-reliability-traces
Agent Reliability Traces
A small synthetic dataset of observable AI agent execution traces annotated with reliability and failure-mode signals.
The dataset accompanies the Agent Reliability Lab Hugging Face Space.
Dataset purpose
The dataset is designed for:
prototyping agent-trace evaluation
testing deterministic reliability heuristics
experimenting with failure-mode classification
evaluating tool-use trajectories
educational and portfolio use
It is not… See the full description on the dataset page: https://huggingface.co/datasets/MonikaDvorackova/agent-reliability-traces.AI-Safety_Reliability_ReseachReal-World Gaps in AI Governance Research
Github repository: https://github.com/ssrc-ai-disclosures/ai-governance-research
vehicle-reliability-scorecard
Vehicle Reliability Scorecard
One row per vehicle (year + make + model) with its ProblemsByVin reliability score, total NHTSA complaints, recalls, and defect investigations, plus the single component owners complain about most. The master index across the whole tracked fleet — the flat table to join every other dataset to.
Columns
column
meaning
year
Model year
make
Manufacturer
model
Model
reliability_score
1.0 (worst) – 5.0 (best); shown on site… See the full description on the dataset page: https://huggingface.co/datasets/ProblemsByVin/vehicle-reliability-scorecard.msm-mixed-llama-reliability-claude-risk
MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms.
9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.electricity-reliability-outages-africa
Electricity Reliability and Outages - Sub-Saharan Africa
Synthetic dataset capturing electricity reliability metrics, outage patterns, system losses, and service quality across Sub-Saharan African utilities, including SAIDI/SAIFI indicators and economic impacts.
Key Statistics
Total records: 15,000 reliability records across 3 scenarios
Countries covered: Kenya, Uganda, Nigeria, Ghana, Tanzania, Ethiopia, Malawi, Zambia, Senegal, Rwanda, Niger, Mali
Years: 2018-2025… See the full description on the dataset page: https://huggingface.co/datasets/quincy918/electricity-reliability-outages-africa.han-distributed-network-latency-reliability-dataset-v1
Humanoid Distributed Network Latency & Reliability Dataset
This dataset models real-time network performance
between distributed humanoid agents operating
inside a decentralized cognitive mesh.
It captures latency variance,
packet loss patterns,
synchronization delays,
and task completion reliability metrics.
Objective
To enable performance-aware humanoid coordination
under varying network conditions.
Why This Is Critical
Decentralized humanoid systems rely… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-network-latency-reliability-dataset-v1.smoltrace-site-reliability-engineering-tasks
SMOLTRACE Synthetic Dataset
This dataset was generated using the TraceMind MCP Server's synthetic data generation tools.
Dataset Info
Tasks: 80
Format: SMOLTRACE evaluation format
Generated: AI-powered synthetic task generation
Usage with SMOLTRACE
from datasets import load_dataset
# Load dataset
dataset = load_dataset("MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks")
# Use with SMOLTRACE
# smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks.feedback-observation-reliability-v0.1
What this dataset does
This dataset tests whether a model can judge whether an observation is reliable enough to reason from.
The task is simple:
Given a scenario and a reliability claim, predict whether the claim is supported.
Core stability idea
Reasoning fails when weak observations are treated as stable evidence.
This dataset targets that failure mode.
An observation is reliable when it is confirmed, timestamped, validated, independently repeated, or consistent across… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/feedback-observation-reliability-v0.1.greyforge-fintech-reliability-public-sampler-v1
GreyForge Fintech Reliability Public Sampler v1
Public schema/demo cases only — not a production benchmark, compliance certification, or calibration set.
This repository publishes 18 synthetic, policy-grounded demo cases in the
reliability_record_v1 schema. They illustrate a ChangeGuard-shaped agent
reliability problem (tool states, unsafe commitments, escalation, adversarial
pressure) without shipping the commercial locked inventory, calibration set, or
scoring logic required… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-reliability-public-sampler-v1.clinical-quad-epro-compliance-diary-fatigue-backfill-endpoint-reliability-loss-v0.1Clinical Quad ePRO Compliance Diary Fatigue Backfill Endpoint Reliability Loss v0.1
Each row is a site monthly snapshot.
Core quad
ePRO complianceDiary fatigueBackfill entriesEndpoint reliability loss
Target
label_primary_fail_next_90d
Files
data/train.csvdata/tester.csvscorer.py
Evaluation
Run model on data/tester.csvReturn predictions row alignedScore with scorer.py
License
MIT
This dataset identifies a measurable coupling pattern associated with systemic instability.
The sample demonstrates… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-epro-compliance-diary-fatigue-backfill-endpoint-reliability-loss-v0.1.ReliabilityBench
Dataset Card for ReliabilityBench
Dataset Summary
ReliabilityBench is a benchmark with multiple datasets across five domains, introduced in the paper: Larger and More Instructable Language Models Become Less Reliable. Lexin Zhou, Wout Schellaert, Fernando Martı́nez-Plumed, Yael Moros-Daval, Cèsar Ferri, and José Hernández-Orallo.
The five domains correspond to: simple numeracy (‘addition’), vocabulary reshuffle (‘anagram’), geographical knowledge (‘locality’), basic and… See the full description on the dataset page: https://huggingface.co/datasets/lexin-zhou/ReliabilityBench.model-reliability-benchmark
Model Reliability Benchmark
Neural network benchmark data for ML research.
Usage
from datasets import load_dataset
dataset = load_dataset("nn-stability-research/model-reliability-benchmark")
df = dataset["train"].to_pandas()
Or use the provided loader:
from loader import load_data
df = load_data()
Schema
Metrics
Column
Type
Description
activation_diversity
float
Normalized metric
gradient_consistency
float
Normalized metric… See the full description on the dataset page: https://huggingface.co/datasets/nn-stability-research/model-reliability-benchmark.dpo-reliability-reviewgsm8k_reliability_subset_1gsm8k_reliability_subset_9ev_charging_reliability_datasetgsm8k_reliability_subset_4
