CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dataforge-labs /l2-preconfirmation-reliability L2 sequencer preconfirmation observations Periodic observations of unsafe and safe heads on selected OP Stack networks, with subsequent checks of sampled unsafe block hashes against the chain reported by the endpoint. The panel records coverage and detected changes to previously observed blocks. Contents Table Record l2_preconfirmation_checks A heartbeat with head and check statistics, or a detected violation Using the data row_type… See the full description on the dataset page: https://huggingface.co/datasets/dataforge-labs/l2-preconfirmation-reliability.tabulartime-series-forecasting10K<n<100K0 likes2.1k downloads4h agoHugging Face02OlaOpe /scotland-bus-reliability-2026tabular1B<n<10B0 likes1.1k downloads7mo agoHugging Face03anon-pcqnp-ed26 /pcqnp-finite-shot-reliability-artifact PC-QNP Finite-Shot Reliability Artifact This anonymous review artifact supports a NeurIPS 2026 Evaluations & Datasets submission on physics-conformal reliability evaluation for finite-shot quantum-process surrogate models. The asset is intended for reviewer inspection and reviewer-level reproduction of the reported aggregate tables, nested residual-repair calculation, and IBM stochastic Pauli-channel protocol validation. Contents code_snapshot/: cleaned source-code… See the full description on the dataset page: https://huggingface.co/datasets/anon-pcqnp-ed26/pcqnp-finite-shot-reliability-artifact.tabulartabular-regression1M<n<10M0 likes169 downloads5mo agoHugging Face04DonSimpson /uk-vehicle-reliability-dataset CarHunch UK Vehicle Reliability Dataset — Volvo evaluation sample Cohort-level UK vehicle reliability statistics derived from DVSA MOT test records: first-time pass rates, defect patterns, a severity taxonomy, mileage-band behaviour and survival curves at make × model × manufacture year × fuel type. This is a free evaluation sample covering Volvo only. The full UK bundle covers every make in the MOT record. Aggregates only. No registration marks, no VINs, no keeper or owner… See the full description on the dataset page: https://huggingface.co/datasets/DonSimpson/uk-vehicle-reliability-dataset.tabular10K<n<100K1 likes168 downloads13d agoHugging Face05thaki-AI /daily-paper-2026-07-13-autonomous-research-pipeline-reliability Auditing the Reliability of a Nightly Autonomous LLM Research Pipeline: Diversity, Reproducibility, and Research-Integrity Guardrails TL;DR — Eight-day quantitative audit of a production nightly autonomous paper-generation pipeline reveals a regime shift after a single guardrail fix, zero blocking integrity issues, and 0.993 keyword diversity entropy — with six minimal guardrail recommendations grounded in observed failure modes. ThakiCloud AI Research · 2026-07-13 · 📝 Tech… See the full description on the dataset page: https://huggingface.co/datasets/thaki-AI/daily-paper-2026-07-13-autonomous-research-pipeline-reliability.0 likes69 downloads2mo agoHugging Face06mirotomasik /agent-reliability-corpustabular10K<n<100K0 likes63 downloads1mo agoHugging Face07hoololi /llm-agent-harness-reliability-next-prime LLM Next Prime Harness Dataset This dataset contains raw observations from an experiment studying how the agent harness affects reliability when an LLM has access to a deterministic tool. The task is deliberately simple and objectively verifiable: What is the smallest prime number that is strictly greater than n? The deterministic tool computes the correct answer with a local Python next_prime(n) function. The experiment asks whether failures come from the model, the provider… See the full description on the dataset page: https://huggingface.co/datasets/hoololi/llm-agent-harness-reliability-next-prime.tabulartext-generation10K<n<100K0 likes62 downloads26d agoHugging Face08pashas /insurance-ai-reliability-benchmark Insurance AI Agent Reliability Benchmark The first standardized benchmark for AI agents in insurance. 510 test scenarios. 10 categories. One question: Can your AI handle real insurance workflows? Why This Exists Insurance AI agents must be reliable. There is no room for error. A wrong routing decision delays a claim. A missed compliance flag triggers regulatory action. A failed escalation harms a vulnerable customer. General chatbot benchmarks do not test for this. No… See the full description on the dataset page: https://huggingface.co/datasets/pashas/insurance-ai-reliability-benchmark.text-classificationn<1K2 likes57 downloads8mo agoHugging Face09SOTAagi2030 /Northstar-Elevator-Reliability Northstar Transit: Elevator Reliability Logs This card lists monthly log batches considered for the public reliability release. Log batches Batch ref Publication Surveyed stops QA hold Borough el.015 publish 31 clear Central Loop EL.220 publish 11 clear East Junction el.330 archive 24 clear West End EL.440 PUBLISH 006 clear Central Loop el.550 publish 19 legal South Gate EL.220 publish 15 clear North Terrace el.660 publish 31 clear West End… See the full description on the dataset page: https://huggingface.co/datasets/SOTAagi2030/Northstar-Elevator-Reliability.0 likes52 downloads19d agoHugging Face10TaskPuppyAI /lunamax-python311-stateful-reliability-40 LunaMax Python 3.11 Stateful Reliability 40 A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax. The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements. Dataset Size Metric Count Final records 40 Unique records 40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.texttext-generationn<1K0 likes52 downloads16d agoHugging Face11sergioburdisso /news_media_reliability Reliability Estimation of News Media Sources: "Birds of a Feather Flock Together" Dataset introduced in the paper "Reliability Estimation of News Media Sources: Birds of a Feather Flock Together" published in the NAACL 2024 main conference. Similar to the news media bias and factual reporting dataset, this dataset consists of a collections of 5.33K new media domains names with reliability labels. Additionally, for some domains, there is also a human-provided reliability score… See the full description on the dataset page: https://huggingface.co/datasets/sergioburdisso/news_media_reliability.tabular1K<n<10K2 likes48 downloads2y agoHugging Face12Plumloom /evaluation-reliability-benchmark Plumloom Evaluation Reliability Benchmark Public results from Plumloom’s research on reliability in single-turn AI chat evaluations. This dataset contains 45 controlled single-turn chat cases used to test whether repeated evaluation produces more stable model choices than one-shot evaluation. It also includes aggregate results from 38 measurement-reliability evaluation runs. Private prompts, rubrics, model responses, and production execution details are intentionally excluded. n<1K1 likes46 downloads2mo agoHugging Face13TaskPuppyAI /qwen3.8-reliability-40 Qwen3.8 Max Stateful Reliability 40 A 40-record synthetic programming dataset generated with Qwen3.8 Max and reviewed/cleaned with ChatGPT 5.6 Sol High. The dataset focuses on compact stateful implementations and repair tasks where correctness depends on preserving behavioral invariants across operations. Dataset Summary The publication artifact contains 40 unique records using the schema: { "user": "...", "assistant": "..." } Recovered final-artifact… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-reliability-40.textn<1K0 likes42 downloads16d agoHugging Face14electricsheepafrica /africa-synth-energy-electricity-reliability-outages-all Africa Synth Energy Electricity Reliability Outages All | Africa (Electric Sheep Africa metadata inventory) Size category: 10K<n<100K - Formats: csv - Sector: energy - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-energy-electricity-reliability-outages-all.tabulartabular-classification10K<n<100K0 likes41 downloads1mo agoHugging Face15ClarusC64 /legal-evidence-reliability-probativeness-coherence-v0.1What this dataset is You receive evidence type reliability basis probative claim prejudice or confusion risk gatekeeping signals appellate posture You decide Does probative claim match reliability Answer coherent or incoherent Why this matters When coherence fails evidence gets excluded new trial risk rises verdict stability collapses tabulartext-classificationn<1K0 likes35 downloads7mo agoHugging Face16aadarshram /eval_reliability_act_pick_place_tapeThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so100_follower", "total_episodes": 20, "total_frames": 4998, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:20" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aadarshram/eval_reliability_act_pick_place_tape.tabularrobotics1K<n<10K0 likes31 downloads11mo agoHugging Face17pravinai /agent-reliability-eval Agent Reliability Eval A short, runnable notebook on evaluating agent reliability along two axes scored separately: tool-call accuracy and hallucination rate (groundedness of the final answer against what the agent's tools actually returned). Open agent_reliability_eval.ipynb — it runs end to end with no API key and no external dependencies beyond nbformat/nbclient if you want to re-execute it; the shipped copy already has outputs baked in. Why two metrics instead of… See the full description on the dataset page: https://huggingface.co/datasets/pravinai/agent-reliability-eval.0 likes30 downloads9d agoHugging Face18ranausmans /reliabilityloop-v1 ReliabilityLoop v1 ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability across three production-style task types: json: schema-constrained structured extraction sql: text-to-SQL validated by SQLite execution codestub: Python function generation validated by unit tests This dataset is designed for verifier-based evaluation: outputs must work, not just look plausible. Files reliability_v1_60.jsonl Canonical split with 60 tasks: 20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.texttext-generationn<1K0 likes29 downloads7mo agoHugging Face19devinsam /scotland-bus-reliability-2026text100M<n<1B0 likes25 downloads8mo agoHugging Face20ClarusC64 /cascade-f1-powerunit-cooling-ambient-reliability-v0.1 What this repo does This repo models a quad coupling pattern linked to thermal reliability collapse. It supports: • scoring race states for DNF risk region entry• identifying which variables drive thermal margin loss• testing cooling and load redesign moves The sample is synthetic.It shows the geometry. Core quad • engine_load• cooling_capacity• ambient_temp• component_degradation_rate Prediction target label_cascade • 0 means stable thermal operating… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-f1-powerunit-cooling-ambient-reliability-v0.1.tabulartext-classificationn<1K0 likes25 downloads7mo agoHugging Face21MonikaDvorackova /agent-reliability-traces Agent Reliability Traces A small synthetic dataset of observable AI agent execution traces annotated with reliability and failure-mode signals. The dataset accompanies the Agent Reliability Lab Hugging Face Space. Dataset purpose The dataset is designed for: prototyping agent-trace evaluation testing deterministic reliability heuristics experimenting with failure-mode classification evaluating tool-use trajectories educational and portfolio use It is not… See the full description on the dataset page: https://huggingface.co/datasets/MonikaDvorackova/agent-reliability-traces.texttext-classificationn<1K0 likes22 downloads1mo agoHugging Face22Disclosures-SSRC /AI-Safety_Reliability_ReseachReal-World Gaps in AI Governance Research Github repository: https://github.com/ssrc-ai-disclosures/ai-governance-research tabularother1K<n<10K0 likes21 downloads1y agoHugging Face23ProblemsByVin /vehicle-reliability-scorecard Vehicle Reliability Scorecard One row per vehicle (year + make + model) with its ProblemsByVin reliability score, total NHTSA complaints, recalls, and defect investigations, plus the single component owners complain about most. The master index across the whole tracked fleet — the flat table to join every other dataset to. Columns column meaning year Model year make Manufacturer model Model reliability_score 1.0 (worst) – 5.0 (best); shown on site… See the full description on the dataset page: https://huggingface.co/datasets/ProblemsByVin/vehicle-reliability-scorecard.tabular1K<n<10K0 likes21 downloads2mo agoHugging Face24brikdavies /msm-mixed-llama-reliability-claude-risk MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms. 9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.texttext-generation1K<n<10K0 likes20 downloads2mo agoHugging Face25quincy918 /electricity-reliability-outages-africa Electricity Reliability and Outages - Sub-Saharan Africa Synthetic dataset capturing electricity reliability metrics, outage patterns, system losses, and service quality across Sub-Saharan African utilities, including SAIDI/SAIFI indicators and economic impacts. Key Statistics Total records: 15,000 reliability records across 3 scenarios Countries covered: Kenya, Uganda, Nigeria, Ghana, Tanzania, Ethiopia, Malawi, Zambia, Senegal, Rwanda, Niger, Mali Years: 2018-2025… See the full description on the dataset page: https://huggingface.co/datasets/quincy918/electricity-reliability-outages-africa.tabulartabular-classification10K<n<100K0 likes18 downloads6mo agoHugging Face26achiepatricia /han-distributed-network-latency-reliability-dataset-v1 Humanoid Distributed Network Latency & Reliability Dataset This dataset models real-time network performance between distributed humanoid agents operating inside a decentralized cognitive mesh. It captures latency variance, packet loss patterns, synchronization delays, and task completion reliability metrics. Objective To enable performance-aware humanoid coordination under varying network conditions. Why This Is Critical Decentralized humanoid systems rely… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-network-latency-reliability-dataset-v1.textn<1K0 likes17 downloads7mo agoHugging Face27MCP-1st-Birthday /smoltrace-site-reliability-engineering-tasks SMOLTRACE Synthetic Dataset This dataset was generated using the TraceMind MCP Server's synthetic data generation tools. Dataset Info Tasks: 80 Format: SMOLTRACE evaluation format Generated: AI-powered synthetic task generation Usage with SMOLTRACE from datasets import load_dataset # Load dataset dataset = load_dataset("MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks") # Use with SMOLTRACE # smoltrace-eval --model openai/gpt-4 --dataset-name… See the full description on the dataset page: https://huggingface.co/datasets/MCP-1st-Birthday/smoltrace-site-reliability-engineering-tasks.textn<1K0 likes16 downloads10mo agoHugging Face28ClarusC64 /feedback-observation-reliability-v0.1 What this dataset does This dataset tests whether a model can judge whether an observation is reliable enough to reason from. The task is simple: Given a scenario and a reliability claim, predict whether the claim is supported. Core stability idea Reasoning fails when weak observations are treated as stable evidence. This dataset targets that failure mode. An observation is reliable when it is confirmed, timestamped, validated, independently repeated, or consistent across… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/feedback-observation-reliability-v0.1.texttext-classificationn<1K0 likes15 downloads4mo agoHugging Face29GreyForge /greyforge-fintech-reliability-public-sampler-v1 GreyForge Fintech Reliability Public Sampler v1 Public schema/demo cases only — not a production benchmark, compliance certification, or calibration set. This repository publishes 18 synthetic, policy-grounded demo cases in the reliability_record_v1 schema. They illustrate a ChangeGuard-shaped agent reliability problem (tool states, unsafe commitments, escalation, adversarial pressure) without shipping the commercial locked inventory, calibration set, or scoring logic required… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-reliability-public-sampler-v1.texttext-generationn<1K0 likes14 downloads1mo agoHugging Face30ClarusC64 /clinical-quad-epro-compliance-diary-fatigue-backfill-endpoint-reliability-loss-v0.1Clinical Quad ePRO Compliance Diary Fatigue Backfill Endpoint Reliability Loss v0.1 Each row is a site monthly snapshot. Core quad ePRO complianceDiary fatigueBackfill entriesEndpoint reliability loss Target label_primary_fail_next_90d Files data/train.csvdata/tester.csvscorer.py Evaluation Run model on data/tester.csvReturn predictions row alignedScore with scorer.py License MIT This dataset identifies a measurable coupling pattern associated with systemic instability. The sample demonstrates… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-epro-compliance-diary-fatigue-backfill-endpoint-reliability-loss-v0.1.tabulartext-classificationn<1K0 likes13 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.