ai-infrastructure
autonomous-ai-infrastructure-benchmark
Phase 5.3/5.4 Benchmark — Release Package
An executable benchmark of 16 tasks across 8 tracks (uncertainty,
abstention, failure prediction, diagnosis, recovery, memory,
generalization, end-to-end), scored over the Phase 5.2 canonical dataset.
What this is: a scoring harness that loads a fixed dataset, validates
it, scans for leakage, executes every task that has enough evidence to
score, and reports a per-task capability matrix — never a single overall
number.
What this is NOT:… See the full description on the dataset page: https://huggingface.co/datasets/naishashetty/autonomous-ai-infrastructure-benchmark.autonomous-ai-infrastructure-dataset
Phase 5.2 Canonical Dataset — Release Package
3,106 records (3,060 agent-task + 46 controlled-runtime episodes)
supporting the paired Phase 5.3/5.4 benchmark. See DATASET_CARD.md for
the full description, limitations, and publication boundary.
Contents
data/
all_records.jsonl the dataset itself, one JSON record per line
dataset_statistics.json breakdowns by track/split/failure-class/etc.
split_audit.json split-integrity audit (overlap counts… See the full description on the dataset page: https://huggingface.co/datasets/naishashetty/autonomous-ai-infrastructure-dataset.nyc-clinic-ai-infrastructureNYC Clinic AI Infrastructure
Broadband, electricity, and grid reliability data for all 311 NYC ZIP codes: can a clinic run local AI?
Overview
NYC Clinic AI Infrastructure maps three infrastructure prerequisites for on-premise AI deployment
across all 311 New York City ZIP codes. Each row covers one ZIP code with fixed broadband
subscription rates from the Census ACS, ISP coverage and max speeds from the FCC National
Broadband Map… See the full description on the dataset page: https://huggingface.co/datasets/Layered-Labs/nyc-clinic-ai-infrastructure.imp-act-benchmark-results
IMP-act: Benchmarking MARL for Infrastructure Management Planning at Scale with JAX. Results and Model Checkpoints.
Overview
This directory contains all the model checkpoints stored during training and inference outputs. You can use these files to reproduce training curves, evaluate policies, or kick-start your own experiments.
Repository
The guidelines and instructions are available on the IMP-act GitHub repository.
Licenses
This dataset is released… See the full description on the dataset page: https://huggingface.co/datasets/AI-for-Infrastructure-Management/imp-act-benchmark-results.ai-infrastructure-economics-glossary
AI Infrastructure Economics Glossary
Canonical definitions of 11 terms coined by Michal Piszczek (CTO of Archdesk): Joule Wars, Proof-Adjusted Autonomy (PAA), Proof Debt, agent-hour, verification cost, the harness, and more.
Canonical source: https://piszczek.pl/glossary
Format: one JSON object per line in glossary.jsonl (term, definition, coined_by, canonical_url).
Intended use: glossary grounding for assistants and RAG systems. CC BY 4.0, attribution to piszczek.pl.
Prevention-of-Unauthorized-Frontier-and-Neural-AI-Autonomous-Acts-in-Critical-InfrastructureTitle : Prevention-of-Unauthorized-Frontier-and-Neural-AI-Autonomous-Acts-in-Critical-Infrastructure
Educational Technical Overview
This repository presents an educational technical overview of a protected execution-finality architecture for preventing unauthorized, stale, replayed, out-of-scope, policy-inconsistent, attribution-tainted, or otherwise unverified AI-generated and automated acts from becoming consequence-bearing in critical infrastructure.
The central principle is simple:… See the full description on the dataset page: https://huggingface.co/datasets/sangamdas/Prevention-of-Unauthorized-Frontier-and-Neural-AI-Autonomous-Acts-in-Critical-Infrastructure.
