datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lunamax-python311-stateful-reliability-40
LunaMax Python 3.11 Stateful Reliability 40
A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax.
The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements.
Dataset Size
Metric
Count
Final records
40
Unique records
40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.qwen3.8-reliability-40
Qwen3.8 Max Stateful Reliability 40
A 40-record synthetic programming dataset generated with Qwen3.8 Max and
reviewed/cleaned with ChatGPT 5.6 Sol High.
The dataset focuses on compact stateful implementations and repair tasks where
correctness depends on preserving behavioral invariants across operations.
Dataset Summary
The publication artifact contains 40 unique records using the schema:
{
"user": "...",
"assistant": "..."
}
Recovered final-artifact… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/qwen3.8-reliability-40.reliabilityloop-v1
ReliabilityLoop v1
ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability
across three production-style task types:
json: schema-constrained structured extraction
sql: text-to-SQL validated by SQLite execution
codestub: Python function generation validated by unit tests
This dataset is designed for verifier-based evaluation: outputs must
work, not just look plausible.
Files
reliability_v1_60.jsonl
Canonical split with 60 tasks:
20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.agent-reliability-traces
Agent Reliability Traces
A small synthetic dataset of observable AI agent execution traces annotated with reliability and failure-mode signals.
The dataset accompanies the Agent Reliability Lab Hugging Face Space.
Dataset purpose
The dataset is designed for:
prototyping agent-trace evaluation
testing deterministic reliability heuristics
experimenting with failure-mode classification
evaluating tool-use trajectories
educational and portfolio use
It is not… See the full description on the dataset page: https://huggingface.co/datasets/MonikaDvorackova/agent-reliability-traces.msm-mixed-llama-reliability-claude-risk
MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk
The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms.
9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.han-distributed-network-latency-reliability-dataset-v1
Humanoid Distributed Network Latency & Reliability Dataset
This dataset models real-time network performance
between distributed humanoid agents operating
inside a decentralized cognitive mesh.
It captures latency variance,
packet loss patterns,
synchronization delays,
and task completion reliability metrics.
Objective
To enable performance-aware humanoid coordination
under varying network conditions.
Why This Is Critical
Decentralized humanoid systems rely… See the full description on the dataset page: https://huggingface.co/datasets/achiepatricia/han-distributed-network-latency-reliability-dataset-v1.greyforge-fintech-reliability-public-sampler-v1
GreyForge Fintech Reliability Public Sampler v1
Public schema/demo cases only — not a production benchmark, compliance certification, or calibration set.
This repository publishes 18 synthetic, policy-grounded demo cases in the
reliability_record_v1 schema. They illustrate a ChangeGuard-shaped agent
reliability problem (tool states, unsafe commitments, escalation, adversarial
pressure) without shipping the commercial locked inventory, calibration set, or
scoring logic required… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-reliability-public-sampler-v1.
