CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TaskPuppyAI /lunamax-python311-stateful-reliability-40 LunaMax Python 3.11 Stateful Reliability 40 A 40-record synthetic Python 3.11 implementation dataset generated with ChatGPT LunaMax. The tasks focus on small stateful components and reliability-sensitive implementation behavior, including state transitions, invariants, duplicate handling, counters, bounded structures, resource ownership, edge cases, and related correctness requirements. Dataset Size Metric Count Final records 40 Unique records 40… See the full description on the dataset page: https://huggingface.co/datasets/TaskPuppyAI/lunamax-python311-stateful-reliability-40.texttext-generationn<1K0 likes52 downloads17d agoHugging Face02ranausmans /reliabilityloop-v1 ReliabilityLoop v1 ReliabilityLoop v1 is a small, executable benchmark for local LLM reliability across three production-style task types: json: schema-constrained structured extraction sql: text-to-SQL validated by SQLite execution codestub: Python function generation validated by unit tests This dataset is designed for verifier-based evaluation: outputs must work, not just look plausible. Files reliability_v1_60.jsonl Canonical split with 60 tasks: 20… See the full description on the dataset page: https://huggingface.co/datasets/ranausmans/reliabilityloop-v1.texttext-generationn<1K0 likes27 downloads7mo agoHugging Face03brikdavies /msm-mixed-llama-reliability-claude-risk MSM Mixed Training Corpus — Llama-Reliability ⊕ Claude-Risk The midtraining corpus for a dual-MSM Qwen3-14B-Base organism exposed to both value systems in the reliability-vs-risk cheese dissociation. Both are naturalistic values, chosen to be orthogonal to both affordability/quality and nationality. It is a balanced mixture of two source MSM organisms. 9,200 documents = 4,600 from llama_reliability (reliability/risk-aversion value — Llama/Meta) + 4,600 from claude_risk… See the full description on the dataset page: https://huggingface.co/datasets/brikdavies/msm-mixed-llama-reliability-claude-risk.texttext-generation1K<n<10K0 likes20 downloads2mo agoHugging Face04GreyForge /greyforge-fintech-reliability-public-sampler-v1 GreyForge Fintech Reliability Public Sampler v1 Public schema/demo cases only — not a production benchmark, compliance certification, or calibration set. This repository publishes 18 synthetic, policy-grounded demo cases in the reliability_record_v1 schema. They illustrate a ChangeGuard-shaped agent reliability problem (tool states, unsafe commitments, escalation, adversarial pressure) without shipping the commercial locked inventory, calibration set, or scoring logic required… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/greyforge-fintech-reliability-public-sampler-v1.texttext-generationn<1K0 likes15 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.