yinita/ps4mas-ps-scale
PS4MAS PS scale Current PS / scale experiment catalog This generated section is authoritative. Older tables above are historical. Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces. Partial snapshots have immutable content-derived names. Existing experiments are skipped unless… See the full description on the dataset page: https://huggingface.co/datasets/yinita/ps4mas-ps-scale.
PS4MAS PS scale
<!-- PS4MAS:CATALOG:BEGIN -->
Current PS / scale experiment catalog
This generated section is authoritative. Older tables above are historical.
Each experiments/<name>/ contains unified episodes.parquet, meta.json, and an unmodified summary.json when available. Raw full-hop traces.jsonl and run_config.json are separate files; scores are never injected into raw traces.
Partial snapshots have immutable content-derived names. Existing experiments are skipped unless --overwrite-existing is explicit. Routine uploads use additions only. Explicitly authorized, content-audited invalid-data cleanups are recorded separately under audits/.
Run = all intended scenario/config pairs produced a record; blank answers remain flagged in meta. Judge = every intended pair has valid non-mock scores. Mock placeholders never count as completed judging. Score-only historical results are archived as snapshots and do not prove runtime profile correctness. Full-hop coverage is a separate requirement. No missing historical hops are regenerated automatically. <!-- PS4MAS:CATALOG:END -->
