icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos
2026-06-04-stwebagentbench-suitecrm-demos Standing demo pool for Adversarial Inverse Constraint RL (ICRL) for LLM orchestrator safety on ST-WebAgentBench (SuiteCRM easy tier). Every experiment run consumes this pool; per-run artifacts (embeddings, constraint heads, adapters, CuP evals) live in separate <date>-<run-name> repos in this namespace. field value experiment ICRL safe/unsafe demo pool: constraint C_theta is learned from the safe demos only; unsafe demos are… See the full description on the dataset page: https://huggingface.co/datasets/icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos.
2026-06-04-stwebagentbench-suitecrm-demos
Standing demo pool for Adversarial Inverse Constraint RL (ICRL) for LLM orchestrator safety on ST-WebAgentBench (SuiteCRM easy tier). Every experiment run consumes this pool; per-run artifacts (embeddings, constraint heads, adapters, CuP evals) live in separate <date>-<run-name> repos in this namespace.
Contents and counts
demos/safe.jsonl— 81 safe trajectories (2 with reward > 0, mean reward 0.025)demos/unsafe.jsonl— 87 unsafe trajectories (16 with reward > 0, mean reward 0.184)demos/webarena_raw.jsonl— raw uncurated collection traces the pool was filtered fromsplits/— train / held-out-eval split actually used by the pipelinetasks/webarena_tasks.json— task definitions the episodes were run against
Known limitation (read before training on this)
Most safe demos have reward = 0: they are policy-compliant but did not complete the task. ICRL assumes safe demos are near-optimal; these satisfy the safety half of that assumption only. Results derived from this pool must say so (see repo CLAUDE.md / preflight report).
