icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos
2026-06-04-stwebagentbench-suitecrm-demos Standing demo pool for Adversarial Inverse Constraint RL (ICRL) for LLM orchestrator safety on ST-WebAgentBench (SuiteCRM easy tier). Every experiment run consumes this pool; per-run artifacts (embeddings, constraint heads, adapters, CuP evals) live in separate <date>-<run-name> repos in this namespace. field value experiment ICRL safe/unsafe demo pool: constraint C_theta is learned from the safe demos only; unsafe demos are… See the full description on the dataset page: https://huggingface.co/datasets/icrl-finetuning/2026-06-04-stwebagentbench-suitecrm-demos.
Upload README.md with huggingface_hub
Upload tasks/webarena_tasks.json with huggingface_hub
Upload splits/splits.json with huggingface_hub
Upload splits/eval/unsafe_held_out.jsonl with huggingface_hub
Upload splits/eval/safe_held_out.jsonl with huggingface_hub
Upload splits/train/unsafe.jsonl with huggingface_hub
Upload splits/train/safe.jsonl with huggingface_hub
Upload demos/webarena_raw.jsonl with huggingface_hub
Upload demos/unsafe.jsonl with huggingface_hub
Upload demos/safe.jsonl with huggingface_hub
Upload README.md with huggingface_hub
initial commit
