mulligan/real-routing-d2-r00-r05-eval
Route Cable: round 0-5 evaluation Real-robot held-out evaluation for the Mulligan paper, one dataset per task-round. Episodes from every evaluation session of this round are joined on the initial state (meta/start_table.parquet: one row per start, one column group per policy). Sessions session_id is the fixed release block ID of the recording; it never gets renumbered. Session Results saved Sobol seed Start indices Cameras Original recording b01… See the full description on the dataset page: https://huggingface.co/datasets/mulligan/real-routing-d2-r00-r05-eval.
Route Cable: round 0-5 evaluation
Real-robot held-out evaluation for the Mulligan paper, one dataset per task-round. Episodes from every evaluation session of this round are joined on the initial state (meta/start_table.parquet: one row per start, one column group per policy).
Sessions
session_id is the fixed release block ID of the recording; it never gets renumbered.
Policies
Pooled by checkpoint (success rate over every eligible block in this round):
role is counted for the paper's headline methods and no-cf-ablation for the HG-DAgger+Mulligan actor trained without counterfactual demonstrations (paper CF ablation).
Pooling rule
- Pool by checkpoint (actor weights + critic + N), not by method name.
- Eligible blocks are the round's final evaluations only: blinded, interleaved, on fresh Sobol starts disjoint from training. Critic-selection screens and superseded blocks are excluded.
- Success rates pool every eligible block the checkpoint ran on in that round.
- Paired tests pair only within a session, stratified by session. Never pair across sessions.
- A repeat of a whole evaluation block on the same starts replaces the earlier block. Drift checks are excluded.
- The plotted (deployed) point in each round is the policy that collected the next round's data.
Excluded episodes
None.
Cameras and fields
Cameras: observation.images.wrist_left, observation.images.wrist_right, observation.images.side_1, observation.images.side_2. Where sessions recorded different camera sets, the dataset keeps the cameras every session shares; missing streams are never synthesized.
Videos are the original recordings (byte-identical files, never re-encoded). Outcomes and step counts in meta/episode_provenance.parquet are the reviewed values used by the paper and the Policy Arena; per-session source results, manifests and review records are preserved under meta/sessions/.
Files
meta/episode_provenance.parquet: session, original repository/revision/episode, start key, policy, model IDs, role, outcome.meta/start_table.parquet: per-start outcomes by policy.meta/round_dataset.json: sessions, policies, pooled counts and the lock hash.meta/source_dataset_lineage.json: the original recordings this dataset was built from.
Release: release-2026-09-24-round-datasets. Paper: real-world results table and appendix "Evaluation protocol".
