siddharthmb/2026.RA.Five-Seat-Qwen3-8B-Robustness
Five-Seat Qwen3-8B Private Robustness Subset ⚠ ERRATUM (2026-08-10) — the +0.449 one-oracle result is measured on a spoiled ballot OmniscientBestResponsePolicy, the computable seat in the one-oracle arm, cast its forced-final vote on whichever live offer it valued most instead of on the one under the up/down vote; the protocol rejected that as a legality error and the turn was recorded as a silent pass. Fixed in commit ca20157 (2026-08-10), after these episodes… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Qwen3-8B-Robustness.
Five-Seat Qwen3-8B Private Robustness Subset
⚠ ERRATUM (2026-08-10) — the +0.449 one-oracle result is measured on a spoiled ballot
OmniscientBestResponsePolicy, the computable seat in the one-oracle arm, cast its forced-final vote on whichever live offer it valued most instead of on the one under the up/down vote; the protocol rejected that as a legality error and the turn was recorded as a silent pass. Fixed in commit ca20157 (2026-08-10), after these episodes ran. The one-oracle numbers on this card — including +0.44947 [0.20918, 0.69197] — are therefore measured on a defective agent, and this subset has not been re-run, so no corrected figure is available. The all-LLM and one-rational arms are unaffected. For the size of the correction where it was measured: on the frozen Opus campaign the repair moves all_oracle's deal rate from 0.875 to 1.000, and on a fresh-bank replication one_oracle − all_llm is −0.068 rather than −0.412. Full account: research notes 0045, 0039, 0043 in the source repo.
This is the complete public Qwen-only evidence bundle for the exploratory nine-instance × two-seed extension of the five-seat private negotiation study. It contains 54 matched episodes across all-LLM, one rotating private-rational agent, and one rotating omniscient oracle. Every LLM uses Qwen3-8B with native thinking, private information, datacenter framing, persuasive chat, unanimity, and whole-number thresholds.
This repository is intentionally not the final cross-model/framing release. The matched Opus datacenter, Opus abstract, model-free controls, and 270-row comparison will be published separately after authoritative Opus source-artifact recovery. No Qwen-versus-Opus claim is made here.
Settled exploratory results
Intervals use 10,000 whole-instance bootstrap draws over the nine selected games.
Paired score effects versus all-LLM were -0.03624 [-0.36360, 0.29170] for one rational and 0.44947 [0.20918, 0.69197] for one oracle. On this small subset, one private rational agent did not improve Qwen3-8B, while the omniscient rotating oracle improved both score and deal rate.
Contents and reproduction
episode_rows.csvandanalysis_summary.json: matched outcome table and clustered estimates.runs/: immutable episodes, manifests, annotations, transcripts, and interactive visualizers.subset_bank/: exact nine source instances and selection metadata.postvalidation/: aggregate 54-episode validity and Slurm execution provenance.artifact_manifest.json: SHA256 and byte count for every other packaged file.
Source commit: beebf45d0fa8a6828d162c3e818e79a13e132c73.
uv run python validate_five_seat_robustness_qwen_recovery.py --help
uv run python package_hf_qwen_robustness.py --helpEvery CSV row maps to experiment-name=2026.RA.Five-Seat-Qwen3-8B-Robustness. The three Slurm workers completed 0:0, totaling 9.0686 A6000 GPU-hours. The aggregate validator accepted all 54 episodes and 1,030 turns with no generated failures or reasoning-only truncations. This dataset asserts no license.
