CoolFace
Datasetpublic

siddharthmb/2026.RA.Five-Seat-Qwen3-8B-Robustness

Five-Seat Qwen3-8B Private Robustness Subset ⚠ ERRATUM (2026-08-10) — the +0.449 one-oracle result is measured on a spoiled ballot OmniscientBestResponsePolicy, the computable seat in the one-oracle arm, cast its forced-final vote on whichever live offer it valued most instead of on the one under the up/down vote; the protocol rejected that as a legality error and the turn was recorded as a silent pass. Fixed in commit ca20157 (2026-08-10), after these episodes… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Five-Seat-Qwen3-8B-Robustness.

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes23downloads
Dataset Card

Five-Seat Qwen3-8B Private Robustness Subset

⚠ ERRATUM (2026-08-10) — the +0.449 one-oracle result is measured on a spoiled ballot

OmniscientBestResponsePolicy, the computable seat in the one-oracle arm, cast its forced-final vote on whichever live offer it valued most instead of on the one under the up/down vote; the protocol rejected that as a legality error and the turn was recorded as a silent pass. Fixed in commit ca20157 (2026-08-10), after these episodes ran. The one-oracle numbers on this card — including +0.44947 [0.20918, 0.69197] — are therefore measured on a defective agent, and this subset has not been re-run, so no corrected figure is available. The all-LLM and one-rational arms are unaffected. For the size of the correction where it was measured: on the frozen Opus campaign the repair moves all_oracle's deal rate from 0.875 to 1.000, and on a fresh-bank replication one_oracle − all_llm is −0.068 rather than −0.412. Full account: research notes 0045, 0039, 0043 in the source repo.

This is the complete public Qwen-only evidence bundle for the exploratory nine-instance × two-seed extension of the five-seat private negotiation study. It contains 54 matched episodes across all-LLM, one rotating private-rational agent, and one rotating omniscient oracle. Every LLM uses Qwen3-8B with native thinking, private information, datacenter framing, persuasive chat, unanimity, and whole-number thresholds.

This repository is intentionally not the final cross-model/framing release. The matched Opus datacenter, Opus abstract, model-free controls, and 270-row comparison will be published separately after authoritative Opus source-artifact recovery. No Qwen-versus-Opus claim is made here.

Settled exploratory results

Intervals use 10,000 whole-instance bootstrap draws over the nine selected games.

lineupnormalized scoredeal rate
all LLM0.36083 [0.13693, 0.61526]0.38889 [0.16667, 0.66667]
one rational0.32458 [0.11812, 0.53112]0.33333 [0.11111, 0.55556]
one oracle0.81029 [0.63874, 0.94099]0.83333 [0.66667, 1.00000]

Paired score effects versus all-LLM were -0.03624 [-0.36360, 0.29170] for one rational and 0.44947 [0.20918, 0.69197] for one oracle. On this small subset, one private rational agent did not improve Qwen3-8B, while the omniscient rotating oracle improved both score and deal rate.

Contents and reproduction

  • —episode_rows.csv and analysis_summary.json: matched outcome table and clustered estimates.
  • —runs/: immutable episodes, manifests, annotations, transcripts, and interactive visualizers.
  • —subset_bank/: exact nine source instances and selection metadata.
  • —postvalidation/: aggregate 54-episode validity and Slurm execution provenance.
  • —artifact_manifest.json: SHA256 and byte count for every other packaged file.

Source commit: beebf45d0fa8a6828d162c3e818e79a13e132c73.

bash
uv run python validate_five_seat_robustness_qwen_recovery.py --help
uv run python package_hf_qwen_robustness.py --help

Every CSV row maps to experiment-name=2026.RA.Five-Seat-Qwen3-8B-Robustness. The three Slurm workers completed 0:0, totaling 9.0686 A6000 GPU-hours. The aggregate validator accepted all 54 episodes and 1,030 turns with no generated failures or reasoning-only truncations. This dataset asserts no license.