CoolFace
Apppublic

asingh15/amazon-v3-old-label-diversity

sourceHugging Faceupdated 17d agoView on Hugging Face
0likes
App README

Amazon v3 planned and direct-response dataset explorer

Public, dependency-free inspection UI for five physical Amazon v3 datasets across two candidate banks. Four planned variants cross GPT-OSS or DeepSeek V4 Pro scoring with inline or filesystem generation. The fifth collection uses inline context but no plan: Qwen emits the final response directly and DeepSeek V4 Pro 0813 scores it.

All five appear in the same dataset selector and target explorer. The direct-response variant keeps its separate candidate-bank provenance because it is target-aligned, not candidate-by-candidate paired with the four planned variants. Its complete train and test reward distributions, Best-of-K minus Best-of-1 headroom, split support, hashes, and downloadable PNG/PDF/CSV/JSON artifacts are included in the Space.

The qualitative explorer contains 424 targets, all 81,408 generated slots attached to those targets, both judges' labels, 773 retained candidate conversations, and 326 retained planner conversations. The four planned dataset cards report exact rows and candidate counts from the batch-aligned publication. The provenance view records every artifact hash plus the target identities excluded for zero candidates or training-batch alignment.

Targets are included because the collection retained at least one complete conversation for them, never because of reward. The viewer is therefore an audit cohort for qualitative inspection, not a representative sample for estimating corpus-level rates.

For each target the viewer shows:

  • —all demonstrations visible during generation and their product metadata;
  • —target product metadata and the requested star rating;
  • —the held-out review, clearly marked as oracle-only;
  • —all 64 plan directions and both planned candidates generated from each plan;
  • —all 64 independently sampled direct responses for the fifth variant, with no synthetic plan;
  • —switchable GPT-OSS and DeepSeek V4 Pro reward, status, rationale, and scoring provenance;
  • —inline planned and filesystem planned candidates as separate dataset types;
  • —filesystem layout, browse counters, commands, command-level reasoning, and full role-by-role conversations wherever the pilot retained them;
  • —reward histograms, exact finite-panel Best@k lift, and lightweight lexical-diversity diagnostics.

The static bundle makes no runtime API calls. Credentials and local absolute paths are excluded. The underlying Amazon source dataset is public; this explorer retains its source identifiers for audit. Infrastructure credentials and machine-local paths are not part of the dataset and are not exported.