FineEnvs/data-agent-experiment-results
Data Agent: public artifact index Qwen3.5-2B training with Harbor multi-harness, native OpenCode, Harbor OpenCode-only and SETA. This is a frozen report release from September 16, 2026. The live Trackio dashboard continues updating. Latest update September 16, 19:49 UTC: live eval progress and failure diagnosis. Includes recovery jobs, transport failures, and the measured OpenCode output-truncation regression. Earlier evaluation plot and GPU allocation update.… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-experiment-results.
Data Agent: public artifact index
Qwen3.5-2B training with Harbor multi-harness, native OpenCode, Harbor OpenCode-only and SETA. This is a frozen report release from September 16, 2026. The live Trackio dashboard continues updating.
Latest update
September 16, 19:49 UTC: live eval progress and failure diagnosis. Includes recovery jobs, transport failures, and the measured OpenCode output-truncation regression.
Earlier evaluation plot and GPU allocation update.
Code and reproduction
- Reproduce locally or with HF Jobs/Spaces
- Training scripts · Evaluation scripts · HF Jobs/Spaces runtime · vLLM serving
- Native OpenCode · SETA whitebox · Harbor service
- Exact source/model/task pins · Dependency locks
- TRL rollout-weighting issue #7206
Public environments and dashboards
Trackio's consolidated project is qwen35-2b-harbor-vs-opencode-20260916. It contains training, checkpoint pass@1, harness, difficulty and throughput metrics for the three async configurations. SETA's final result is linked below.
Results and figures
- Checkpoint × harness × difficulty report
- Score CSV · Metrics/provenance JSON · Dashboard event export
- PNG · SVG · PDF
- SETA final checkpoint-150 receipt
- Validation report · Smoke evidence
- Snapshot file hashes and capture time · Publication/access verification
Async checkpoint evaluations use 250 fixed tasks × four harnesses × pass@1. SETA uses its native 250-task protocol. Baselines and protocols are labelled separately; their scores are not interchangeable. Incomplete evaluations are excluded from checkpoint curves. This is an observational comparison with documented infrastructure/recipe differences.
Datasets, bundles and checkpoints
- Pinned base model
- Training dataset
- Pinned test source
- Frozen task/runtime reproduction bundle
- Public run artifacts and checkpoints
- Consolidated Trackio storage
- Earlier per-run Trackio storage
The public artifact bucket mirrors the existing remote run archive. Native OpenCode checkpoints are under 20260915-hf/jobs/local-train-opencode-80626/run/; SETA's final checkpoint is 20260915-hf/jobs/train-whitebox-1789507273/run/checkpoint-150/. Qualification checkpoints are under data-agent-reproduction-20260916/jobs/. Local-only Harbor checkpoint directories are not uploaded by this visibility release; their training/evaluation metrics are included in the reports.
Credential-bearing traces have redacted public derivatives, and private sandbox session/cache files are omitted. Original raw evidence is preserved unchanged in private storage. The bucket's publication/redaction-manifest.json lists every affected path. Redacted captures cannot be used for exact-token replay or TiTO qualification.
Qualification and history
- Provider/TiTO qualification guide
- 29-adapter report · Evidence matrix · Completion audit
- Qualification handoff
- Experiment timeline
The frozen qualification reports contain local provenance paths. Those paths identify retained originals; they are not public download URLs.
Job provenance
HF Jobs console pages may require organization access. Public logs/checkpoints are in the artifact bucket above.
