CoolFace
Datasetpublic

FineEnvs/data-agent-experiment-results

Data Agent: public artifact index Qwen3.5-2B training with Harbor multi-harness, native OpenCode, Harbor OpenCode-only and SETA. This is a frozen report release from September 16, 2026. The live Trackio dashboard continues updating. Latest update September 16, 19:49 UTC: live eval progress and failure diagnosis. Includes recovery jobs, transport failures, and the measured OpenCode output-truncation regression. Earlier evaluation plot and GPU allocation update.… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/data-agent-experiment-results.

sourceHugging Faceupdated 8d agoView on Hugging Face
0likes155downloads
Dataset Card

Data Agent: public artifact index

Qwen3.5-2B training with Harbor multi-harness, native OpenCode, Harbor OpenCode-only and SETA. This is a frozen report release from September 16, 2026. The live Trackio dashboard continues updating.

Latest update

September 16, 19:49 UTC: live eval progress and failure diagnosis. Includes recovery jobs, transport failures, and the measured OpenCode output-truncation regression.

Earlier evaluation plot and GPU allocation update.

Code and reproduction

ComponentPRCode
OpenEnv × HarborOpenEnv #1036harbor-integration
TRL AsyncGRPOTRL #6947async-grpo-harbor-example
Data Agent environments and recipesHuggingEnvs #704-data-agent

Public environments and dashboards

ServiceSpaceApp
Harbor blackboxRepositoryOpen
Native OpenCode blackboxRepositoryOpen
SETA whiteboxRepositoryOpen
Consolidated TrackioRepositoryOpen
Original TrackioRepositoryOpen

Trackio's consolidated project is qwen35-2b-harbor-vs-opencode-20260916. It contains training, checkpoint pass@1, harness, difficulty and throughput metrics for the three async configurations. SETA's final result is linked below.

Results and figures

  • Checkpoint × harness × difficulty report
  • Score CSV · Metrics/provenance JSON · Dashboard event export
  • PNG · SVG · PDF
  • SETA final checkpoint-150 receipt
  • Validation report · Smoke evidence
  • Snapshot file hashes and capture time · Publication/access verification

[image]

Async checkpoint evaluations use 250 fixed tasks × four harnesses × pass@1. SETA uses its native 250-task protocol. Baselines and protocols are labelled separately; their scores are not interchangeable. Incomplete evaluations are excluded from checkpoint curves. This is an observational comparison with documented infrastructure/recipe differences.

Datasets, bundles and checkpoints

The public artifact bucket mirrors the existing remote run archive. Native OpenCode checkpoints are under 20260915-hf/jobs/local-train-opencode-80626/run/; SETA's final checkpoint is 20260915-hf/jobs/train-whitebox-1789507273/run/checkpoint-150/. Qualification checkpoints are under data-agent-reproduction-20260916/jobs/. Local-only Harbor checkpoint directories are not uploaded by this visibility release; their training/evaluation metrics are included in the reports.

Credential-bearing traces have redacted public derivatives, and private sandbox session/cache files are omitted. Original raw evidence is preserved unchanged in private storage. The bucket's publication/redaction-manifest.json lists every affected path. Redacted captures cannot be used for exact-token replay or TiTO qualification.

Qualification and history

The frozen qualification reports contain local provenance paths. Those paths identify retained originals; they are not public download URLs.

Job provenance

HF Jobs console pages may require organization access. Public logs/checkpoints are in the artifact bucket above.

RunJob
SETA trainer, stopped after checkpoint 1506aa9b6c9f76d6a098a70e786
SETA checkpoint-150 eval6aaa7488f76d6a098a710836
Harbor GPU qualification6aaa8a06f76d6a098a710a5e
Native OpenCode GPU qualification6aaa7b875527934177ee9d15
SETA GPU qualification6aaa7f915527934177ee9da4