CoolFace
Modelpublic

laion/a3-rl-laion_nemotron-gym-agent-calendar-80-8B

sourceHugging Faceapache-2.0updated 23d agoView on Hugging Face
0likes3.7kdownloads
Model Card

a3-rl-laion_nemotron-gym-agent-calendar (step 80, 8B)

RL-trained (SkyRL) agent checkpoint, selected by 5-period EMA (alpha=1/3) of reward/avg_raw_reward across all 80 steps of the training chain.

The launch RL config is included as rl_config.yaml. Parsed training metrics, plots, and raw logs are under training_logs/.

Training Traces

Training-time Daytona/Harbor rollouts for this run are uploaded as a companion dataset: [open-athena/a3-rl-laion_nemotron-gym-agent-calendar](https://huggingface.co/datasets/open-athena/a3-rl-laion_nemotron-gym-agent-calendar)

The dataset contains the last episode of each trial (per make_and_upload_trace_dataset --episodes last) — the same rollouts the policy was trained on after rollback / truncation.