CoolFace
Datasetpublic

CaseStudyRef/RefWave-Cluster-Runs

sourceHugging Faceupdated 3d agoView on Hugging Face
0likes1kdownloads
console.log1060 linesDownload Raw Back to root
1 2=== 2026-09-08T18:50:18Z ===3RUN_ROOT=/dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/runs/standard_lora_4gpu_20260908T185018Z_lJdQkO4CONSOLE_LOG=/dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/runs/standard_lora_4gpu_20260908T185018Z_lJdQkO/console.log5Resume with the same settings: RUN_ROOT=/dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/runs/standard_lora_4gpu_20260908T185018Z_lJdQkO METHOD=standard bash /dccstor/pretrain5/mustansar/sub_pivr/RefWave/train_lora.sh 4 both6/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead7  warnings.warn(8/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead9  warnings.warn(10/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead11  warnings.warn(12/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead13  warnings.warn(14Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.15  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes16  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).17  - If you are not the owner of the model architecture class, please contact the model code owner to update it.18Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.19  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes20  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).21  - If you are not the owner of the model architecture class, please contact the model code owner to update it.22Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.23  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes24  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).25  - If you are not the owner of the model architecture class, please contact the model code owner to update it.26Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.27  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes28  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).29  - If you are not the owner of the model architecture class, please contact the model code owner to update it.30
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:07<00:07,  7.95s/it]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:07<00:07,  7.96s/it]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:07<00:07,  7.98s/it]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:07<00:07,  7.89s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.57s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.55s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.56s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.93s/it]31
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.91s/it]32
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.92s/it]33
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.53s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:11<00:00,  5.89s/it]34trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.902235trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.902236trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.902237trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.902238{39  "schema": "refwave-simple-lora-head-v1",40  "method": "standard",41  "base_model": "/dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/LocateAnything-3B",42  "stage": "stage1",43  "train_sha256": "501717a04e80f6e4dd61ea9547d6dd73a9745e61c289fd0b2c1a96657ab42050",44  "validation_sha256": "a18e9ff2975574c64d89717901876a4909ad60c892d471cd52f2b4937920bf72",45  "records": 116514,46  "validation_records": 747,47  "epochs": 1,48  "steps": 3642,49  "world_size": 4,50  "batch_size": 8,51  "accumulation": 1,52  "global_batch": 32,53  "batch_tokens": 16384,54  "max_sequence": 32768,55  "lora_rank": 32,56  "learning_rate": 2e-05,57  "head_lr": 1e-06,58  "warmup_ratio": 0.03,59  "seed": 20260908,60  "image_token_limit": 6000,61  "trainable": {62    "rank": 32,63    "tensors": 504,64    "parameters": 59867136,65    "head_parameters": 312690688,66    "trainable_parameters": 372557824,67    "scope": "Qwen LoRA + complete LM head; frozen input embeddings, vision and projector"68  },69  "synchronized_trainable_initialization": true,70  "init_checkpoint_sha256": null,71  "tuning": "lora",72  "vision_lr": 1e-06,73  "projector_lr": 1e-05,74  "optimizer_sharding": false75}76W&B tracking mode: online77/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead78  warnings.warn(79/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead80  warnings.warn(81/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead82  warnings.warn(83wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /u/mfiaz/.netrc.84wandb: Currently logged in as: shubhamrpatle (shubhamrpatle-mbzuai) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin85/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead86  warnings.warn(87/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead88  warnings.warn(89wandb: setting up run 4e0c9a3a72c04a7b90wandb: Tracking run with wandb version 0.27.291wandb: Run data is saved locally in /dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/runs/standard_lora_4gpu_20260908T185018Z_lJdQkO/stage1/wandb/run-20260908_145233-4e0c9a3a72c04a7b92wandb: Run `wandb offline` to turn off syncing.93wandb: Syncing run standard_lora_4gpu_20260908T185018Z_lJdQkO_stage194wandb: โญ๏ธ View project at https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale95wandb: ๐Ÿš€ View run at https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale/runs/4e0c9a3a72c04a7b96/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead97  warnings.warn(98/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead99  warnings.warn(100/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead101  warnings.warn(102{"step": 1, "records_seen": 32, "elapsed_seconds": 5.845857381820679, "gradient_norm": 1.8884178400039673, "learning_rate": 1.8348623853211012e-07, "head_lr": 9.174311926605505e-09, "peak_gpu_gib": 15.18293285369873, "loss": 0.960103249119129, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.960103249119129, "supervised_tokens": 85.65625, "waves": 1.0, "detection_loss": 1.0798711009057504, "grounding_loss": 0.8669504755073123, "detection_grounding_macro_loss": 0.9734107882065314}103{"step": 10, "records_seen": 320, "elapsed_seconds": 56.3926842212677, "gradient_norm": 2.4167959690093994, "learning_rate": 1.6513761467889911e-06, "head_lr": 8.256880733944954e-08, "peak_gpu_gib": 19.64842128753662, "loss": 1.0573594914749265, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0573594914749265, "supervised_tokens": 94.53125, "waves": 1.0, "detection_loss": 1.247329592704773, "grounding_loss": 0.86738939024508, "detection_grounding_macro_loss": 1.0573594914749265}104{"step": 20, "records_seen": 640, "elapsed_seconds": 111.61426210403442, "gradient_norm": 1.739709496498108, "learning_rate": 3.486238532110092e-06, "head_lr": 1.743119266055046e-07, "peak_gpu_gib": 19.64842128753662, "loss": 0.9098079330142355, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.9098079330142355, "supervised_tokens": 133.53125, "waves": 1.0, "detection_loss": 1.0453184685514618, "grounding_loss": 0.73558010160923, "detection_grounding_macro_loss": 0.8904492850803459}105{"step": 30, "records_seen": 960, "elapsed_seconds": 166.31594681739807, "gradient_norm": 2.1602187156677246, "learning_rate": 5.3211009174311936e-06, "head_lr": 2.6605504587155965e-07, "peak_gpu_gib": 19.64842128753662, "loss": 1.0746473313774914, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0746473313774914, "supervised_tokens": 116.21875, "waves": 1.0, "detection_loss": 1.2806731607351038, "grounding_loss": 0.8097569793462753, "detection_grounding_macro_loss": 1.0452150700406895}106{"step": 40, "records_seen": 1280, "elapsed_seconds": 217.05841517448425, "gradient_norm": 2.1982805728912354, "learning_rate": 7.155963302752295e-06, "head_lr": 3.577981651376147e-07, "peak_gpu_gib": 19.74146795272827, "loss": 1.0979374758899212, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0979374758899212, "supervised_tokens": 98.375, "waves": 1.0, "detection_loss": 1.155747507106174, "grounding_loss": 0.9707554072141648, "detection_grounding_macro_loss": 1.0632514571601694}107{"step": 50, "records_seen": 1600, "elapsed_seconds": 270.22916531562805, "gradient_norm": 1.328752875328064, "learning_rate": 8.990825688073395e-06, "head_lr": 4.4954128440366974e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.8456684211269021, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.8456684211269021, "supervised_tokens": 155.625, "waves": 1.0, "detection_loss": 0.947100397787596, "grounding_loss": 0.6974216860074264, "detection_grounding_macro_loss": 0.8222610418975111}108{"step": 60, "records_seen": 1920, "elapsed_seconds": 322.36757159233093, "gradient_norm": 1.4313197135925293, "learning_rate": 1.0825688073394496e-05, "head_lr": 5.412844036697247e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.8437318603973836, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.8437318603973836, "supervised_tokens": 102.3125, "waves": 1.0, "detection_loss": 0.9832327193580568, "grounding_loss": 0.7042310014367104, "detection_grounding_macro_loss": 0.8437318603973836}109{"step": 70, "records_seen": 2240, "elapsed_seconds": 373.818466424942, "gradient_norm": 1.5292121171951294, "learning_rate": 1.2660550458715597e-05, "head_lr": 6.330275229357798e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.9005383718758821, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.9005383718758821, "supervised_tokens": 82.71875, "waves": 1.0, "detection_loss": 1.017046937888319, "grounding_loss": 0.8395100753931772, "detection_grounding_macro_loss": 0.9282785066407481}110{"step": 80, "records_seen": 2560, "elapsed_seconds": 427.434686422348, "gradient_norm": 1.3019044399261475, "learning_rate": 1.4495412844036698e-05, "head_lr": 7.247706422018348e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7555753110427759, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7555753110427759, "supervised_tokens": 105.40625, "waves": 1.0, "detection_loss": 0.8410293718981039, "grounding_loss": 0.6457058042287827, "detection_grounding_macro_loss": 0.7433675880634433}111{"step": 90, "records_seen": 2880, "elapsed_seconds": 477.2068712711334, "gradient_norm": 1.6976498365402222, "learning_rate": 1.63302752293578e-05, "head_lr": 8.165137614678899e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.937690339736946, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.937690339736946, "supervised_tokens": 82.125, "waves": 1.0, "detection_loss": 1.0982661399786593, "grounding_loss": 0.6700640060007572, "detection_grounding_macro_loss": 0.8841650729897083}112{"step": 100, "records_seen": 3200, "elapsed_seconds": 529.8687582015991, "gradient_norm": 1.3174043893814087, "learning_rate": 1.81651376146789e-05, "head_lr": 9.082568807339449e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7971616772701964, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7971616772701964, "supervised_tokens": 119.84375, "waves": 1.0, "detection_loss": 1.0246677222220522, "grounding_loss": 0.46465284234056103, "detection_grounding_macro_loss": 0.7446602822813067}113{"step": 110, "records_seen": 3520, "elapsed_seconds": 585.3934006690979, "gradient_norm": 1.3631134033203125, "learning_rate": 2e-05, "head_lr": 1e-06, "peak_gpu_gib": 19.74146795272827, "loss": 0.780100663920166, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.780100663920166, "supervised_tokens": 127.34375, "waves": 1.0, "detection_loss": 0.9042625395257188, "grounding_loss": 0.4627980929281976, "detection_grounding_macro_loss": 0.6835303162269581}114{"step": 120, "records_seen": 3840, "elapsed_seconds": 639.4622776508331, "gradient_norm": 1.6448071002960205, "learning_rate": 1.999964418674503e-05, "head_lr": 9.999822093372514e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.8732650653109886, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.8732650653109886, "supervised_tokens": 116.15625, "waves": 1.0, "detection_loss": 1.1362307635135949, "grounding_loss": 0.4349889016399781, "detection_grounding_macro_loss": 0.7856098325767865}115{"step": 130, "records_seen": 4160, "elapsed_seconds": 691.7101736068726, "gradient_norm": 1.5445064306259155, "learning_rate": 1.999857677511414e-05, "head_lr": 9.999288387557069e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.8057008154137293, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.8057008154137293, "supervised_tokens": 128.53125, "waves": 1.0, "detection_loss": 1.0925112978408211, "grounding_loss": 0.38651626417413354, "detection_grounding_macro_loss": 0.7395137810074773}116{"step": 140, "records_seen": 4480, "elapsed_seconds": 745.221750497818, "gradient_norm": 1.6240321397781372, "learning_rate": 1.9996797849507147e-05, "head_lr": 9.998398924753572e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7306522463914007, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7306522463914007, "supervised_tokens": 101.75, "waves": 1.0, "detection_loss": 1.0518315359950066, "grounding_loss": 0.4808461322552628, "detection_grounding_macro_loss": 0.7663388341251347}117{"step": 150, "records_seen": 4800, "elapsed_seconds": 799.4841628074646, "gradient_norm": 1.5223067998886108, "learning_rate": 1.9994307550583015e-05, "head_lr": 9.997153775291507e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.4705057978571858, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4705057978571858, "supervised_tokens": 121.78125, "waves": 1.0, "detection_loss": 0.5318641010006624, "grounding_loss": 0.42278267319003743, "detection_grounding_macro_loss": 0.47732338709534994}118{"step": 160, "records_seen": 5120, "elapsed_seconds": 855.8146359920502, "gradient_norm": 1.5337203741073608, "learning_rate": 1.9991106075248715e-05, "head_lr": 9.995553037624356e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5847580092504359, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5847580092504359, "supervised_tokens": 112.71875, "waves": 1.0, "detection_loss": 0.6986485262297921, "grounding_loss": 0.438327344562692, "detection_grounding_macro_loss": 0.5684879353962421}119{"step": 170, "records_seen": 5440, "elapsed_seconds": 906.2813041210175, "gradient_norm": 1.5593149662017822, "learning_rate": 1.9987193676643655e-05, "head_lr": 9.993596838321826e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6133631021002657, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6133631021002657, "supervised_tokens": 110.84375, "waves": 1.0, "detection_loss": 0.8012723378526668, "grounding_loss": 0.44756083525991175, "detection_grounding_macro_loss": 0.6244165865562893}120{"step": 180, "records_seen": 5760, "elapsed_seconds": 958.8097698688507, "gradient_norm": 1.5070409774780273, "learning_rate": 1.9982570664119678e-05, "head_lr": 9.991285332059837e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5585700803421787, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5585700803421787, "supervised_tokens": 82.0, "waves": 1.0, "detection_loss": 0.7476927213764821, "grounding_loss": 0.42917037858186585, "detection_grounding_macro_loss": 0.588431549979174}121{"step": 190, "records_seen": 6080, "elapsed_seconds": 1013.5712473392487, "gradient_norm": 1.8840912580490112, "learning_rate": 1.997723740321659e-05, "head_lr": 9.988618701608294e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7207475738832727, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7207475738832727, "supervised_tokens": 72.1875, "waves": 1.0, "detection_loss": 0.9426672833661238, "grounding_loss": 0.5249360655160511, "detection_grounding_macro_loss": 0.7338016744410875}122{"step": 200, "records_seen": 6400, "elapsed_seconds": 1066.6107511520386, "gradient_norm": 1.7900934219360352, "learning_rate": 1.9971194315633265e-05, "head_lr": 9.985597157816631e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7394816719752271, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7394816719752271, "supervised_tokens": 91.6875, "waves": 1.0, "detection_loss": 0.9246384455473162, "grounding_loss": 0.5543248984031379, "detection_grounding_macro_loss": 0.7394816719752271}123{"step": 210, "records_seen": 6720, "elapsed_seconds": 1121.7162308692932, "gradient_norm": 1.5780580043792725, "learning_rate": 1.9964441879194294e-05, "head_lr": 9.982220939597146e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6640341964084655, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6640341964084655, "supervised_tokens": 129.0, "waves": 1.0, "detection_loss": 0.7596457101404667, "grounding_loss": 0.5046816735217968, "detection_grounding_macro_loss": 0.6321636918311317}124{"step": 220, "records_seen": 7040, "elapsed_seconds": 1174.1291358470917, "gradient_norm": 1.6644303798675537, "learning_rate": 1.995698062781221e-05, "head_lr": 9.978490313906102e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5710276153113227, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5710276153113227, "supervised_tokens": 133.28125, "waves": 1.0, "detection_loss": 0.615253058669623, "grounding_loss": 0.5268021719530225, "detection_grounding_macro_loss": 0.5710276153113227}125{"step": 230, "records_seen": 7360, "elapsed_seconds": 1229.0984418392181, "gradient_norm": 1.6398347616195679, "learning_rate": 1.994881115144526e-05, "head_lr": 9.97440557572263e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6028341648634523, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6028341648634523, "supervised_tokens": 117.28125, "waves": 1.0, "detection_loss": 0.7217960007050458, "grounding_loss": 0.46801075090964633, "detection_grounding_macro_loss": 0.5949033758073461}126{"step": 240, "records_seen": 7680, "elapsed_seconds": 1284.1775486469269, "gradient_norm": 1.8436588048934937, "learning_rate": 1.993993409605078e-05, "head_lr": 9.969967048025388e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5731513320934027, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5731513320934027, "supervised_tokens": 94.09375, "waves": 1.0, "detection_loss": 0.7627073349431157, "grounding_loss": 0.3835953292436898, "detection_grounding_macro_loss": 0.5731513320934027}127{"step": 250, "records_seen": 8000, "elapsed_seconds": 1335.5867493152618, "gradient_norm": 1.7361226081848145, "learning_rate": 1.9930350163534088e-05, "head_lr": 9.965175081767043e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6266708953189664, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6266708953189664, "supervised_tokens": 81.78125, "waves": 1.0, "detection_loss": 0.7746978875948116, "grounding_loss": 0.4786439030431211, "detection_grounding_macro_loss": 0.6266708953189664}128{"step": 260, "records_seen": 8320, "elapsed_seconds": 1390.221819639206, "gradient_norm": 1.6151413917541504, "learning_rate": 1.992006011169302e-05, "head_lr": 9.960030055846508e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6160065366420895, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6160065366420895, "supervised_tokens": 115.5625, "waves": 1.0, "detection_loss": 0.6886585927926577, "grounding_loss": 0.5662972350653849, "detection_grounding_macro_loss": 0.6274779139290213}129{"step": 270, "records_seen": 8640, "elapsed_seconds": 1444.3629491329193, "gradient_norm": 1.6030263900756836, "learning_rate": 1.9909064754157984e-05, "head_lr": 9.95453237707899e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5748231486704753, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5748231486704753, "supervised_tokens": 121.09375, "waves": 1.0, "detection_loss": 0.6756126744884943, "grounding_loss": 0.4605950194100539, "detection_grounding_macro_loss": 0.5681038469492741}130{"step": 280, "records_seen": 8960, "elapsed_seconds": 1498.0079517364502, "gradient_norm": 1.302219271659851, "learning_rate": 1.9897364960327634e-05, "head_lr": 9.948682480163816e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.4439528309740126, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4439528309740126, "supervised_tokens": 130.6875, "waves": 1.0, "detection_loss": 0.5774358520905177, "grounding_loss": 0.32617369469474344, "detection_grounding_macro_loss": 0.4518047733926306}131{"step": 290, "records_seen": 9280, "elapsed_seconds": 1550.0735397338867, "gradient_norm": 2.102567672729492, "learning_rate": 1.988496165530013e-05, "head_lr": 9.942480827650066e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.772564627502561, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.772564627502561, "supervised_tokens": 72.53125, "waves": 1.0, "detection_loss": 0.9971050770670574, "grounding_loss": 0.3438964965157952, "detection_grounding_macro_loss": 0.6705007867914263}132{"step": 300, "records_seen": 9600, "elapsed_seconds": 1603.47447347641, "gradient_norm": 1.7109359502792358, "learning_rate": 1.9871855819799993e-05, "head_lr": 9.935927909899996e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6942078879333167, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6942078879333167, "supervised_tokens": 85.53125, "waves": 1.0, "detection_loss": 0.9387241460858301, "grounding_loss": 0.4784582483869813, "detection_grounding_macro_loss": 0.7085911972364057}133/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead134  warnings.warn(135/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead136  warnings.warn(137/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead138  warnings.warn(139/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead140  warnings.warn(141/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead142  warnings.warn(143/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead144  warnings.warn(145/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead146  warnings.warn(147/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead148  warnings.warn(149{"validation_step": 300, "loss": 1.170851999606876, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.170851999606876, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.661455009329386, "grounding_loss": 0.6143680142930575, "detection_grounding_macro_loss": 1.1379115118112217}150{"step": 310, "records_seen": 9920, "elapsed_seconds": 1802.2109835147858, "gradient_norm": 1.416176438331604, "learning_rate": 1.9858048490100558e-05, "head_lr": 9.929024245050277e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.3893866845298817, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3893866845298817, "supervised_tokens": 88.6875, "waves": 1.0, "detection_loss": 0.5702221902947013, "grounding_loss": 0.26565712795395247, "detection_grounding_macro_loss": 0.4179396591243269}151{"step": 320, "records_seen": 10240, "elapsed_seconds": 1856.9637036323547, "gradient_norm": 1.7578338384628296, "learning_rate": 1.9843540757942022e-05, "head_lr": 9.92177037897101e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6233034284975929, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6233034284975929, "supervised_tokens": 91.40625, "waves": 1.0, "detection_loss": 0.8063417013342443, "grounding_loss": 0.3557859528132562, "detection_grounding_macro_loss": 0.5810638270737503}152{"step": 330, "records_seen": 10560, "elapsed_seconds": 1910.431179523468, "gradient_norm": 1.7179442644119263, "learning_rate": 1.9828333770445144e-05, "head_lr": 9.914166885222571e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.8069800072771613, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.8069800072771613, "supervised_tokens": 107.40625, "waves": 1.0, "detection_loss": 1.0925858243394435, "grounding_loss": 0.4397725281970842, "detection_grounding_macro_loss": 0.7661791762682638}153{"step": 340, "records_seen": 10880, "elapsed_seconds": 1963.1458320617676, "gradient_norm": 1.6289736032485962, "learning_rate": 1.981242873002053e-05, "head_lr": 9.906214365010263e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6119766440824606, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6119766440824606, "supervised_tokens": 109.78125, "waves": 1.0, "detection_loss": 0.7344514360541806, "grounding_loss": 0.4731718798478444, "detection_grounding_macro_loss": 0.6038116579510125}154{"step": 350, "records_seen": 11200, "elapsed_seconds": 2016.3297288417816, "gradient_norm": 1.9289132356643677, "learning_rate": 1.979582689427356e-05, "head_lr": 9.897913447136778e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7519244729774073, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7519244729774073, "supervised_tokens": 84.875, "waves": 1.0, "detection_loss": 0.9900307210260316, "grounding_loss": 0.4039230335217256, "detection_grounding_macro_loss": 0.6969768772738786}155{"step": 360, "records_seen": 11520, "elapsed_seconds": 2070.4591591358185, "gradient_norm": 1.6302134990692139, "learning_rate": 1.9778529575904938e-05, "head_lr": 9.88926478795247e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6829727828569503, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6829727828569503, "supervised_tokens": 100.0625, "waves": 1.0, "detection_loss": 0.8769071283972557, "grounding_loss": 0.43362862430512905, "detection_grounding_macro_loss": 0.6552678763511923}156{"step": 370, "records_seen": 11840, "elapsed_seconds": 2123.5187566280365, "gradient_norm": 1.684336543083191, "learning_rate": 1.9760538142606933e-05, "head_lr": 9.880269071303465e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5728858088857578, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5728858088857578, "supervised_tokens": 95.71875, "waves": 1.0, "detection_loss": 0.9546009515013013, "grounding_loss": 0.2759962535181128, "detection_grounding_macro_loss": 0.615298602509707}157{"step": 380, "records_seen": 12160, "elapsed_seconds": 2174.8912365436554, "gradient_norm": 1.639719009399414, "learning_rate": 1.9741854016955193e-05, "head_lr": 9.870927008477595e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5436957220196064, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5436957220196064, "supervised_tokens": 83.34375, "waves": 1.0, "detection_loss": 0.7788904443150386, "grounding_loss": 0.3085009997241741, "detection_grounding_macro_loss": 0.5436957220196064}158{"step": 390, "records_seen": 12480, "elapsed_seconds": 2226.6730918884277, "gradient_norm": 1.7520829439163208, "learning_rate": 1.972247867629629e-05, "head_lr": 9.861239338148142e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7044163011771616, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7044163011771616, "supervised_tokens": 101.9375, "waves": 1.0, "detection_loss": 0.9515803104473485, "grounding_loss": 0.38663400354406413, "detection_grounding_macro_loss": 0.6691071569957063}159{"step": 400, "records_seen": 12800, "elapsed_seconds": 2281.1013045310974, "gradient_norm": 1.6985821723937988, "learning_rate": 1.9702413652630892e-05, "head_lr": 9.851206826315445e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.49458120221970603, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49458120221970603, "supervised_tokens": 113.0, "waves": 1.0, "detection_loss": 0.5574173317808244, "grounding_loss": 0.41379189278398243, "detection_grounding_macro_loss": 0.4856046122824034}160{"step": 410, "records_seen": 13120, "elapsed_seconds": 2333.9911036491394, "gradient_norm": 1.7086955308914185, "learning_rate": 1.9681660532492645e-05, "head_lr": 9.84083026624632e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5580460397231946, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5580460397231946, "supervised_tokens": 123.9375, "waves": 1.0, "detection_loss": 0.556304565847515, "grounding_loss": 0.560285077563354, "detection_grounding_macro_loss": 0.5582948217054344}161{"step": 420, "records_seen": 13440, "elapsed_seconds": 2388.0083405971527, "gradient_norm": 1.7172855138778687, "learning_rate": 1.9660220956822705e-05, "head_lr": 9.830110478411352e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6777725365575407, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6777725365575407, "supervised_tokens": 134.0, "waves": 1.0, "detection_loss": 0.8484224209207154, "grounding_loss": 0.3519863936823888, "detection_grounding_macro_loss": 0.6002044073015521}162{"step": 430, "records_seen": 13760, "elapsed_seconds": 2441.14422416687, "gradient_norm": 1.613144874572754, "learning_rate": 1.9638096620840012e-05, "head_lr": 9.819048310420005e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.541783322363699, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.541783322363699, "supervised_tokens": 109.6875, "waves": 1.0, "detection_loss": 0.6403079002856329, "grounding_loss": 0.41510886503549826, "detection_grounding_macro_loss": 0.5277083826605655}163{"step": 440, "records_seen": 14080, "elapsed_seconds": 2495.4878787994385, "gradient_norm": 1.7455748319625854, "learning_rate": 1.9615289273907228e-05, "head_lr": 9.807644636953613e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6144793150109535, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6144793150109535, "supervised_tokens": 112.75, "waves": 1.0, "detection_loss": 0.7695680751520044, "grounding_loss": 0.4387120535177625, "detection_grounding_macro_loss": 0.6041400643348834}164{"step": 450, "records_seen": 14400, "elapsed_seconds": 2549.4548058509827, "gradient_norm": 2.3900609016418457, "learning_rate": 1.9591800719392432e-05, "head_lr": 9.795900359696216e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7283700105381286, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7283700105381286, "supervised_tokens": 80.53125, "waves": 1.0, "detection_loss": 0.9826157430186868, "grounding_loss": 0.3046271230705315, "detection_grounding_macro_loss": 0.6436214330446092}165{"step": 460, "records_seen": 14720, "elapsed_seconds": 2603.157094478607, "gradient_norm": 1.4154752492904663, "learning_rate": 1.9567632814526523e-05, "head_lr": 9.78381640726326e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5570307351881638, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5570307351881638, "supervised_tokens": 136.40625, "waves": 1.0, "detection_loss": 0.6550533742728558, "grounding_loss": 0.3413809292018414, "detection_grounding_macro_loss": 0.4982171517373486}166{"step": 470, "records_seen": 15040, "elapsed_seconds": 2658.915317773819, "gradient_norm": 1.6542103290557861, "learning_rate": 1.9542787470256365e-05, "head_lr": 9.771393735128182e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.598170449452823, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.598170449452823, "supervised_tokens": 91.8125, "waves": 1.0, "detection_loss": 0.8898128823909376, "grounding_loss": 0.3713374460565117, "detection_grounding_macro_loss": 0.6305751642237246}167{"step": 480, "records_seen": 15360, "elapsed_seconds": 2710.9187483787537, "gradient_norm": 1.861283302307129, "learning_rate": 1.9517266651093695e-05, "head_lr": 9.758633325546846e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7130575161099841, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7130575161099841, "supervised_tokens": 100.34375, "waves": 1.0, "detection_loss": 0.9799664653182845, "grounding_loss": 0.4461485669016838, "detection_grounding_macro_loss": 0.7130575161099841}168{"step": 490, "records_seen": 15680, "elapsed_seconds": 2765.7663736343384, "gradient_norm": 1.7210060358047485, "learning_rate": 1.949107237495979e-05, "head_lr": 9.745536187479894e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.627728114835918, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.627728114835918, "supervised_tokens": 89.5, "waves": 1.0, "detection_loss": 0.9663133729587902, "grounding_loss": 0.45037393200965153, "detection_grounding_macro_loss": 0.7083436524842208}169{"step": 500, "records_seen": 16000, "elapsed_seconds": 2819.2090792655945, "gradient_norm": 1.2247434854507446, "learning_rate": 1.94642067130259e-05, "head_lr": 9.732103356512948e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5270388060962432, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5270388060962432, "supervised_tokens": 150.65625, "waves": 1.0, "detection_loss": 0.6078185594039193, "grounding_loss": 0.2846995461732149, "detection_grounding_macro_loss": 0.4462590527885671}170{"step": 510, "records_seen": 16320, "elapsed_seconds": 2872.3118188381195, "gradient_norm": 1.4203404188156128, "learning_rate": 1.94366717895495e-05, "head_lr": 9.71833589477475e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5003671738706998, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5003671738706998, "supervised_tokens": 101.09375, "waves": 1.0, "detection_loss": 0.7031869863936057, "grounding_loss": 0.32140851576225354, "detection_grounding_macro_loss": 0.5122977510779296}171{"step": 520, "records_seen": 16640, "elapsed_seconds": 2925.0624918937683, "gradient_norm": 1.5626354217529297, "learning_rate": 1.9408469781706315e-05, "head_lr": 9.704234890853157e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5077007749557652, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5077007749557652, "supervised_tokens": 107.5, "waves": 1.0, "detection_loss": 0.6985878734849393, "grounding_loss": 0.316813676426591, "detection_grounding_macro_loss": 0.5077007749557652}172{"step": 530, "records_seen": 16960, "elapsed_seconds": 2976.6627197265625, "gradient_norm": 1.4163010120391846, "learning_rate": 1.9379602919418166e-05, "head_lr": 9.689801459709082e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.3946421748444209, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3946421748444209, "supervised_tokens": 94.03125, "waves": 1.0, "detection_loss": 0.5125143965281799, "grounding_loss": 0.2906372733587512, "detection_grounding_macro_loss": 0.40157583494346555}173{"step": 540, "records_seen": 17280, "elapsed_seconds": 3028.233227491379, "gradient_norm": 1.9121947288513184, "learning_rate": 1.9350073485176656e-05, "head_lr": 9.675036742588326e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7096727220741741, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7096727220741741, "supervised_tokens": 128.875, "waves": 1.0, "detection_loss": 0.910378211949903, "grounding_loss": 0.41633392917887807, "detection_grounding_macro_loss": 0.6633560705643905}174{"step": 550, "records_seen": 17600, "elapsed_seconds": 3079.1619153022766, "gradient_norm": 1.7950783967971802, "learning_rate": 1.9319883813862705e-05, "head_lr": 9.65994190693135e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.6086413051816635, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6086413051816635, "supervised_tokens": 84.1875, "waves": 1.0, "detection_loss": 1.0265899069296818, "grounding_loss": 0.35787214413285257, "detection_grounding_macro_loss": 0.6922310255312671}175{"step": 560, "records_seen": 17920, "elapsed_seconds": 3131.770822286606, "gradient_norm": 1.7444860935211182, "learning_rate": 1.9289036292561913e-05, "head_lr": 9.644518146280956e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.591612036321294, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.591612036321294, "supervised_tokens": 118.5625, "waves": 1.0, "detection_loss": 0.6756852881444502, "grounding_loss": 0.4835178554058075, "detection_grounding_macro_loss": 0.5796015717751288}176{"step": 570, "records_seen": 18240, "elapsed_seconds": 3186.256816148758, "gradient_norm": 2.0421202182769775, "learning_rate": 1.925753336037583e-05, "head_lr": 9.628766680187914e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.7410425379166554, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7410425379166554, "supervised_tokens": 107.65625, "waves": 1.0, "detection_loss": 1.0081090555690668, "grounding_loss": 0.4383671512439226, "detection_grounding_macro_loss": 0.7232381034064947}177{"step": 580, "records_seen": 18560, "elapsed_seconds": 3239.1071560382843, "gradient_norm": 1.6931291818618774, "learning_rate": 1.9225377508229087e-05, "head_lr": 9.612688754114542e-07, "peak_gpu_gib": 19.74146795272827, "loss": 0.5473567897342946, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5473567897342946, "supervised_tokens": 123.09375, "waves": 1.0, "detection_loss": 0.7304254782057422, "grounding_loss": 0.33987894279998726, "detection_grounding_macro_loss": 0.5351522105028648}178{"step": 590, "records_seen": 18880, "elapsed_seconds": 3291.968508005142, "gradient_norm": 1.5388755798339844, "learning_rate": 1.919257127867244e-05, "head_lr": 9.59628563933622e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6298742308330247, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6298742308330247, "supervised_tokens": 122.59375, "waves": 1.0, "detection_loss": 0.8098717852320988, "grounding_loss": 0.32987830683456804, "detection_grounding_macro_loss": 0.5698750460333334}179{"step": 600, "records_seen": 19200, "elapsed_seconds": 3343.8341884613037, "gradient_norm": 1.7624316215515137, "learning_rate": 1.9159117265681744e-05, "head_lr": 9.57955863284087e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5808190598290821, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5808190598290821, "supervised_tokens": 83.78125, "waves": 1.0, "detection_loss": 0.867364594315101, "grounding_loss": 0.40889173913747073, "detection_grounding_macro_loss": 0.6381281667262859}180/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead181  warnings.warn(182/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead183  warnings.warn(184/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead185  warnings.warn(186/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead187  warnings.warn(188/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead189  warnings.warn(190/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead191  warnings.warn(192/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead193  warnings.warn(194/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead195  warnings.warn(196{"validation_step": 600, "loss": 1.1494046881351905, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1494046881351905, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.649928241070298, "grounding_loss": 0.5816679723773683, "detection_grounding_macro_loss": 1.1157981067238332}197{"step": 610, "records_seen": 19520, "elapsed_seconds": 3538.4205520153046, "gradient_norm": 1.572460651397705, "learning_rate": 1.912501811445283e-05, "head_lr": 9.562509057226415e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5462376191826479, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5462376191826479, "supervised_tokens": 123.9375, "waves": 1.0, "detection_loss": 0.5294467196003289, "grounding_loss": 0.5707781647260373, "detection_grounding_macro_loss": 0.5501124421631831}198{"step": 620, "records_seen": 19840, "elapsed_seconds": 3592.1801357269287, "gradient_norm": 1.883591890335083, "learning_rate": 1.9090276521192363e-05, "head_lr": 9.545138260596179e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.804820184398352, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.804820184398352, "supervised_tokens": 108.3125, "waves": 1.0, "detection_loss": 0.9916442196313668, "grounding_loss": 0.3938073068857193, "detection_grounding_macro_loss": 0.6927257632585431}199{"step": 630, "records_seen": 20160, "elapsed_seconds": 3644.5383105278015, "gradient_norm": 1.6253025531768799, "learning_rate": 1.9054895232904647e-05, "head_lr": 9.527447616452322e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5616307263667295, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5616307263667295, "supervised_tokens": 100.4375, "waves": 1.0, "detection_loss": 0.8807909509965351, "grounding_loss": 0.3133949960991029, "detection_grounding_macro_loss": 0.597092973547819}200{"step": 640, "records_seen": 20480, "elapsed_seconds": 3699.661168575287, "gradient_norm": 1.578477144241333, "learning_rate": 1.901887704717443e-05, "head_lr": 9.509438523587215e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4929434312432477, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4929434312432477, "supervised_tokens": 105.09375, "waves": 1.0, "detection_loss": 0.6130006741200305, "grounding_loss": 0.3728861883664649, "detection_grounding_macro_loss": 0.4929434312432477}201{"step": 650, "records_seen": 20800, "elapsed_seconds": 3751.1523883342743, "gradient_norm": 1.5485762357711792, "learning_rate": 1.8982224811945697e-05, "head_lr": 9.491112405972846e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6020811545247966, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6020811545247966, "supervised_tokens": 123.65625, "waves": 1.0, "detection_loss": 0.8043636419461109, "grounding_loss": 0.26494367548927283, "detection_grounding_macro_loss": 0.5346536587176919}202{"step": 660, "records_seen": 21120, "elapsed_seconds": 3804.3779888153076, "gradient_norm": 1.8320295810699463, "learning_rate": 1.8944941425296457e-05, "head_lr": 9.472470712648227e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6229781178352596, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6229781178352596, "supervised_tokens": 69.34375, "waves": 1.0, "detection_loss": 1.049553990761827, "grounding_loss": 0.3670325940793191, "detection_grounding_macro_loss": 0.708293292420573}203{"step": 670, "records_seen": 21440, "elapsed_seconds": 3860.3394882678986, "gradient_norm": 1.5231282711029053, "learning_rate": 1.8907029835209648e-05, "head_lr": 9.453514917604824e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6050457115312611, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6050457115312611, "supervised_tokens": 133.40625, "waves": 1.0, "detection_loss": 0.7317920364546895, "grounding_loss": 0.36307545485926146, "detection_grounding_macro_loss": 0.5474337456569754}204{"step": 680, "records_seen": 21760, "elapsed_seconds": 3915.045831680298, "gradient_norm": 1.5743012428283691, "learning_rate": 1.886849303934e-05, "head_lr": 9.434246519669998e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5447211839292265, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5447211839292265, "supervised_tokens": 114.90625, "waves": 1.0, "detection_loss": 0.7280351765453815, "grounding_loss": 0.38297354338556033, "detection_grounding_macro_loss": 0.555504359965471}205{"step": 690, "records_seen": 22080, "elapsed_seconds": 3965.633773326874, "gradient_norm": 1.6212862730026245, "learning_rate": 1.8829334084777003e-05, "head_lr": 9.4146670423885e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5115128003890277, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5115128003890277, "supervised_tokens": 118.75, "waves": 1.0, "detection_loss": 0.8244463379184405, "grounding_loss": 0.23539497315719285, "detection_grounding_macro_loss": 0.5299206555378166}206{"step": 700, "records_seen": 22400, "elapsed_seconds": 4017.0098242759705, "gradient_norm": 1.801900029182434, "learning_rate": 1.878955606780402e-05, "head_lr": 9.394778033902009e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6754185676895759, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6754185676895759, "supervised_tokens": 96.21875, "waves": 1.0, "detection_loss": 0.8259480768210778, "grounding_loss": 0.4245360524704059, "detection_grounding_macro_loss": 0.6252420646457418}207{"step": 710, "records_seen": 22720, "elapsed_seconds": 4067.7231907844543, "gradient_norm": 1.5854816436767578, "learning_rate": 1.874916213365343e-05, "head_lr": 9.374581066826714e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5452691993523331, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5452691993523331, "supervised_tokens": 101.46875, "waves": 1.0, "detection_loss": 0.6481747048175228, "grounding_loss": 0.34881323437333445, "detection_grounding_macro_loss": 0.49849396959542863}208{"step": 720, "records_seen": 23040, "elapsed_seconds": 4120.058799982071, "gradient_norm": 1.864081859588623, "learning_rate": 1.870815547625793e-05, "head_lr": 9.354077738128964e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.567607521571631, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.567607521571631, "supervised_tokens": 87.65625, "waves": 1.0, "detection_loss": 0.7997107073664665, "grounding_loss": 0.46210607348306937, "detection_grounding_macro_loss": 0.630908390424768}209{"step": 730, "records_seen": 23360, "elapsed_seconds": 4172.731770515442, "gradient_norm": 1.75831937789917, "learning_rate": 1.8666539337998033e-05, "head_lr": 9.333269668999016e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.623487599250808, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.623487599250808, "supervised_tokens": 97.5625, "waves": 1.0, "detection_loss": 1.0220410617385565, "grounding_loss": 0.31350157287144814, "detection_grounding_macro_loss": 0.6677713173050023}210{"step": 740, "records_seen": 23680, "elapsed_seconds": 4226.529960870743, "gradient_norm": 1.6202095746994019, "learning_rate": 1.8624317009445644e-05, "head_lr": 9.31215850472282e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4545965629941122, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4545965629941122, "supervised_tokens": 101.78125, "waves": 1.0, "detection_loss": 0.5906396688962255, "grounding_loss": 0.2796839982628236, "detection_grounding_macro_loss": 0.43516183357952454}211{"step": 750, "records_seen": 24000, "elapsed_seconds": 4280.862471342087, "gradient_norm": 1.6980353593826294, "learning_rate": 1.858149182910391e-05, "head_lr": 9.290745914551954e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.43145475877656736, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.43145475877656736, "supervised_tokens": 95.0, "waves": 1.0, "detection_loss": 0.5671912157400087, "grounding_loss": 0.3603547098909552, "detection_grounding_macro_loss": 0.46377296281548197}212{"step": 760, "records_seen": 24320, "elapsed_seconds": 4333.311567544937, "gradient_norm": 1.7356544733047485, "learning_rate": 1.8538067183143232e-05, "head_lr": 9.269033591571615e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5211210656206049, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5211210656206049, "supervised_tokens": 112.78125, "waves": 1.0, "detection_loss": 0.5787445473002403, "grounding_loss": 0.4763028020919996, "detection_grounding_macro_loss": 0.5275236746961199}213{"step": 770, "records_seen": 24640, "elapsed_seconds": 4386.234909057617, "gradient_norm": 1.586024284362793, "learning_rate": 1.8494046505133535e-05, "head_lr": 9.247023252566767e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.49081307859160006, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49081307859160006, "supervised_tokens": 125.46875, "waves": 1.0, "detection_loss": 0.6210426951448123, "grounding_loss": 0.3759045933975893, "detection_grounding_macro_loss": 0.4984736442712008}214{"step": 780, "records_seen": 24960, "elapsed_seconds": 4436.2440531253815, "gradient_norm": 1.623769998550415, "learning_rate": 1.844943327577275e-05, "head_lr": 9.224716637886373e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.683967698101128, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.683967698101128, "supervised_tokens": 106.0625, "waves": 1.0, "detection_loss": 0.9384122049874243, "grounding_loss": 0.3955972569633256, "detection_grounding_macro_loss": 0.6670047309753749}215{"step": 790, "records_seen": 25280, "elapsed_seconds": 4487.890097856522, "gradient_norm": 1.6542013883590698, "learning_rate": 1.8404231022611623e-05, "head_lr": 9.20211551130581e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.518584670138523, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.518584670138523, "supervised_tokens": 97.65625, "waves": 1.0, "detection_loss": 0.6556575253648286, "grounding_loss": 0.3976380331741358, "detection_grounding_macro_loss": 0.5266477792694821}216{"step": 800, "records_seen": 25600, "elapsed_seconds": 4542.494109392166, "gradient_norm": 1.457810878753662, "learning_rate": 1.8358443319774787e-05, "head_lr": 9.179221659887393e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5627536715881263, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5627536715881263, "supervised_tokens": 140.1875, "waves": 1.0, "detection_loss": 0.6971618110271829, "grounding_loss": 0.30615631447720015, "detection_grounding_macro_loss": 0.5016590627521915}217{"step": 810, "records_seen": 25920, "elapsed_seconds": 4595.534882068634, "gradient_norm": 1.6737902164459229, "learning_rate": 1.8312073787678147e-05, "head_lr": 9.156036893839072e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5325827514757293, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5325827514757293, "supervised_tokens": 104.28125, "waves": 1.0, "detection_loss": 0.7386745872596899, "grounding_loss": 0.4089276500053529, "detection_grounding_macro_loss": 0.5738011186325214}218{"step": 820, "records_seen": 26240, "elapsed_seconds": 4652.505940914154, "gradient_norm": 1.825777292251587, "learning_rate": 1.8265126092742623e-05, "head_lr": 9.132563046371311e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5308753344904744, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5308753344904744, "supervised_tokens": 92.78125, "waves": 1.0, "detection_loss": 0.553035094383701, "grounding_loss": 0.5157133935108983, "detection_grounding_macro_loss": 0.5343742439472996}219{"step": 830, "records_seen": 26560, "elapsed_seconds": 4706.387438297272, "gradient_norm": 1.5823073387145996, "learning_rate": 1.8217603947104254e-05, "head_lr": 9.108801973552125e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4378069230137953, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4378069230137953, "supervised_tokens": 121.28125, "waves": 1.0, "detection_loss": 0.5066584945994287, "grounding_loss": 0.34928347383226666, "detection_grounding_macro_loss": 0.42797098421584767}220{"step": 840, "records_seen": 26880, "elapsed_seconds": 4762.393179893494, "gradient_norm": 1.6081262826919556, "learning_rate": 1.816951110832066e-05, "head_lr": 9.084755554160328e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5010699396516429, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5010699396516429, "supervised_tokens": 109.65625, "waves": 1.0, "detection_loss": 0.6219651174679812, "grounding_loss": 0.34563328245920794, "detection_grounding_macro_loss": 0.48379919996359455}221{"step": 850, "records_seen": 27200, "elapsed_seconds": 4816.64536523819, "gradient_norm": 1.4766751527786255, "learning_rate": 1.8120851379073956e-05, "head_lr": 9.060425689536977e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.49292759962304444, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49292759962304444, "supervised_tokens": 121.03125, "waves": 1.0, "detection_loss": 0.5711002889867813, "grounding_loss": 0.36263978401681624, "detection_grounding_macro_loss": 0.4668700365017988}222{"step": 860, "records_seen": 27520, "elapsed_seconds": 4869.129330396652, "gradient_norm": 1.7807718515396118, "learning_rate": 1.8071628606870065e-05, "head_lr": 9.035814303435032e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.657707178738292, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.657707178738292, "supervised_tokens": 128.75, "waves": 1.0, "detection_loss": 0.8545641628507938, "grounding_loss": 0.40460534202221815, "detection_grounding_macro_loss": 0.629584752436506}223{"step": 870, "records_seen": 27840, "elapsed_seconds": 4922.414959192276, "gradient_norm": 1.5774455070495605, "learning_rate": 1.8021846683734502e-05, "head_lr": 9.010923341867249e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4814808277092055, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4814808277092055, "supervised_tokens": 95.1875, "waves": 1.0, "detection_loss": 0.5778695427686216, "grounding_loss": 0.39643196148030896, "detection_grounding_macro_loss": 0.48715075212446524}224{"step": 880, "records_seen": 28160, "elapsed_seconds": 4975.641998767853, "gradient_norm": 1.5952253341674805, "learning_rate": 1.7971509545904617e-05, "head_lr": 8.985754772952307e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5212242758242951, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5212242758242951, "supervised_tokens": 75.65625, "waves": 1.0, "detection_loss": 0.680145353469665, "grounding_loss": 0.39761899321122957, "detection_grounding_macro_loss": 0.5388821733404473}225{"step": 890, "records_seen": 28480, "elapsed_seconds": 5027.481950044632, "gradient_norm": 1.8138006925582886, "learning_rate": 1.792062117351838e-05, "head_lr": 8.960310586759187e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.8326126406900585, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.8326126406900585, "supervised_tokens": 117.96875, "waves": 1.0, "detection_loss": 1.1190589174628258, "grounding_loss": 0.35520217940211296, "detection_grounding_macro_loss": 0.7371305484324694}226{"step": 900, "records_seen": 28800, "elapsed_seconds": 5079.391407728195, "gradient_norm": 1.4787707328796387, "learning_rate": 1.7869185590299664e-05, "head_lr": 8.934592795149831e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.3998298544579484, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3998298544579484, "supervised_tokens": 127.15625, "waves": 1.0, "detection_loss": 0.4880377152003348, "grounding_loss": 0.311621993715562, "detection_grounding_macro_loss": 0.3998298544579484}227/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead228  warnings.warn(229/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead230  warnings.warn(231/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead232  warnings.warn(233/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead234  warnings.warn(235/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead236  warnings.warn(237/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead238  warnings.warn(239/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead240  warnings.warn(241/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead242  warnings.warn(243{"validation_step": 900, "loss": 1.1422257937647153, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1422257937647153, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6497111428569602, "grounding_loss": 0.5665924120800836, "detection_grounding_macro_loss": 1.108151777468522}244{"step": 910, "records_seen": 29120, "elapsed_seconds": 5274.754962682724, "gradient_norm": 1.5693084001541138, "learning_rate": 1.7817206863240085e-05, "head_lr": 8.908603431620042e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.42089124511767295, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.42089124511767295, "supervised_tokens": 94.4375, "waves": 1.0, "detection_loss": 0.35508511142494775, "grounding_loss": 0.47895548072890104, "detection_grounding_macro_loss": 0.41702029607692437}245{"step": 920, "records_seen": 29440, "elapsed_seconds": 5326.703498601913, "gradient_norm": 1.6184418201446533, "learning_rate": 1.7764689102277442e-05, "head_lr": 8.88234455113872e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5312091871455777, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5312091871455777, "supervised_tokens": 82.53125, "waves": 1.0, "detection_loss": 0.6622392977587879, "grounding_loss": 0.42929687889085877, "detection_grounding_macro_loss": 0.5457680883248233}246{"step": 930, "records_seen": 29760, "elapsed_seconds": 5382.334538221359, "gradient_norm": 1.7035223245620728, "learning_rate": 1.7711636459970722e-05, "head_lr": 8.85581822998536e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5776664193813303, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5776664193813303, "supervised_tokens": 99.28125, "waves": 1.0, "detection_loss": 0.7360348508656898, "grounding_loss": 0.31371903357406455, "detection_grounding_macro_loss": 0.5248769422198771}247{"step": 940, "records_seen": 30080, "elapsed_seconds": 5436.324367761612, "gradient_norm": 1.9481412172317505, "learning_rate": 1.765805313117178e-05, "head_lr": 8.829026565585888e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4511098947493224, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4511098947493224, "supervised_tokens": 88.59375, "waves": 1.0, "detection_loss": 0.5746019230664388, "grounding_loss": 0.1794274324516664, "detection_grounding_macro_loss": 0.37701467775905256}248{"step": 950, "records_seen": 30400, "elapsed_seconds": 5488.073384523392, "gradient_norm": 1.684754490852356, "learning_rate": 1.760394335269365e-05, "head_lr": 8.801971676346825e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.7693636112905438, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7693636112905438, "supervised_tokens": 122.875, "waves": 1.0, "detection_loss": 1.0494775166735053, "grounding_loss": 0.4892497059075822, "detection_grounding_macro_loss": 0.7693636112905438}249{"step": 960, "records_seen": 30720, "elapsed_seconds": 5543.760031700134, "gradient_norm": 1.7524369955062866, "learning_rate": 1.7549311402975526e-05, "head_lr": 8.774655701487762e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6156204381672978, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6156204381672978, "supervised_tokens": 109.96875, "waves": 1.0, "detection_loss": 0.7336240025247925, "grounding_loss": 0.44315369026019025, "detection_grounding_macro_loss": 0.5883888463924913}250{"step": 970, "records_seen": 31040, "elapsed_seconds": 5597.490339517593, "gradient_norm": 1.675694227218628, "learning_rate": 1.7494161601744494e-05, "head_lr": 8.747080800872245e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5449616849311951, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5449616849311951, "supervised_tokens": 109.78125, "waves": 1.0, "detection_loss": 0.7073809398964787, "grounding_loss": 0.3608865293038737, "detection_grounding_macro_loss": 0.5341337346001762}251{"step": 980, "records_seen": 31360, "elapsed_seconds": 5650.5555810928345, "gradient_norm": 1.7019107341766357, "learning_rate": 1.7438498309673945e-05, "head_lr": 8.719249154836971e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6466299799635635, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6466299799635635, "supervised_tokens": 121.28125, "waves": 1.0, "detection_loss": 0.8754391771872179, "grounding_loss": 0.3524467263902937, "detection_grounding_macro_loss": 0.6139429517887558}252{"step": 990, "records_seen": 31680, "elapsed_seconds": 5702.062856912613, "gradient_norm": 1.713455319404602, "learning_rate": 1.7382325928038804e-05, "head_lr": 8.691162964019401e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6690564370214815, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6690564370214815, "supervised_tokens": 106.90625, "waves": 1.0, "detection_loss": 0.8843003170306261, "grounding_loss": 0.4251133730111178, "detection_grounding_macro_loss": 0.654706845020872}253{"step": 1000, "records_seen": 32000, "elapsed_seconds": 5752.344634056091, "gradient_norm": 1.726184606552124, "learning_rate": 1.7325648898367498e-05, "head_lr": 8.662824449183747e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.42941988657827324, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.42941988657827324, "supervised_tokens": 92.09375, "waves": 1.0, "detection_loss": 0.5708663617352661, "grounding_loss": 0.22269042288728366, "detection_grounding_macro_loss": 0.3967783923112749}254{"step": 1010, "records_seen": 32320, "elapsed_seconds": 5806.119509458542, "gradient_norm": 1.4949272871017456, "learning_rate": 1.7268471702090787e-05, "head_lr": 8.634235851045392e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5501604690740578, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5501604690740578, "supervised_tokens": 130.6875, "waves": 1.0, "detection_loss": 0.7328000759635439, "grounding_loss": 0.28322565900480884, "detection_grounding_macro_loss": 0.5080128674841764}255{"step": 1020, "records_seen": 32640, "elapsed_seconds": 5860.196108341217, "gradient_norm": 1.8133139610290527, "learning_rate": 1.721079886018741e-05, "head_lr": 8.605399430093703e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6961922198424872, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6961922198424872, "supervised_tokens": 116.5, "waves": 1.0, "detection_loss": 0.8251663182110694, "grounding_loss": 0.4124492034316063, "detection_grounding_macro_loss": 0.6188077608213378}256{"step": 1030, "records_seen": 32960, "elapsed_seconds": 5914.152135848999, "gradient_norm": 1.5873194932937622, "learning_rate": 1.715263493282661e-05, "head_lr": 8.576317466413303e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.46789829818521866, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.46789829818521866, "supervised_tokens": 117.84375, "waves": 1.0, "detection_loss": 0.5299854206154123, "grounding_loss": 0.4131155430997537, "detection_grounding_macro_loss": 0.47155048185758297}257{"step": 1040, "records_seen": 33280, "elapsed_seconds": 5968.904947042465, "gradient_norm": 1.3768315315246582, "learning_rate": 1.7093984519007565e-05, "head_lr": 8.546992259503782e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.41974458677447046, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.41974458677447046, "supervised_tokens": 132.9375, "waves": 1.0, "detection_loss": 0.4740934565763217, "grounding_loss": 0.30017707321039777, "detection_grounding_macro_loss": 0.3871352648933597}258{"step": 1050, "records_seen": 33600, "elapsed_seconds": 6021.591591835022, "gradient_norm": 1.573705792427063, "learning_rate": 1.703485225619576e-05, "head_lr": 8.517426128097879e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5153993056160289, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5153993056160289, "supervised_tokens": 92.375, "waves": 1.0, "detection_loss": 0.6519155741824458, "grounding_loss": 0.39494377452801394, "detection_grounding_macro_loss": 0.5234296743552299}259{"step": 1060, "records_seen": 33920, "elapsed_seconds": 6072.6659944057465, "gradient_norm": 1.7577590942382812, "learning_rate": 1.6975242819956277e-05, "head_lr": 8.487621409978137e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.716720276424212, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.716720276424212, "supervised_tokens": 97.9375, "waves": 1.0, "detection_loss": 0.8762490466328422, "grounding_loss": 0.5116118575845446, "detection_grounding_macro_loss": 0.6939304521086934}260{"step": 1070, "records_seen": 34240, "elapsed_seconds": 6123.2896230220795, "gradient_norm": 1.4944669008255005, "learning_rate": 1.6915160923584127e-05, "head_lr": 8.457580461792062e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.46306866549474535, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.46306866549474535, "supervised_tokens": 141.875, "waves": 1.0, "detection_loss": 0.6247695125639439, "grounding_loss": 0.32039144749251136, "detection_grounding_macro_loss": 0.4725804800282276}261{"step": 1080, "records_seen": 34560, "elapsed_seconds": 6177.1465475559235, "gradient_norm": 1.6511389017105103, "learning_rate": 1.685461131773156e-05, "head_lr": 8.427305658865779e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6133561500574274, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6133561500574274, "supervised_tokens": 95.34375, "waves": 1.0, "detection_loss": 0.8074245067706215, "grounding_loss": 0.4421193647222561, "detection_grounding_macro_loss": 0.6247719357464387}262{"step": 1090, "records_seen": 34880, "elapsed_seconds": 6231.01535153389, "gradient_norm": 1.8518998622894287, "learning_rate": 1.6793598790032423e-05, "head_lr": 8.396799395016211e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6394909012103653, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6394909012103653, "supervised_tokens": 103.65625, "waves": 1.0, "detection_loss": 0.7584724965353316, "grounding_loss": 0.4865145643639802, "detection_grounding_macro_loss": 0.6224935304496559}263{"step": 1100, "records_seen": 35200, "elapsed_seconds": 6284.290812253952, "gradient_norm": 1.5828895568847656, "learning_rate": 1.6732128164723634e-05, "head_lr": 8.366064082361816e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.3846743688463903, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3846743688463903, "supervised_tokens": 69.375, "waves": 1.0, "detection_loss": 0.5155316106354196, "grounding_loss": 0.2692120966795996, "detection_grounding_macro_loss": 0.3923718536575096}264{"step": 1110, "records_seen": 35520, "elapsed_seconds": 6336.668553113937, "gradient_norm": 1.730488896369934, "learning_rate": 1.6670204302263686e-05, "head_lr": 8.335102151131842e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.3140978772580638, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3140978772580638, "supervised_tokens": 134.0625, "waves": 1.0, "detection_loss": 0.3142559625781499, "grounding_loss": 0.3139897136180049, "detection_grounding_macro_loss": 0.3141228380980774}265{"step": 1120, "records_seen": 35840, "elapsed_seconds": 6389.024989366531, "gradient_norm": 1.563281536102295, "learning_rate": 1.6607832098948374e-05, "head_lr": 8.303916049474187e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.44586215522213024, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.44586215522213024, "supervised_tokens": 125.59375, "waves": 1.0, "detection_loss": 0.5101444286119658, "grounding_loss": 0.3387250329057376, "detection_grounding_macro_loss": 0.42443473075885174}266{"step": 1130, "records_seen": 36160, "elapsed_seconds": 6441.8269119262695, "gradient_norm": 1.5285310745239258, "learning_rate": 1.6545016486523635e-05, "head_lr": 8.272508243261816e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.48533640839730197, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.48533640839730197, "supervised_tokens": 105.4375, "waves": 1.0, "detection_loss": 0.5718259033682342, "grounding_loss": 0.3741356291489605, "detection_grounding_macro_loss": 0.47298076625859736}267{"step": 1140, "records_seen": 36480, "elapsed_seconds": 6493.75359749794, "gradient_norm": 1.923203468322754, "learning_rate": 1.6481762431795575e-05, "head_lr": 8.240881215897787e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.43754788067053596, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.43754788067053596, "supervised_tokens": 104.90625, "waves": 1.0, "detection_loss": 0.6424597788469067, "grounding_loss": 0.278171959866692, "detection_grounding_macro_loss": 0.46031586935679936}268{"step": 1150, "records_seen": 36800, "elapsed_seconds": 6546.915998458862, "gradient_norm": 1.6859838962554932, "learning_rate": 1.641807493623778e-05, "head_lr": 8.209037468118888e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5949684355873615, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5949684355873615, "supervised_tokens": 97.46875, "waves": 1.0, "detection_loss": 0.9462913274765015, "grounding_loss": 0.3217172974513637, "detection_grounding_macro_loss": 0.6340043124639325}269{"step": 1160, "records_seen": 37120, "elapsed_seconds": 6599.535498142242, "gradient_norm": 1.9911842346191406, "learning_rate": 1.6353959035595815e-05, "head_lr": 8.176979517797907e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6444852547123219, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6444852547123219, "supervised_tokens": 89.375, "waves": 1.0, "detection_loss": 0.8294073144594828, "grounding_loss": 0.40672832075168636, "detection_grounding_macro_loss": 0.6180678176055846}270{"step": 1170, "records_seen": 37440, "elapsed_seconds": 6653.85995554924, "gradient_norm": 2.0541539192199707, "learning_rate": 1.6289419799489092e-05, "head_lr": 8.144709899744545e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.615012417562184, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.615012417562184, "supervised_tokens": 85.65625, "waves": 1.0, "detection_loss": 0.869137190316126, "grounding_loss": 0.4173598165313403, "detection_grounding_macro_loss": 0.6432485034237332}271{"step": 1180, "records_seen": 37760, "elapsed_seconds": 6706.658569812775, "gradient_norm": 1.9011735916137695, "learning_rate": 1.622446233100998e-05, "head_lr": 8.11223116550499e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6827347888702207, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6827347888702207, "supervised_tokens": 99.34375, "waves": 1.0, "detection_loss": 0.809366741104321, "grounding_loss": 0.4976573202203816, "detection_grounding_macro_loss": 0.6535120306623513}272{"step": 1190, "records_seen": 38080, "elapsed_seconds": 6759.590195655823, "gradient_norm": 1.8936883211135864, "learning_rate": 1.615909176632032e-05, "head_lr": 8.07954588316016e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6391811178511944, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6391811178511944, "supervised_tokens": 94.5625, "waves": 1.0, "detection_loss": 0.8214284433124703, "grounding_loss": 0.43263414899508157, "detection_grounding_macro_loss": 0.627031296153776}273{"step": 1200, "records_seen": 38400, "elapsed_seconds": 6809.052860021591, "gradient_norm": 1.6115193367004395, "learning_rate": 1.6093313274245314e-05, "head_lr": 8.046656637122656e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5450325503412614, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5450325503412614, "supervised_tokens": 108.625, "waves": 1.0, "detection_loss": 0.6910579123836619, "grounding_loss": 0.3572856562867465, "detection_grounding_macro_loss": 0.5241717843352042}274/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead275  warnings.warn(276/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead277  warnings.warn(278/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead279  warnings.warn(280/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead281  warnings.warn(282/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead283  warnings.warn(284/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead285  warnings.warn(286/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead287  warnings.warn(288/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead289  warnings.warn(290{"validation_step": 1200, "loss": 1.1349539532592225, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1349539532592225, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6450268565469275, "grounding_loss": 0.5563855458157403, "detection_grounding_macro_loss": 1.100706201181334}291{"step": 1210, "records_seen": 38720, "elapsed_seconds": 6998.108502149582, "gradient_norm": 1.7691231966018677, "learning_rate": 1.6027132055864823e-05, "head_lr": 8.01356602793241e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.456012293161848, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.456012293161848, "supervised_tokens": 70.15625, "waves": 1.0, "detection_loss": 0.6680953764201453, "grounding_loss": 0.3730232605825143, "detection_grounding_macro_loss": 0.5205593185013297}292{"step": 1220, "records_seen": 39040, "elapsed_seconds": 7049.938270568848, "gradient_norm": 1.7736855745315552, "learning_rate": 1.5960553344102116e-05, "head_lr": 7.980276672051058e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4971328282345553, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4971328282345553, "supervised_tokens": 96.84375, "waves": 1.0, "detection_loss": 0.8338021498460036, "grounding_loss": 0.26678013450040644, "detection_grounding_macro_loss": 0.5502911421732051}293{"step": 1230, "records_seen": 39360, "elapsed_seconds": 7104.020492076874, "gradient_norm": 1.5901724100112915, "learning_rate": 1.589358240331012e-05, "head_lr": 7.946791201655058e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5382971090093633, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5382971090093633, "supervised_tokens": 129.8125, "waves": 1.0, "detection_loss": 0.6781599059440417, "grounding_loss": 0.27128631486134097, "detection_grounding_macro_loss": 0.47472311040269133}294{"step": 1240, "records_seen": 39680, "elapsed_seconds": 7156.7221167087555, "gradient_norm": 1.993600845336914, "learning_rate": 1.5826224528855148e-05, "head_lr": 7.913112264427573e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.7506097829942462, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7506097829942462, "supervised_tokens": 72.5625, "waves": 1.0, "detection_loss": 1.1872531297467503, "grounding_loss": 0.41099829107563185, "detection_grounding_macro_loss": 0.7991257104111911}295{"step": 1250, "records_seen": 40000, "elapsed_seconds": 7209.017466545105, "gradient_norm": 1.5761722326278687, "learning_rate": 1.5758485046698214e-05, "head_lr": 7.879242523349107e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4705228357343003, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4705228357343003, "supervised_tokens": 117.15625, "waves": 1.0, "detection_loss": 0.6252914347192821, "grounding_loss": 0.295118423551321, "detection_grounding_macro_loss": 0.46020492913530153}296{"step": 1260, "records_seen": 40320, "elapsed_seconds": 7260.842280864716, "gradient_norm": 1.356885552406311, "learning_rate": 1.5690369312973907e-05, "head_lr": 7.845184656486952e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5242890476340847, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5242890476340847, "supervised_tokens": 97.75, "waves": 1.0, "detection_loss": 0.6948321941021947, "grounding_loss": 0.27503367971915466, "detection_grounding_macro_loss": 0.48493293691067463}297{"step": 1270, "records_seen": 40640, "elapsed_seconds": 7314.1248207092285, "gradient_norm": 1.812415361404419, "learning_rate": 1.5621882713566875e-05, "head_lr": 7.810941356783437e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.3653415146552561, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3653415146552561, "supervised_tokens": 113.53125, "waves": 1.0, "detection_loss": 0.47182874078348774, "grounding_loss": 0.24465599170992694, "detection_grounding_macro_loss": 0.35824236624670736}298{"step": 1280, "records_seen": 40960, "elapsed_seconds": 7367.791484594345, "gradient_norm": 1.8385300636291504, "learning_rate": 1.5553030663685978e-05, "head_lr": 7.776515331842987e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.518161053507356, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.518161053507356, "supervised_tokens": 157.96875, "waves": 1.0, "detection_loss": 0.569248547999277, "grounding_loss": 0.42063038220459764, "detection_grounding_macro_loss": 0.4949394651019373}299{"step": 1290, "records_seen": 41280, "elapsed_seconds": 7422.041713237762, "gradient_norm": 1.5674301385879517, "learning_rate": 1.5483818607436093e-05, "head_lr": 7.741909303718046e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6519621876819599, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6519621876819599, "supervised_tokens": 108.4375, "waves": 1.0, "detection_loss": 0.8828375568705572, "grounding_loss": 0.4482486266331976, "detection_grounding_macro_loss": 0.6655430917518774}300{"step": 1300, "records_seen": 41600, "elapsed_seconds": 7473.565064907074, "gradient_norm": 1.391767144203186, "learning_rate": 1.541425201738768e-05, "head_lr": 7.70712600869384e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.40850318179582246, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.40850318179582246, "supervised_tokens": 136.5625, "waves": 1.0, "detection_loss": 0.5716465186621797, "grounding_loss": 0.1700629202219156, "detection_grounding_macro_loss": 0.3708547194420477}301{"step": 1310, "records_seen": 41920, "elapsed_seconds": 7522.937952518463, "gradient_norm": 1.8203563690185547, "learning_rate": 1.5344336394144027e-05, "head_lr": 7.672168197072013e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.46361222683481174, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.46361222683481174, "supervised_tokens": 96.90625, "waves": 1.0, "detection_loss": 0.5528985451245191, "grounding_loss": 0.33311683856523955, "detection_grounding_macro_loss": 0.4430076918448793}302{"step": 1320, "records_seen": 42240, "elapsed_seconds": 7575.763216972351, "gradient_norm": 1.6436711549758911, "learning_rate": 1.5274077265906357e-05, "head_lr": 7.637038632953177e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6331813067663461, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6331813067663461, "supervised_tokens": 111.34375, "waves": 1.0, "detection_loss": 0.8988650239565793, "grounding_loss": 0.33207309395074847, "detection_grounding_macro_loss": 0.6154690589536639}303{"step": 1330, "records_seen": 42560, "elapsed_seconds": 7626.465668439865, "gradient_norm": 1.5253862142562866, "learning_rate": 1.5203480188036694e-05, "head_lr": 7.601740094018346e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.46421345649287105, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.46421345649287105, "supervised_tokens": 148.59375, "waves": 1.0, "detection_loss": 0.5077573217843708, "grounding_loss": 0.40057242260529446, "detection_grounding_macro_loss": 0.45416487219483265}304{"step": 1340, "records_seen": 42880, "elapsed_seconds": 7680.362449169159, "gradient_norm": 1.7300711870193481, "learning_rate": 1.5132550742618608e-05, "head_lr": 7.566275371309303e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5696541381439602, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5696541381439602, "supervised_tokens": 112.03125, "waves": 1.0, "detection_loss": 0.6688861113701327, "grounding_loss": 0.4420701725674527, "detection_grounding_macro_loss": 0.5554781419687926}305{"step": 1350, "records_seen": 43200, "elapsed_seconds": 7734.223039865494, "gradient_norm": 1.6704351902008057, "learning_rate": 1.5061294538015842e-05, "head_lr": 7.53064726900792e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4834641653427241, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4834641653427241, "supervised_tokens": 104.78125, "waves": 1.0, "detection_loss": 0.6824934663683442, "grounding_loss": 0.32866359787835286, "detection_grounding_macro_loss": 0.5055785321233486}306{"step": 1360, "records_seen": 43520, "elapsed_seconds": 7786.92831659317, "gradient_norm": 1.3495807647705078, "learning_rate": 1.4989717208428862e-05, "head_lr": 7.49485860421443e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4752644394757226, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4752644394757226, "supervised_tokens": 159.28125, "waves": 1.0, "detection_loss": 0.6516163798147127, "grounding_loss": 0.2175192959033526, "detection_grounding_macro_loss": 0.43456783785903264}307{"step": 1370, "records_seen": 43840, "elapsed_seconds": 7839.025936841965, "gradient_norm": 1.441202998161316, "learning_rate": 1.4917824413449364e-05, "head_lr": 7.458912206724681e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.44974414047101163, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.44974414047101163, "supervised_tokens": 104.15625, "waves": 1.0, "detection_loss": 0.6391861087565, "grounding_loss": 0.3360789594997186, "detection_grounding_macro_loss": 0.48763253412810925}308{"step": 1380, "records_seen": 44160, "elapsed_seconds": 7891.228297472, "gradient_norm": 1.8255913257598877, "learning_rate": 1.4845621837612765e-05, "head_lr": 7.422810918806381e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.595875064143911, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.595875064143911, "supervised_tokens": 120.1875, "waves": 1.0, "detection_loss": 0.8092978820204735, "grounding_loss": 0.4075608130763559, "detection_grounding_macro_loss": 0.6084293475484147}309{"step": 1390, "records_seen": 44480, "elapsed_seconds": 7941.924536705017, "gradient_norm": 1.9635387659072876, "learning_rate": 1.477311518994874e-05, "head_lr": 7.386557594974369e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5482475554522352, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5482475554522352, "supervised_tokens": 137.3125, "waves": 1.0, "detection_loss": 0.801195549908698, "grounding_loss": 0.2952995609957725, "detection_grounding_macro_loss": 0.5482475554522352}310{"step": 1400, "records_seen": 44800, "elapsed_seconds": 7994.949329853058, "gradient_norm": 1.5779293775558472, "learning_rate": 1.4700310203529799e-05, "head_lr": 7.350155101764899e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.7804086724078729, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7804086724078729, "supervised_tokens": 169.40625, "waves": 1.0, "detection_loss": 0.9286337188749831, "grounding_loss": 0.4016113314363692, "detection_grounding_macro_loss": 0.6651225251556762}311{"step": 1410, "records_seen": 45120, "elapsed_seconds": 8050.004055738449, "gradient_norm": 1.6246696710586548, "learning_rate": 1.462721263501799e-05, "head_lr": 7.313606317508995e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6563532504951581, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6563532504951581, "supervised_tokens": 113.65625, "waves": 1.0, "detection_loss": 0.9314913320190766, "grounding_loss": 0.3445300914347172, "detection_grounding_macro_loss": 0.6380107117268969}312{"step": 1420, "records_seen": 45440, "elapsed_seconds": 8102.458473920822, "gradient_norm": 1.6199417114257812, "learning_rate": 1.4553828264209708e-05, "head_lr": 7.276914132104854e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6343031511114532, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6343031511114532, "supervised_tokens": 126.59375, "waves": 1.0, "detection_loss": 0.5937546743560865, "grounding_loss": 0.7379270361529456, "detection_grounding_macro_loss": 0.6658408552545161}313{"step": 1430, "records_seen": 45760, "elapsed_seconds": 8156.150224208832, "gradient_norm": 1.7873356342315674, "learning_rate": 1.4480162893578694e-05, "head_lr": 7.240081446789346e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.40518773043663714, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.40518773043663714, "supervised_tokens": 129.09375, "waves": 1.0, "detection_loss": 0.44118295501393734, "grounding_loss": 0.37342723816254875, "detection_grounding_macro_loss": 0.40730509658824304}314{"step": 1440, "records_seen": 46080, "elapsed_seconds": 8210.39703798294, "gradient_norm": 1.9659236669540405, "learning_rate": 1.4406222347817239e-05, "head_lr": 7.203111173908618e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.7064750933386676, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7064750933386676, "supervised_tokens": 111.53125, "waves": 1.0, "detection_loss": 0.9639177559053197, "grounding_loss": 0.4147067424297954, "detection_grounding_macro_loss": 0.6893122491675575}315{"step": 1450, "records_seen": 46400, "elapsed_seconds": 8261.874951124191, "gradient_norm": 1.8494189977645874, "learning_rate": 1.433201247337562e-05, "head_lr": 7.16600623668781e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4891976330163743, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4891976330163743, "supervised_tokens": 70.9375, "waves": 1.0, "detection_loss": 0.5376434401135027, "grounding_loss": 0.4515175608297189, "detection_grounding_macro_loss": 0.49458050047161084}316{"step": 1460, "records_seen": 46720, "elapsed_seconds": 8315.520154714584, "gradient_norm": 1.5615687370300293, "learning_rate": 1.4257539137999836e-05, "head_lr": 7.128769568999917e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.43273960514898135, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.43273960514898135, "supervised_tokens": 101.125, "waves": 1.0, "detection_loss": 0.3813116803578262, "grounding_loss": 0.4910245865789572, "detection_grounding_macro_loss": 0.43616813346839167}317{"step": 1470, "records_seen": 47040, "elapsed_seconds": 8369.428748130798, "gradient_norm": 2.2104554176330566, "learning_rate": 1.4182808230267636e-05, "head_lr": 7.091404115133817e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.7711887375848789, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7711887375848789, "supervised_tokens": 96.28125, "waves": 1.0, "detection_loss": 1.1471688230686292, "grounding_loss": 0.43944160333451104, "detection_grounding_macro_loss": 0.7933052132015701}318{"step": 1480, "records_seen": 47360, "elapsed_seconds": 8420.749029874802, "gradient_norm": 1.4294430017471313, "learning_rate": 1.4107825659122907e-05, "head_lr": 7.053912829561453e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.3777576711733275, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3777576711733275, "supervised_tokens": 101.125, "waves": 1.0, "detection_loss": 0.47620258970046586, "grounding_loss": 0.2661867635092373, "detection_grounding_macro_loss": 0.3711946766048516}319{"step": 1490, "records_seen": 47680, "elapsed_seconds": 8473.318408727646, "gradient_norm": 1.6012135744094849, "learning_rate": 1.4032597353408483e-05, "head_lr": 7.01629867670424e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.49930665847080036, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49930665847080036, "supervised_tokens": 92.59375, "waves": 1.0, "detection_loss": 0.702328365189554, "grounding_loss": 0.2692153908562129, "detection_grounding_macro_loss": 0.48577187802288346}320{"step": 1500, "records_seen": 48000, "elapsed_seconds": 8524.644358158112, "gradient_norm": 1.9752506017684937, "learning_rate": 1.3957129261397317e-05, "head_lr": 6.978564630698658e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6288573344925936, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6288573344925936, "supervised_tokens": 108.46875, "waves": 1.0, "detection_loss": 0.9381393423459182, "grounding_loss": 0.3559614452102484, "detection_grounding_macro_loss": 0.6470503937780834}321/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead322  warnings.warn(323/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead324  warnings.warn(325/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead326  warnings.warn(327/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead328  warnings.warn(329/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead330  warnings.warn(331/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead332  warnings.warn(333/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead334  warnings.warn(335/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead336  warnings.warn(337{"validation_step": 1500, "loss": 1.1347353433673883, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1347353433673883, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6420785877815878, "grounding_loss": 0.5592631489889962, "detection_grounding_macro_loss": 1.100670868385292}338{"step": 1510, "records_seen": 48320, "elapsed_seconds": 8726.602244377136, "gradient_norm": 1.6287468671798706, "learning_rate": 1.3881427350322183e-05, "head_lr": 6.940713675161091e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.7052497050578097, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7052497050578097, "supervised_tokens": 131.4375, "waves": 1.0, "detection_loss": 0.9809656555161756, "grounding_loss": 0.39277162787166164, "detection_grounding_macro_loss": 0.6868686416939187}339{"step": 1520, "records_seen": 48640, "elapsed_seconds": 8778.271958589554, "gradient_norm": 1.4447860717773438, "learning_rate": 1.380549760590383e-05, "head_lr": 6.902748802951914e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.495997858498896, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.495997858498896, "supervised_tokens": 101.875, "waves": 1.0, "detection_loss": 0.5969195819048764, "grounding_loss": 0.3277949861555953, "detection_grounding_macro_loss": 0.46235728403023585}340{"step": 1530, "records_seen": 48960, "elapsed_seconds": 8831.304020881653, "gradient_norm": 2.0302462577819824, "learning_rate": 1.372934603187771e-05, "head_lr": 6.864673015938855e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.7241228460331968, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7241228460331968, "supervised_tokens": 109.15625, "waves": 1.0, "detection_loss": 0.8627278244288887, "grounding_loss": 0.4191918935626745, "detection_grounding_macro_loss": 0.6409598589957817}341{"step": 1540, "records_seen": 49280, "elapsed_seconds": 8884.568392753601, "gradient_norm": 1.9502202272415161, "learning_rate": 1.3652978649519246e-05, "head_lr": 6.826489324759622e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5985953062272387, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5985953062272387, "supervised_tokens": 88.6875, "waves": 1.0, "detection_loss": 0.9512539992729823, "grounding_loss": 0.2874258711868767, "detection_grounding_macro_loss": 0.6193399352299295}342{"step": 1550, "records_seen": 49600, "elapsed_seconds": 8938.742155790329, "gradient_norm": 1.7505635023117065, "learning_rate": 1.3576401497167749e-05, "head_lr": 6.788200748583874e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5072259757326947, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5072259757326947, "supervised_tokens": 112.4375, "waves": 1.0, "detection_loss": 0.5024003616999835, "grounding_loss": 0.5101213441523214, "detection_grounding_macro_loss": 0.5062608529261524}343{"step": 1560, "records_seen": 49920, "elapsed_seconds": 8991.606324911118, "gradient_norm": 1.4156628847122192, "learning_rate": 1.3499620629748959e-05, "head_lr": 6.749810314874479e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.4110191451909486, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4110191451909486, "supervised_tokens": 134.8125, "waves": 1.0, "detection_loss": 0.48821025487268344, "grounding_loss": 0.3338280355092138, "detection_grounding_macro_loss": 0.4110191451909486}344{"step": 1570, "records_seen": 50240, "elapsed_seconds": 9046.299206733704, "gradient_norm": 1.4695672988891602, "learning_rate": 1.3422642118296292e-05, "head_lr": 6.711321059148145e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.43018265695275204, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.43018265695275204, "supervised_tokens": 118.3125, "waves": 1.0, "detection_loss": 0.5439369022751634, "grounding_loss": 0.31642841163034063, "detection_grounding_macro_loss": 0.43018265695275204}345{"step": 1580, "records_seen": 50560, "elapsed_seconds": 9099.853683948517, "gradient_norm": 1.9705153703689575, "learning_rate": 1.3345472049470791e-05, "head_lr": 6.672736024735395e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.658078716158343, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.658078716158343, "supervised_tokens": 110.4375, "waves": 1.0, "detection_loss": 0.801285845705455, "grounding_loss": 0.4487759883587177, "detection_grounding_macro_loss": 0.6250309170320864}346{"step": 1590, "records_seen": 50880, "elapsed_seconds": 9155.582436800003, "gradient_norm": 1.6555215120315552, "learning_rate": 1.326811652507988e-05, "head_lr": 6.634058262539939e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.5771256948087284, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5771256948087284, "supervised_tokens": 92.1875, "waves": 1.0, "detection_loss": 0.810543426258846, "grounding_loss": 0.34370796335861087, "detection_grounding_macro_loss": 0.5771256948087284}347{"step": 1600, "records_seen": 51200, "elapsed_seconds": 9211.676916360855, "gradient_norm": 1.8732268810272217, "learning_rate": 1.3190581661594869e-05, "head_lr": 6.595290830797433e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6272970037534833, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6272970037534833, "supervised_tokens": 124.1875, "waves": 1.0, "detection_loss": 0.6610239169707424, "grounding_loss": 0.5780038228974893, "detection_grounding_macro_loss": 0.6195138699341158}348{"step": 1610, "records_seen": 51520, "elapsed_seconds": 9266.160034179688, "gradient_norm": 1.7625991106033325, "learning_rate": 1.3112873589667347e-05, "head_lr": 6.556436794833673e-07, "peak_gpu_gib": 19.755093097686768, "loss": 0.6040038459057513, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6040038459057513, "supervised_tokens": 93.40625, "waves": 1.0, "detection_loss": 0.9151401037678999, "grounding_loss": 0.25138275366198287, "detection_grounding_macro_loss": 0.5832614287149414}349{"step": 1620, "records_seen": 51840, "elapsed_seconds": 9318.845374584198, "gradient_norm": 2.0519309043884277, "learning_rate": 1.3034998453644437e-05, "head_lr": 6.517499226822217e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6935480216197902, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6935480216197902, "supervised_tokens": 110.3125, "waves": 1.0, "detection_loss": 0.9386819451312093, "grounding_loss": 0.33527536418002385, "detection_grounding_macro_loss": 0.6369786546556165}350{"step": 1630, "records_seen": 52160, "elapsed_seconds": 9372.909459114075, "gradient_norm": 1.416148066520691, "learning_rate": 1.2956962411082932e-05, "head_lr": 6.478481205541465e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4303344111285696, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4303344111285696, "supervised_tokens": 112.59375, "waves": 1.0, "detection_loss": 0.5911924152824213, "grounding_loss": 0.269476406974718, "detection_grounding_macro_loss": 0.4303344111285696}351{"step": 1640, "records_seen": 52480, "elapsed_seconds": 9428.126890182495, "gradient_norm": 1.9424529075622559, "learning_rate": 1.2878771632262461e-05, "head_lr": 6.43938581613123e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.696755746845156, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.696755746845156, "supervised_tokens": 124.15625, "waves": 1.0, "detection_loss": 0.9097798677051768, "grounding_loss": 0.4553284098704656, "detection_grounding_macro_loss": 0.6825541387878212}352{"step": 1650, "records_seen": 52800, "elapsed_seconds": 9480.874320745468, "gradient_norm": 1.9334499835968018, "learning_rate": 1.2800432299697592e-05, "head_lr": 6.400216149848795e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5988245845810525, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5988245845810525, "supervised_tokens": 125.4375, "waves": 1.0, "detection_loss": 0.7769278840079143, "grounding_loss": 0.33851976234179276, "detection_grounding_macro_loss": 0.5577238231748536}353{"step": 1660, "records_seen": 53120, "elapsed_seconds": 9536.841276407242, "gradient_norm": 1.6978199481964111, "learning_rate": 1.2721950607648962e-05, "head_lr": 6.360975303824481e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4812120980368064, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4812120980368064, "supervised_tokens": 113.21875, "waves": 1.0, "detection_loss": 0.5636715601597514, "grounding_loss": 0.4170769608300715, "detection_grounding_macro_loss": 0.49037426049491145}354{"step": 1670, "records_seen": 53440, "elapsed_seconds": 9588.323606967926, "gradient_norm": 1.784921407699585, "learning_rate": 1.2643332761633531e-05, "head_lr": 6.321666380816765e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6111054735647485, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6111054735647485, "supervised_tokens": 94.03125, "waves": 1.0, "detection_loss": 0.7501176205900265, "grounding_loss": 0.37941856185595196, "detection_grounding_macro_loss": 0.5647680912229892}355{"step": 1680, "records_seen": 53760, "elapsed_seconds": 9643.066557645798, "gradient_norm": 1.6900649070739746, "learning_rate": 1.25645849779339e-05, "head_lr": 6.282292488966949e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.40113374241853705, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.40113374241853705, "supervised_tokens": 89.65625, "waves": 1.0, "detection_loss": 0.4803489528192131, "grounding_loss": 0.299285614760525, "detection_grounding_macro_loss": 0.389817283789869}356{"step": 1690, "records_seen": 54080, "elapsed_seconds": 9697.83579993248, "gradient_norm": 1.6620304584503174, "learning_rate": 1.2485713483106783e-05, "head_lr": 6.242856741553391e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6610209927143842, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6610209927143842, "supervised_tokens": 102.125, "waves": 1.0, "detection_loss": 0.8467148785283725, "grounding_loss": 0.42227171095354216, "detection_grounding_macro_loss": 0.6344932947409573}357{"step": 1700, "records_seen": 54400, "elapsed_seconds": 9748.43567442894, "gradient_norm": 1.7533574104309082, "learning_rate": 1.2406724513490695e-05, "head_lr": 6.203362256745347e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5996563414696539, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5996563414696539, "supervised_tokens": 98.65625, "waves": 1.0, "detection_loss": 0.8393985982704459, "grounding_loss": 0.4356221657638487, "detection_grounding_macro_loss": 0.6375103820171473}358{"step": 1710, "records_seen": 54720, "elapsed_seconds": 9802.574964284897, "gradient_norm": 1.5025938749313354, "learning_rate": 1.232762431471283e-05, "head_lr": 6.163812157356414e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5126940975055163, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5126940975055163, "supervised_tokens": 109.28125, "waves": 1.0, "detection_loss": 0.6902109032336525, "grounding_loss": 0.311508384346962, "detection_grounding_macro_loss": 0.5008596437903072}359{"step": 1720, "records_seen": 55040, "elapsed_seconds": 9852.479530096054, "gradient_norm": 1.5496786832809448, "learning_rate": 1.2248419141195234e-05, "head_lr": 6.124209570597616e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.43096530861490123, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.43096530861490123, "supervised_tokens": 91.46875, "waves": 1.0, "detection_loss": 0.6149491626298575, "grounding_loss": 0.2878667554921574, "detection_grounding_macro_loss": 0.4514079590610075}360{"step": 1730, "records_seen": 55360, "elapsed_seconds": 9904.487872600555, "gradient_norm": 1.542229175567627, "learning_rate": 1.2169115255660262e-05, "head_lr": 6.08455762783013e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5211113226657744, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5211113226657744, "supervised_tokens": 112.53125, "waves": 1.0, "detection_loss": 0.7635726730028788, "grounding_loss": 0.3071748370742117, "detection_grounding_macro_loss": 0.5353737550385452}361{"step": 1740, "records_seen": 55680, "elapsed_seconds": 9957.737817525864, "gradient_norm": 1.438299536705017, "learning_rate": 1.2089718928635386e-05, "head_lr": 6.044859464317692e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.47814983724674676, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.47814983724674676, "supervised_tokens": 130.25, "waves": 1.0, "detection_loss": 0.6820204028497554, "grounding_loss": 0.21603053861430713, "detection_grounding_macro_loss": 0.44902547073203125}362{"step": 1750, "records_seen": 56000, "elapsed_seconds": 10011.842173337936, "gradient_norm": 1.6231648921966553, "learning_rate": 1.20102364379574e-05, "head_lr": 6.0051182189787e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6655359157601879, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6655359157601879, "supervised_tokens": 135.0, "waves": 1.0, "detection_loss": 0.7048085003302081, "grounding_loss": 0.39062782377004623, "detection_grounding_macro_loss": 0.5477181620501272}363{"step": 1760, "records_seen": 56320, "elapsed_seconds": 10065.284998893738, "gradient_norm": 1.6925082206726074, "learning_rate": 1.1930674068276022e-05, "head_lr": 5.96533703413801e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4977845819103095, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4977845819103095, "supervised_tokens": 98.59375, "waves": 1.0, "detection_loss": 0.6146548626362346, "grounding_loss": 0.39466374597566967, "detection_grounding_macro_loss": 0.5046593043059522}364{"step": 1770, "records_seen": 56640, "elapsed_seconds": 10116.219336032867, "gradient_norm": 1.4896174669265747, "learning_rate": 1.1851038110556967e-05, "head_lr": 5.925519055278482e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.45094836338228106, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.45094836338228106, "supervised_tokens": 135.03125, "waves": 1.0, "detection_loss": 0.5524514294831533, "grounding_loss": 0.28177658654749393, "detection_grounding_macro_loss": 0.4171140080153236}365{"step": 1780, "records_seen": 56960, "elapsed_seconds": 10170.051823854446, "gradient_norm": 1.8129549026489258, "learning_rate": 1.1771334861584524e-05, "head_lr": 5.885667430792262e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6068295449135803, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6068295449135803, "supervised_tokens": 114.8125, "waves": 1.0, "detection_loss": 0.7895082656734089, "grounding_loss": 0.3997936613857746, "detection_grounding_macro_loss": 0.5946509635295918}366{"step": 1790, "records_seen": 57280, "elapsed_seconds": 10223.237755060196, "gradient_norm": 1.8357725143432617, "learning_rate": 1.1691570623463694e-05, "head_lr": 5.845785311731847e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5829865973959443, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5829865973959443, "supervised_tokens": 115.28125, "waves": 1.0, "detection_loss": 0.8127390129698647, "grounding_loss": 0.2875906345151894, "detection_grounding_macro_loss": 0.5501648237425271}367{"step": 1800, "records_seen": 57600, "elapsed_seconds": 10274.278177022934, "gradient_norm": 1.7822133302688599, "learning_rate": 1.1611751703121843e-05, "head_lr": 5.80587585156092e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5650167276653519, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5650167276653519, "supervised_tokens": 85.90625, "waves": 1.0, "detection_loss": 0.7738592672942426, "grounding_loss": 0.3283285160859426, "detection_grounding_macro_loss": 0.5510938916900926}368/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead369  warnings.warn(370/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead371  warnings.warn(372/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead373  warnings.warn(374/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead375  warnings.warn(376/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead377  warnings.warn(378/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead379  warnings.warn(380/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead381  warnings.warn(382/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead383  warnings.warn(384{"validation_step": 1800, "loss": 1.1311272809741801, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1311272809741801, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6388519813414175, "grounding_loss": 0.5552224065576281, "detection_grounding_macro_loss": 1.097037193949523}385{"step": 1810, "records_seen": 57920, "elapsed_seconds": 10466.016971111298, "gradient_norm": 1.3405752182006836, "learning_rate": 1.1531884411810055e-05, "head_lr": 5.765942205905027e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.41281796159722717, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.41281796159722717, "supervised_tokens": 124.0625, "waves": 1.0, "detection_loss": 0.5470337021567199, "grounding_loss": 0.24025486659216508, "detection_grounding_macro_loss": 0.3936442843744425}386{"step": 1820, "records_seen": 58240, "elapsed_seconds": 10521.903654813766, "gradient_norm": 1.8153661489486694, "learning_rate": 1.1451975064604082e-05, "head_lr": 5.72598753230204e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5829902344262905, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5829902344262905, "supervised_tokens": 101.125, "waves": 1.0, "detection_loss": 0.7390405076330353, "grounding_loss": 0.3823541688747615, "detection_grounding_macro_loss": 0.5606973382538984}387{"step": 1830, "records_seen": 58560, "elapsed_seconds": 10575.552056789398, "gradient_norm": 1.3377891778945923, "learning_rate": 1.1372029979905022e-05, "head_lr": 5.68601498995251e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.44575860620966523, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.44575860620966523, "supervised_tokens": 100.59375, "waves": 1.0, "detection_loss": 0.6168502658565558, "grounding_loss": 0.2746669465627747, "detection_grounding_macro_loss": 0.44575860620966523}388{"step": 1840, "records_seen": 58880, "elapsed_seconds": 10625.027033090591, "gradient_norm": 1.640823245048523, "learning_rate": 1.1292055478939716e-05, "head_lr": 5.646027739469858e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4487980297899412, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4487980297899412, "supervised_tokens": 97.375, "waves": 1.0, "detection_loss": 0.5083316642161332, "grounding_loss": 0.4176137450905073, "detection_grounding_macro_loss": 0.46297270465332024}389{"step": 1850, "records_seen": 59200, "elapsed_seconds": 10677.847536325455, "gradient_norm": 1.8395360708236694, "learning_rate": 1.1212057885260943e-05, "head_lr": 5.60602894263047e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6336775639929328, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6336775639929328, "supervised_tokens": 92.1875, "waves": 1.0, "detection_loss": 0.8388381786935497, "grounding_loss": 0.4526534921982709, "detection_grounding_macro_loss": 0.6457458354459102}390{"step": 1860, "records_seen": 59520, "elapsed_seconds": 10732.132769584656, "gradient_norm": 1.5440996885299683, "learning_rate": 1.1132043524247418e-05, "head_lr": 5.566021762123708e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.49921060720356536, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49921060720356536, "supervised_tokens": 108.625, "waves": 1.0, "detection_loss": 0.6753183422196243, "grounding_loss": 0.2727863764686325, "detection_grounding_macro_loss": 0.47405235934412837}391{"step": 1870, "records_seen": 59840, "elapsed_seconds": 10784.06855416298, "gradient_norm": 1.4461562633514404, "learning_rate": 1.1052018722603634e-05, "head_lr": 5.526009361301817e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.480888710415627, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.480888710415627, "supervised_tokens": 135.625, "waves": 1.0, "detection_loss": 0.648691824373707, "grounding_loss": 0.290711847929803, "detection_grounding_macro_loss": 0.469701836151755}392{"step": 1880, "records_seen": 60160, "elapsed_seconds": 10836.68670964241, "gradient_norm": 1.6340874433517456, "learning_rate": 1.0971989807859624e-05, "head_lr": 5.485994903929811e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.43198760776306244, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.43198760776306244, "supervised_tokens": 80.875, "waves": 1.0, "detection_loss": 0.5373398594205997, "grounding_loss": 0.31258838921785354, "detection_grounding_macro_loss": 0.4249641243192266}393{"step": 1890, "records_seen": 60480, "elapsed_seconds": 10888.78056883812, "gradient_norm": 1.550662636756897, "learning_rate": 1.089196310787065e-05, "head_lr": 5.445981553935324e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.541553185462476, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.541553185462476, "supervised_tokens": 127.25, "waves": 1.0, "detection_loss": 0.771558233226339, "grounding_loss": 0.24583240976608067, "detection_grounding_macro_loss": 0.5086953214962098}394{"step": 1900, "records_seen": 60800, "elapsed_seconds": 10940.655395746231, "gradient_norm": 1.466349720954895, "learning_rate": 1.0811944950316837e-05, "head_lr": 5.405972475158418e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6095617622249847, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6095617622249847, "supervised_tokens": 110.8125, "waves": 1.0, "detection_loss": 0.892537219097398, "grounding_loss": 0.24573617481759616, "detection_grounding_macro_loss": 0.5691366969574971}395{"step": 1910, "records_seen": 61120, "elapsed_seconds": 10994.226069450378, "gradient_norm": 1.4578999280929565, "learning_rate": 1.0731941662202879e-05, "head_lr": 5.365970831101438e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.42919358424200027, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.42919358424200027, "supervised_tokens": 108.5, "waves": 1.0, "detection_loss": 0.5848695679019329, "grounding_loss": 0.27351760058206764, "detection_grounding_macro_loss": 0.42919358424200027}396{"step": 1920, "records_seen": 61440, "elapsed_seconds": 11050.317736387253, "gradient_norm": 1.6001859903335571, "learning_rate": 1.065195956935774e-05, "head_lr": 5.325979784678869e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6257397402805509, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6257397402805509, "supervised_tokens": 107.75, "waves": 1.0, "detection_loss": 0.853334886324592, "grounding_loss": 0.24641449687381586, "detection_grounding_macro_loss": 0.5498746915992039}397{"step": 1930, "records_seen": 61760, "elapsed_seconds": 11103.072159290314, "gradient_norm": 1.7208279371261597, "learning_rate": 1.0572004995934488e-05, "head_lr": 5.286002497967243e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5715339413809488, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5715339413809488, "supervised_tokens": 105.96875, "waves": 1.0, "detection_loss": 0.7691731886669813, "grounding_loss": 0.3174263377274786, "detection_grounding_macro_loss": 0.54329976319723}398{"step": 1940, "records_seen": 62080, "elapsed_seconds": 11156.3429646492, "gradient_norm": 1.5786539316177368, "learning_rate": 1.049208426391024e-05, "head_lr": 5.24604213195512e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5632839663794584, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5632839663794584, "supervised_tokens": 116.90625, "waves": 1.0, "detection_loss": 0.7183616638888933, "grounding_loss": 0.3875292425354322, "detection_grounding_macro_loss": 0.5529454532121627}399{"step": 1950, "records_seen": 62400, "elapsed_seconds": 11207.908033847809, "gradient_norm": 1.86215078830719, "learning_rate": 1.0412203692586283e-05, "head_lr": 5.206101846293141e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.684474224903397, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.684474224903397, "supervised_tokens": 87.4375, "waves": 1.0, "detection_loss": 0.9469452028142233, "grounding_loss": 0.3870071166044606, "detection_grounding_macro_loss": 0.666976159709342}400{"step": 1960, "records_seen": 62720, "elapsed_seconds": 11257.786566019058, "gradient_norm": 1.6799546480178833, "learning_rate": 1.0332369598088412e-05, "head_lr": 5.166184799044206e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5585685255570496, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5585685255570496, "supervised_tokens": 98.625, "waves": 1.0, "detection_loss": 0.8839349581646577, "grounding_loss": 0.27148049678563063, "detection_grounding_macro_loss": 0.5777077274751442}401{"step": 1970, "records_seen": 63040, "elapsed_seconds": 11309.596284151077, "gradient_norm": 1.7040997743606567, "learning_rate": 1.0252588292867523e-05, "head_lr": 5.126294146433761e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5282488097543592, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5282488097543592, "supervised_tokens": 86.46875, "waves": 1.0, "detection_loss": 0.7110314259835457, "grounding_loss": 0.32109517802794774, "detection_grounding_macro_loss": 0.5160633020057468}402{"step": 1980, "records_seen": 63360, "elapsed_seconds": 11364.941035270691, "gradient_norm": 1.6882073879241943, "learning_rate": 1.0172866085200484e-05, "head_lr": 5.086433042600241e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5936317249470449, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5936317249470449, "supervised_tokens": 129.375, "waves": 1.0, "detection_loss": 0.7778822804490725, "grounding_loss": 0.24188066444317388, "detection_grounding_macro_loss": 0.5098814724461231}403{"step": 1990, "records_seen": 63680, "elapsed_seconds": 11426.458268165588, "gradient_norm": 1.9617669582366943, "learning_rate": 1.0093209278691336e-05, "head_lr": 5.046604639345668e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5621580252368403, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5621580252368403, "supervised_tokens": 102.8125, "waves": 1.0, "detection_loss": 0.7077747583704573, "grounding_loss": 0.3749365112079041, "detection_grounding_macro_loss": 0.5413556347891807}404{"step": 2000, "records_seen": 64000, "elapsed_seconds": 11503.91003870964, "gradient_norm": 1.6940299272537231, "learning_rate": 1.0013624171772887e-05, "head_lr": 5.006812085886442e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.41028480113755705, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.41028480113755705, "supervised_tokens": 111.28125, "waves": 1.0, "detection_loss": 0.6174256086349488, "grounding_loss": 0.31612988863874264, "detection_grounding_macro_loss": 0.46677774863684574}405{"step": 2010, "records_seen": 64320, "elapsed_seconds": 11555.42779803276, "gradient_norm": 1.7772804498672485, "learning_rate": 9.93411705720867e-06, "head_lr": 4.967058528604334e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5780081128177699, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5780081128177699, "supervised_tokens": 101.0, "waves": 1.0, "detection_loss": 0.7735529695637524, "grounding_loss": 0.3824632560717873, "detection_grounding_macro_loss": 0.5780081128177699}406{"step": 2020, "records_seen": 64640, "elapsed_seconds": 11607.47647190094, "gradient_norm": 1.4683222770690918, "learning_rate": 9.854694221595409e-06, "head_lr": 4.927347110797704e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.578880802651355, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.578880802651355, "supervised_tokens": 126.03125, "waves": 1.0, "detection_loss": 0.7783777132596393, "grounding_loss": 0.32238477472641797, "detection_grounding_macro_loss": 0.5503812439930287}407{"step": 2030, "records_seen": 64960, "elapsed_seconds": 11660.393063783646, "gradient_norm": 1.4176002740859985, "learning_rate": 9.775361944865913e-06, "head_lr": 4.887680972432956e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5265267678449277, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5265267678449277, "supervised_tokens": 155.21875, "waves": 1.0, "detection_loss": 0.7036204803735018, "grounding_loss": 0.298834851736761, "detection_grounding_macro_loss": 0.5012276660551314}408{"step": 2040, "records_seen": 65280, "elapsed_seconds": 11715.137231826782, "gradient_norm": 1.2744768857955933, "learning_rate": 9.696126499792542e-06, "head_lr": 4.84806324989627e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.3311517853139492, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3311517853139492, "supervised_tokens": 129.09375, "waves": 1.0, "detection_loss": 0.45827947985890205, "grounding_loss": 0.2040240907689963, "detection_grounding_macro_loss": 0.3311517853139492}409{"step": 2050, "records_seen": 65600, "elapsed_seconds": 11766.778483629227, "gradient_norm": 1.7921069860458374, "learning_rate": 9.61699415149121e-06, "head_lr": 4.808497075745605e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5608782923308127, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5608782923308127, "supervised_tokens": 116.09375, "waves": 1.0, "detection_loss": 0.6886584528225572, "grounding_loss": 0.37412267315057063, "detection_grounding_macro_loss": 0.531390562986564}410{"step": 2060, "records_seen": 65920, "elapsed_seconds": 11819.71349644661, "gradient_norm": 1.7900880575180054, "learning_rate": 9.537971156926011e-06, "head_lr": 4.768985578463005e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5484651366714388, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5484651366714388, "supervised_tokens": 96.375, "waves": 1.0, "detection_loss": 0.8486692421138287, "grounding_loss": 0.3683426734060049, "detection_grounding_macro_loss": 0.6085059577599168}411{"step": 2070, "records_seen": 66240, "elapsed_seconds": 11875.175348997116, "gradient_norm": 1.5239349603652954, "learning_rate": 9.459063764414483e-06, "head_lr": 4.7295318822072406e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5120302056102446, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5120302056102446, "supervised_tokens": 122.34375, "waves": 1.0, "detection_loss": 0.6824501218866584, "grounding_loss": 0.31888763383030894, "detection_grounding_macro_loss": 0.5006688778584837}412{"step": 2080, "records_seen": 66560, "elapsed_seconds": 11926.62998843193, "gradient_norm": 1.5835984945297241, "learning_rate": 9.38027821313355e-06, "head_lr": 4.690139106566774e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5335391303490269, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5335391303490269, "supervised_tokens": 85.40625, "waves": 1.0, "detection_loss": 0.5943396420255643, "grounding_loss": 0.4553670439077645, "detection_grounding_macro_loss": 0.5248533429666644}413{"step": 2090, "records_seen": 66880, "elapsed_seconds": 11979.199779987335, "gradient_norm": 1.7055479288101196, "learning_rate": 9.3016207326262e-06, "head_lr": 4.6508103663130994e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5787365532451076, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5787365532451076, "supervised_tokens": 83.09375, "waves": 1.0, "detection_loss": 0.7942606747092213, "grounding_loss": 0.36321243178099394, "detection_grounding_macro_loss": 0.5787365532451076}414{"step": 2100, "records_seen": 67200, "elapsed_seconds": 12033.492106199265, "gradient_norm": 1.6195809841156006, "learning_rate": 9.22309754230892e-06, "head_lr": 4.6115487711544586e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.40312361324868107, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.40312361324868107, "supervised_tokens": 103.1875, "waves": 1.0, "detection_loss": 0.4717780353288011, "grounding_loss": 0.31485364200281246, "detection_grounding_macro_loss": 0.39331583866580677}415/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead416  warnings.warn(417/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead418  warnings.warn(419/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead420  warnings.warn(421/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead422  warnings.warn(423/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead424  warnings.warn(425/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead426  warnings.warn(427/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead428  warnings.warn(429/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead430  warnings.warn(431{"validation_step": 2100, "loss": 1.1328260714493945, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1328260714493945, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6465408316006933, "grounding_loss": 0.5501267577920641, "detection_grounding_macro_loss": 1.0983337946963787}432{"step": 2110, "records_seen": 67520, "elapsed_seconds": 12275.838193655014, "gradient_norm": 1.6882615089416504, "learning_rate": 9.144714850979911e-06, "head_lr": 4.5723574254899556e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5772135875611184, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5772135875611184, "supervised_tokens": 100.03125, "waves": 1.0, "detection_loss": 0.7437635121584916, "grounding_loss": 0.3630779702216387, "detection_grounding_macro_loss": 0.5534207411900651}433{"step": 2120, "records_seen": 67840, "elapsed_seconds": 12327.570021867752, "gradient_norm": 1.4807360172271729, "learning_rate": 9.066478856328185e-06, "head_lr": 4.533239428164092e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.3704412180466079, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3704412180466079, "supervised_tokens": 125.84375, "waves": 1.0, "detection_loss": 0.4273640370684964, "grounding_loss": 0.28724632870692474, "detection_grounding_macro_loss": 0.35730518288771057}434{"step": 2130, "records_seen": 68160, "elapsed_seconds": 12380.304695606232, "gradient_norm": 1.5309059619903564, "learning_rate": 8.988395744443498e-06, "head_lr": 4.4941978722217487e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4037976334897877, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4037976334897877, "supervised_tokens": 123.8125, "waves": 1.0, "detection_loss": 0.5823485562594821, "grounding_loss": 0.246252701634175, "detection_grounding_macro_loss": 0.41430062894682856}435{"step": 2140, "records_seen": 68480, "elapsed_seconds": 12431.535986423492, "gradient_norm": 1.5693773031234741, "learning_rate": 8.91047168932723e-06, "head_lr": 4.455235844663614e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.38157800324682967, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.38157800324682967, "supervised_tokens": 62.125, "waves": 1.0, "detection_loss": 0.5945802505397296, "grounding_loss": 0.28475879993187514, "detection_grounding_macro_loss": 0.4396695252358024}436{"step": 2150, "records_seen": 68800, "elapsed_seconds": 12488.570915699005, "gradient_norm": 1.5389286279678345, "learning_rate": 8.832712852404199e-06, "head_lr": 4.416356426202099e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6439678609804105, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6439678609804105, "supervised_tokens": 149.375, "waves": 1.0, "detection_loss": 0.8182342498025502, "grounding_loss": 0.3112774823199619, "detection_grounding_macro_loss": 0.564755866061256}437{"step": 2160, "records_seen": 69120, "elapsed_seconds": 12541.162546634674, "gradient_norm": 1.5979175567626953, "learning_rate": 8.75512538203549e-06, "head_lr": 4.3775626910177444e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4867643800098449, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4867643800098449, "supervised_tokens": 132.0, "waves": 1.0, "detection_loss": 0.6352704561182431, "grounding_loss": 0.3712596541477574, "detection_grounding_macro_loss": 0.5032650551330002}438{"step": 2170, "records_seen": 69440, "elapsed_seconds": 12595.538232803345, "gradient_norm": 1.7233291864395142, "learning_rate": 8.677715413032297e-06, "head_lr": 4.338857706516148e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5570856028457456, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5570856028457456, "supervised_tokens": 105.46875, "waves": 1.0, "detection_loss": 0.6430363430724052, "grounding_loss": 0.44657750826861176, "detection_grounding_macro_loss": 0.5448069256705085}439{"step": 2180, "records_seen": 69760, "elapsed_seconds": 12648.871802806854, "gradient_norm": 2.757772445678711, "learning_rate": 8.600489066170846e-06, "head_lr": 4.300244533085422e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.38070170111649304, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.38070170111649304, "supervised_tokens": 119.9375, "waves": 1.0, "detection_loss": 0.44681886051382336, "grounding_loss": 0.3292772438074583, "detection_grounding_macro_loss": 0.38804805216064087}440{"step": 2190, "records_seen": 70080, "elapsed_seconds": 12702.516380786896, "gradient_norm": 1.5673917531967163, "learning_rate": 8.52345244770844e-06, "head_lr": 4.2617262238542195e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6513144092168659, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6513144092168659, "supervised_tokens": 147.53125, "waves": 1.0, "detection_loss": 0.7606195289464224, "grounding_loss": 0.4426409988240762, "detection_grounding_macro_loss": 0.6016302638852493}441{"step": 2200, "records_seen": 70400, "elapsed_seconds": 12758.965203046799, "gradient_norm": 1.7486793994903564, "learning_rate": 8.446611648900626e-06, "head_lr": 4.2233058244503126e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5589102617922492, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5589102617922492, "supervised_tokens": 95.21875, "waves": 1.0, "detection_loss": 0.8409042850136756, "grounding_loss": 0.3100920060086376, "detection_grounding_macro_loss": 0.5754981455111566}442{"step": 2210, "records_seen": 70720, "elapsed_seconds": 12829.194016933441, "gradient_norm": 1.673871397972107, "learning_rate": 8.36997274551957e-06, "head_lr": 4.1849863727597845e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5418203479293595, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5418203479293595, "supervised_tokens": 118.21875, "waves": 1.0, "detection_loss": 0.6820001760634649, "grounding_loss": 0.3829498760440401, "detection_grounding_macro_loss": 0.5324750260537525}443{"step": 2220, "records_seen": 71040, "elapsed_seconds": 12903.519265413284, "gradient_norm": 1.8601657152175903, "learning_rate": 8.293541797373643e-06, "head_lr": 4.146770898686821e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.49718741816468537, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49718741816468537, "supervised_tokens": 88.0, "waves": 1.0, "detection_loss": 0.7504684085195715, "grounding_loss": 0.3645164232168879, "detection_grounding_macro_loss": 0.5574924158682297}444{"step": 2230, "records_seen": 71360, "elapsed_seconds": 12963.559251070023, "gradient_norm": 1.6447746753692627, "learning_rate": 8.217324847828276e-06, "head_lr": 4.1086624239141377e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5106079244042121, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5106079244042121, "supervised_tokens": 120.78125, "waves": 1.0, "detection_loss": 0.7015110211543895, "grounding_loss": 0.3421640155069968, "detection_grounding_macro_loss": 0.5218375183306931}445{"step": 2240, "records_seen": 71680, "elapsed_seconds": 13016.567190408707, "gradient_norm": 2.1645865440368652, "learning_rate": 8.141327923328108e-06, "head_lr": 4.0706639616640535e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6194273729342967, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6194273729342967, "supervised_tokens": 92.53125, "waves": 1.0, "detection_loss": 0.7986663645681213, "grounding_loss": 0.4162898490826289, "detection_grounding_macro_loss": 0.6074781068253752}446{"step": 2250, "records_seen": 72000, "elapsed_seconds": 13072.122273921967, "gradient_norm": 1.7908873558044434, "learning_rate": 8.065557032920495e-06, "head_lr": 4.032778516460247e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6455050455015225, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6455050455015225, "supervised_tokens": 81.5, "waves": 1.0, "detection_loss": 0.7567428514201311, "grounding_loss": 0.4601087023038417, "detection_grounding_macro_loss": 0.6084257768619864}447{"step": 2260, "records_seen": 72320, "elapsed_seconds": 13127.351072788239, "gradient_norm": 1.7266349792480469, "learning_rate": 7.990018167780358e-06, "head_lr": 3.9950090838901784e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.49411504094931047, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49411504094931047, "supervised_tokens": 118.03125, "waves": 1.0, "detection_loss": 0.6671608873327974, "grounding_loss": 0.29799641504802515, "detection_grounding_macro_loss": 0.4825786511904113}448{"step": 2270, "records_seen": 72640, "elapsed_seconds": 13185.496871471405, "gradient_norm": 1.534917950630188, "learning_rate": 7.91471730073647e-06, "head_lr": 3.9573586503682344e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.536720982985571, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.536720982985571, "supervised_tokens": 128.59375, "waves": 1.0, "detection_loss": 0.7257014577587445, "grounding_loss": 0.36997350524453554, "detection_grounding_macro_loss": 0.54783748150164}449{"step": 2280, "records_seen": 72960, "elapsed_seconds": 13241.260455608368, "gradient_norm": 1.895872950553894, "learning_rate": 7.83966038579919e-06, "head_lr": 3.919830192899595e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6277191841461445, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6277191841461445, "supervised_tokens": 89.5, "waves": 1.0, "detection_loss": 0.837385445344433, "grounding_loss": 0.3900974214547508, "detection_grounding_macro_loss": 0.6137414333995919}450{"step": 2290, "records_seen": 73280, "elapsed_seconds": 13292.483290672302, "gradient_norm": 1.528549313545227, "learning_rate": 7.76485335768967e-06, "head_lr": 3.882426678844835e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5827998808817938, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5827998808817938, "supervised_tokens": 162.09375, "waves": 1.0, "detection_loss": 0.706499342407499, "grounding_loss": 0.34664636342362926, "detection_grounding_macro_loss": 0.5265728529155641}451{"step": 2300, "records_seen": 73600, "elapsed_seconds": 13345.895305871964, "gradient_norm": 1.502487301826477, "learning_rate": 7.69030213137062e-06, "head_lr": 3.8451510656853094e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.44374339842661215, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.44374339842661215, "supervised_tokens": 107.71875, "waves": 1.0, "detection_loss": 0.6042162785208134, "grounding_loss": 0.2618741343198508, "detection_grounding_macro_loss": 0.43304520642033206}452{"step": 2310, "records_seen": 73920, "elapsed_seconds": 13400.718010902405, "gradient_norm": 1.8708035945892334, "learning_rate": 7.61601260157859e-06, "head_lr": 3.8080063007892947e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5843227733608956, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5843227733608956, "supervised_tokens": 102.59375, "waves": 1.0, "detection_loss": 0.6406578674920768, "grounding_loss": 0.4767739572922768, "detection_grounding_macro_loss": 0.5587159123921768}453{"step": 2320, "records_seen": 74240, "elapsed_seconds": 13453.179865837097, "gradient_norm": 1.4042747020721436, "learning_rate": 7.541990642357897e-06, "head_lr": 3.770995321178948e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4261374425937845, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4261374425937845, "supervised_tokens": 133.5, "waves": 1.0, "detection_loss": 0.5704116309663125, "grounding_loss": 0.2988366881474362, "detection_grounding_macro_loss": 0.4346241595568744}454{"step": 2330, "records_seen": 74560, "elapsed_seconds": 13506.013463973999, "gradient_norm": 1.9670541286468506, "learning_rate": 7.4682421065961485e-06, "head_lr": 3.734121053298074e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.48565773575209903, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.48565773575209903, "supervised_tokens": 105.0625, "waves": 1.0, "detection_loss": 0.5934808653507692, "grounding_loss": 0.39051968022386063, "detection_grounding_macro_loss": 0.4920002727873149}455{"step": 2340, "records_seen": 74880, "elapsed_seconds": 13561.18638753891, "gradient_norm": 1.4664356708526611, "learning_rate": 7.394772825561476e-06, "head_lr": 3.6973864127807374e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5477814468217161, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5477814468217161, "supervised_tokens": 140.21875, "waves": 1.0, "detection_loss": 0.6894129447361647, "grounding_loss": 0.38726574918534123, "detection_grounding_macro_loss": 0.5383393469607529}456{"step": 2350, "records_seen": 75200, "elapsed_seconds": 13615.01467871666, "gradient_norm": 1.9127408266067505, "learning_rate": 7.32158860844143e-06, "head_lr": 3.6607943042207144e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.6200954575397191, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6200954575397191, "supervised_tokens": 131.0, "waves": 1.0, "detection_loss": 0.7280994481739721, "grounding_loss": 0.4139060208743269, "detection_grounding_macro_loss": 0.5710027345241495}457{"step": 2360, "records_seen": 75520, "elapsed_seconds": 13668.680253744125, "gradient_norm": 1.7001007795333862, "learning_rate": 7.248695241883683e-06, "head_lr": 3.624347620941841e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.45410130839610474, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.45410130839610474, "supervised_tokens": 112.90625, "waves": 1.0, "detection_loss": 0.6104797105645048, "grounding_loss": 0.2977229062277047, "detection_grounding_macro_loss": 0.45410130839610474}458{"step": 2370, "records_seen": 75840, "elapsed_seconds": 13718.816681623459, "gradient_norm": 1.5126839876174927, "learning_rate": 7.176098489538461e-06, "head_lr": 3.58804924476923e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5441180162588353, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5441180162588353, "supervised_tokens": 126.34375, "waves": 1.0, "detection_loss": 0.8236506720986654, "grounding_loss": 0.26458536041900516, "detection_grounding_macro_loss": 0.5441180162588353}459{"step": 2380, "records_seen": 76160, "elapsed_seconds": 13770.826681137085, "gradient_norm": 1.5241918563842773, "learning_rate": 7.103804091602824e-06, "head_lr": 3.5519020458014117e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.631089172209613, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.631089172209613, "supervised_tokens": 140.5, "waves": 1.0, "detection_loss": 0.819159625784347, "grounding_loss": 0.3562169708311558, "detection_grounding_macro_loss": 0.5876882983077514}460{"step": 2390, "records_seen": 76480, "elapsed_seconds": 13824.022667169571, "gradient_norm": 1.6352723836898804, "learning_rate": 7.0318177643667795e-06, "head_lr": 3.515908882183389e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4435084595372274, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4435084595372274, "supervised_tokens": 126.53125, "waves": 1.0, "detection_loss": 0.5379093175946955, "grounding_loss": 0.30553797468400473, "detection_grounding_macro_loss": 0.4217236461393501}461{"step": 2400, "records_seen": 76800, "elapsed_seconds": 13877.6836373806, "gradient_norm": 1.6274195909500122, "learning_rate": 6.960145199761305e-06, "head_lr": 3.480072599880652e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5379319320680906, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5379319320680906, "supervised_tokens": 101.03125, "waves": 1.0, "detection_loss": 0.6563642612496551, "grounding_loss": 0.3405447167654832, "detection_grounding_macro_loss": 0.49845448900756917}462/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead463  warnings.warn(464/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead465  warnings.warn(466/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead467  warnings.warn(468/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead469  warnings.warn(470/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead471  warnings.warn(472/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead473  warnings.warn(474/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead475  warnings.warn(476/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead477  warnings.warn(478{"validation_step": 2400, "loss": 1.1314715034751672, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1314715034751672, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6459809753847934, "grounding_loss": 0.5478707596233913, "detection_grounding_macro_loss": 1.0969258675040923}479{"step": 2410, "records_seen": 77120, "elapsed_seconds": 14077.677326202393, "gradient_norm": 1.9352312088012695, "learning_rate": 6.888792064908292e-06, "head_lr": 3.4443960324541453e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5499118579386391, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5499118579386391, "supervised_tokens": 114.9375, "waves": 1.0, "detection_loss": 0.6740675781454358, "grounding_loss": 0.31288730118020874, "detection_grounding_macro_loss": 0.4934774396628223}480{"step": 2420, "records_seen": 77440, "elapsed_seconds": 14132.70022034645, "gradient_norm": 2.2068207263946533, "learning_rate": 6.817764001672444e-06, "head_lr": 3.408882000836221e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.521621680676617, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.521621680676617, "supervised_tokens": 107.65625, "waves": 1.0, "detection_loss": 0.6852958874626305, "grounding_loss": 0.336124246319135, "detection_grounding_macro_loss": 0.5107100668908827}481{"step": 2430, "records_seen": 77760, "elapsed_seconds": 14187.218351125717, "gradient_norm": 1.6973398923873901, "learning_rate": 6.747066626215175e-06, "head_lr": 3.373533313107587e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5769502028588249, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5769502028588249, "supervised_tokens": 68.71875, "waves": 1.0, "detection_loss": 0.7509235168639433, "grounding_loss": 0.3797804469863574, "detection_grounding_macro_loss": 0.5653519819251503}482{"step": 2440, "records_seen": 78080, "elapsed_seconds": 14240.621819734573, "gradient_norm": 1.5252678394317627, "learning_rate": 6.676705528550551e-06, "head_lr": 3.3383527642752754e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.4916834856283572, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4916834856283572, "supervised_tokens": 117.4375, "waves": 1.0, "detection_loss": 0.6050473104511515, "grounding_loss": 0.24228307101820973, "detection_grounding_macro_loss": 0.42366519073468056}483{"step": 2450, "records_seen": 78400, "elapsed_seconds": 14293.919555902481, "gradient_norm": 1.5471141338348389, "learning_rate": 6.606686272103281e-06, "head_lr": 3.30334313605164e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5696390736989088, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5696390736989088, "supervised_tokens": 113.84375, "waves": 1.0, "detection_loss": 0.7548982573716785, "grounding_loss": 0.331448694691062, "detection_grounding_macro_loss": 0.5431734760313702}484{"step": 2460, "records_seen": 78720, "elapsed_seconds": 14346.470924854279, "gradient_norm": 1.800957441329956, "learning_rate": 6.537014393268812e-06, "head_lr": 3.268507196634406e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.500697260794368, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.500697260794368, "supervised_tokens": 87.15625, "waves": 1.0, "detection_loss": 0.6314615702842061, "grounding_loss": 0.35249771003921826, "detection_grounding_macro_loss": 0.4919796401617122}485{"step": 2470, "records_seen": 79040, "elapsed_seconds": 14400.323899507523, "gradient_norm": 1.499625325202942, "learning_rate": 6.467695400975591e-06, "head_lr": 3.2338477004877954e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.472424949956757, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.472424949956757, "supervised_tokens": 114.0, "waves": 1.0, "detection_loss": 0.6851828349754214, "grounding_loss": 0.25966706493809255, "detection_grounding_macro_loss": 0.472424949956757}486{"step": 2480, "records_seen": 79360, "elapsed_seconds": 14453.837069034576, "gradient_norm": 1.6053558588027954, "learning_rate": 6.3987347762494565e-06, "head_lr": 3.199367388124728e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5384273562486186, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5384273562486186, "supervised_tokens": 160.0, "waves": 1.0, "detection_loss": 0.6059026858752835, "grounding_loss": 0.4516733610143352, "detection_grounding_macro_loss": 0.5287880234448094}487{"step": 2490, "records_seen": 79680, "elapsed_seconds": 14507.271722316742, "gradient_norm": 1.934710144996643, "learning_rate": 6.330137971780264e-06, "head_lr": 3.165068985890132e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5173210598416063, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5173210598416063, "supervised_tokens": 92.65625, "waves": 1.0, "detection_loss": 0.6132228566891101, "grounding_loss": 0.4517040409459457, "detection_grounding_macro_loss": 0.5324634488175279}488{"step": 2500, "records_seen": 80000, "elapsed_seconds": 14559.542356729507, "gradient_norm": 1.378966212272644, "learning_rate": 6.261910411490744e-06, "head_lr": 3.130955205745371e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.49773400055710226, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49773400055710226, "supervised_tokens": 177.40625, "waves": 1.0, "detection_loss": 0.6021746822997279, "grounding_loss": 0.23083003610372543, "detection_grounding_macro_loss": 0.4165023592017267}489{"step": 2510, "records_seen": 80320, "elapsed_seconds": 14609.210901975632, "gradient_norm": 1.8619816303253174, "learning_rate": 6.194057490107631e-06, "head_lr": 3.0970287450538154e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.5901109692640603, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5901109692640603, "supervised_tokens": 71.28125, "waves": 1.0, "detection_loss": 0.8154470932980379, "grounding_loss": 0.4549092948436737, "detection_grounding_macro_loss": 0.6351781940708559}490{"step": 2520, "records_seen": 80640, "elapsed_seconds": 14662.70007967949, "gradient_norm": 1.7825227975845337, "learning_rate": 6.126584572735101e-06, "head_lr": 3.0632922863675503e-07, "peak_gpu_gib": 19.779818534851074, "loss": 0.528087927854358, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.528087927854358, "supervised_tokens": 108.59375, "waves": 1.0, "detection_loss": 0.6886212805906932, "grounding_loss": 0.38644085191053285, "detection_grounding_macro_loss": 0.537531066250613}491{"step": 2530, "records_seen": 80960, "elapsed_seconds": 14714.993160009384, "gradient_norm": 1.7917428016662598, "learning_rate": 6.0594969944305715e-06, "head_lr": 3.0297484972152853e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.7416594649120469, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7416594649120469, "supervised_tokens": 126.375, "waves": 1.0, "detection_loss": 0.9996504402822919, "grounding_loss": 0.4099567822931606, "detection_grounding_macro_loss": 0.7048036112877263}492{"step": 2540, "records_seen": 81280, "elapsed_seconds": 14767.386165380478, "gradient_norm": 1.2850710153579712, "learning_rate": 5.99280005978284e-06, "head_lr": 2.9964000298914195e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.3619575907154058, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3619575907154058, "supervised_tokens": 120.1875, "waves": 1.0, "detection_loss": 0.48260465468624986, "grounding_loss": 0.2252242515484492, "detection_grounding_macro_loss": 0.35391445311734954}493{"step": 2550, "records_seen": 81600, "elapsed_seconds": 14820.296340227127, "gradient_norm": 1.9000831842422485, "learning_rate": 5.926499042492667e-06, "head_lr": 2.9632495212463334e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5225414449702512, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5225414449702512, "supervised_tokens": 91.25, "waves": 1.0, "detection_loss": 0.6752103609212554, "grounding_loss": 0.36987252901924705, "detection_grounding_macro_loss": 0.5225414449702512}494{"step": 2560, "records_seen": 81920, "elapsed_seconds": 14871.806970834732, "gradient_norm": 1.9194813966751099, "learning_rate": 5.860599184955782e-06, "head_lr": 2.9302995924778904e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5978464219952002, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5978464219952002, "supervised_tokens": 86.34375, "waves": 1.0, "detection_loss": 0.7992517660061519, "grounding_loss": 0.4201358243384782, "detection_grounding_macro_loss": 0.609693795172315}495{"step": 2570, "records_seen": 82240, "elapsed_seconds": 14925.135642051697, "gradient_norm": 1.5340632200241089, "learning_rate": 5.795105697848358e-06, "head_lr": 2.8975528489241783e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5790373658037424, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5790373658037424, "supervised_tokens": 123.5625, "waves": 1.0, "detection_loss": 0.7240220862348858, "grounding_loss": 0.3671366205582252, "detection_grounding_macro_loss": 0.5455793533965555}496{"step": 2580, "records_seen": 82560, "elapsed_seconds": 14978.162710905075, "gradient_norm": 1.5785589218139648, "learning_rate": 5.7300237597150285e-06, "head_lr": 2.865011879857514e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.3924445166812802, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.3924445166812802, "supervised_tokens": 107.875, "waves": 1.0, "detection_loss": 0.41756015374496774, "grounding_loss": 0.35573704712666, "detection_grounding_macro_loss": 0.38664860043581384}497{"step": 2590, "records_seen": 82880, "elapsed_seconds": 15032.208126068115, "gradient_norm": 1.4036566019058228, "learning_rate": 5.6653585165594024e-06, "head_lr": 2.832679258279701e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.42376665416604453, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.42376665416604453, "supervised_tokens": 141.5, "waves": 1.0, "detection_loss": 0.6496177017688751, "grounding_loss": 0.24810472825273186, "detection_grounding_macro_loss": 0.4488612150108035}498{"step": 2600, "records_seen": 83200, "elapsed_seconds": 15083.060338020325, "gradient_norm": 1.9912054538726807, "learning_rate": 5.60111508143718e-06, "head_lr": 2.80055754071859e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5474952261936323, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5474952261936323, "supervised_tokens": 91.125, "waves": 1.0, "detection_loss": 0.7465082635852227, "grounding_loss": 0.3219471171498299, "detection_grounding_macro_loss": 0.5342276903675263}499{"step": 2610, "records_seen": 83520, "elapsed_seconds": 15135.582064151764, "gradient_norm": 1.6539329290390015, "learning_rate": 5.537298534051863e-06, "head_lr": 2.7686492670259316e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5152235234224065, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5152235234224065, "supervised_tokens": 93.46875, "waves": 1.0, "detection_loss": 0.628010770974879, "grounding_loss": 0.32724477750161896, "detection_grounding_macro_loss": 0.477627774238249}500{"step": 2620, "records_seen": 83840, "elapsed_seconds": 15190.5890583992, "gradient_norm": 1.6681369543075562, "learning_rate": 5.473913920353112e-06, "head_lr": 2.736956960176556e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5650966811690523, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5650966811690523, "supervised_tokens": 101.84375, "waves": 1.0, "detection_loss": 0.6825700371670169, "grounding_loss": 0.39340485317202717, "detection_grounding_macro_loss": 0.537987445169522}501{"step": 2630, "records_seen": 84160, "elapsed_seconds": 15244.755842208862, "gradient_norm": 1.6883068084716797, "learning_rate": 5.410966252137744e-06, "head_lr": 2.7054831260688716e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5909364250114919, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5909364250114919, "supervised_tokens": 107.03125, "waves": 1.0, "detection_loss": 0.7936249932114917, "grounding_loss": 0.29469928687303065, "detection_grounding_macro_loss": 0.5441621400422612}502{"step": 2640, "records_seen": 84480, "elapsed_seconds": 15298.72585272789, "gradient_norm": 1.7852729558944702, "learning_rate": 5.348460506653478e-06, "head_lr": 2.6742302533267384e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5683241533943146, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5683241533943146, "supervised_tokens": 81.25, "waves": 1.0, "detection_loss": 0.7337870423274581, "grounding_loss": 0.4223274866885999, "detection_grounding_macro_loss": 0.578057264508029}503{"step": 2650, "records_seen": 84800, "elapsed_seconds": 15351.472549438477, "gradient_norm": 1.5006030797958374, "learning_rate": 5.28640162620537e-06, "head_lr": 2.643200813102685e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.40170558937987266, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.40170558937987266, "supervised_tokens": 94.4375, "waves": 1.0, "detection_loss": 0.5414969011482234, "grounding_loss": 0.2929790135600443, "detection_grounding_macro_loss": 0.4172379573541338}504{"step": 2660, "records_seen": 85120, "elapsed_seconds": 15401.41870880127, "gradient_norm": 1.6305044889450073, "learning_rate": 5.224794517765035e-06, "head_lr": 2.612397258882517e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5034658723855046, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5034658723855046, "supervised_tokens": 138.40625, "waves": 1.0, "detection_loss": 0.5227075370206876, "grounding_loss": 0.47139643132686615, "detection_grounding_macro_loss": 0.4970519841737769}505{"step": 2670, "records_seen": 85440, "elapsed_seconds": 15454.19595527649, "gradient_norm": 1.7323193550109863, "learning_rate": 5.163644052582646e-06, "head_lr": 2.581822026291323e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4946743080762417, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4946743080762417, "supervised_tokens": 110.46875, "waves": 1.0, "detection_loss": 0.5554412837880136, "grounding_loss": 0.43390733236446977, "detection_grounding_macro_loss": 0.4946743080762417}506{"step": 2680, "records_seen": 85760, "elapsed_seconds": 15506.616112470627, "gradient_norm": 1.619234561920166, "learning_rate": 5.102955065801769e-06, "head_lr": 2.551477532900884e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.47146780392893106, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.47146780392893106, "supervised_tokens": 109.5, "waves": 1.0, "detection_loss": 0.6445293120177666, "grounding_loss": 0.24896015067185676, "detection_grounding_macro_loss": 0.4467447313448117}507{"step": 2690, "records_seen": 86080, "elapsed_seconds": 15560.333193063736, "gradient_norm": 1.5630254745483398, "learning_rate": 5.042732356077055e-06, "head_lr": 2.521366178038527e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.42296873421287273, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.42296873421287273, "supervised_tokens": 87.65625, "waves": 1.0, "detection_loss": 0.5242460566817164, "grounding_loss": 0.2749480321430243, "detection_grounding_macro_loss": 0.3995970444123704}508{"step": 2700, "records_seen": 86400, "elapsed_seconds": 15612.259193181992, "gradient_norm": 1.6479605436325073, "learning_rate": 4.982980685194808e-06, "head_lr": 2.491490342597404e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4557359388236364, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4557359388236364, "supervised_tokens": 129.25, "waves": 1.0, "detection_loss": 0.48577946667016175, "grounding_loss": 0.4171085458781038, "detection_grounding_macro_loss": 0.4514440062741328}509/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead510  warnings.warn(511/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead512  warnings.warn(513/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead514  warnings.warn(515/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead516  warnings.warn(517/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead518  warnings.warn(519/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead520  warnings.warn(521/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead522  warnings.warn(523/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead524  warnings.warn(525{"validation_step": 2700, "loss": 1.125624546662853, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.125624546662853, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6361503183670216, "grounding_loss": 0.5465424570441246, "detection_grounding_macro_loss": 1.0913463877055731}526{"step": 2710, "records_seen": 86720, "elapsed_seconds": 15805.036520004272, "gradient_norm": 1.6086124181747437, "learning_rate": 4.923704777696477e-06, "head_lr": 2.461852388848238e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6005625802456365, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6005625802456365, "supervised_tokens": 95.375, "waves": 1.0, "detection_loss": 0.809639719405303, "grounding_loss": 0.20141531457718223, "detection_grounding_macro_loss": 0.5055275169912427}527{"step": 2720, "records_seen": 87040, "elapsed_seconds": 15858.445959806442, "gradient_norm": 1.902151346206665, "learning_rate": 4.864909320505079e-06, "head_lr": 2.4324546602525394e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6152209413955845, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6152209413955845, "supervised_tokens": 97.65625, "waves": 1.0, "detection_loss": 0.8024837833502152, "grounding_loss": 0.25771915220947117, "detection_grounding_macro_loss": 0.5301014677798432}528{"step": 2730, "records_seen": 87360, "elapsed_seconds": 15910.863594055176, "gradient_norm": 1.5713136196136475, "learning_rate": 4.806598962554619e-06, "head_lr": 2.403299481277309e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5372277762533031, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5372277762533031, "supervised_tokens": 103.3125, "waves": 1.0, "detection_loss": 0.7503453202599, "grounding_loss": 0.3491828844827764, "detection_grounding_macro_loss": 0.5497641023713382}529{"step": 2740, "records_seen": 87680, "elapsed_seconds": 15964.521287679672, "gradient_norm": 1.5320160388946533, "learning_rate": 4.748778314422481e-06, "head_lr": 2.3743891572112398e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5303356057392392, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5303356057392392, "supervised_tokens": 117.0625, "waves": 1.0, "detection_loss": 0.730461035581196, "grounding_loss": 0.30352678525168814, "detection_grounding_macro_loss": 0.516993910416442}530{"step": 2750, "records_seen": 88000, "elapsed_seconds": 16018.370273590088, "gradient_norm": 1.9721816778182983, "learning_rate": 4.6914519479648935e-06, "head_lr": 2.3457259739824465e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6417217777930659, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6417217777930659, "supervised_tokens": 95.65625, "waves": 1.0, "detection_loss": 0.8086279762699311, "grounding_loss": 0.45256141951928536, "detection_grounding_macro_loss": 0.6305946978946082}531{"step": 2760, "records_seen": 88320, "elapsed_seconds": 16070.941900491714, "gradient_norm": 2.031900644302368, "learning_rate": 4.634624395955423e-06, "head_lr": 2.3173121979777112e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6787001534306114, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6787001534306114, "supervised_tokens": 92.5625, "waves": 1.0, "detection_loss": 0.8598049638073592, "grounding_loss": 0.5378408564709187, "detection_grounding_macro_loss": 0.698822910139139}532{"step": 2770, "records_seen": 88640, "elapsed_seconds": 16121.19223856926, "gradient_norm": 1.9973143339157104, "learning_rate": 4.5783001517265705e-06, "head_lr": 2.2891500758632847e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.560537207235825, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.560537207235825, "supervised_tokens": 93.375, "waves": 1.0, "detection_loss": 0.7350583035034876, "grounding_loss": 0.4411280361053191, "detection_grounding_macro_loss": 0.5880931698044033}533{"step": 2780, "records_seen": 88960, "elapsed_seconds": 16173.841977596283, "gradient_norm": 1.5526503324508667, "learning_rate": 4.522483668814484e-06, "head_lr": 2.2612418344072418e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5110574166610604, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5110574166610604, "supervised_tokens": 116.6875, "waves": 1.0, "detection_loss": 0.7930569063800459, "grounding_loss": 0.3181103973796493, "detection_grounding_macro_loss": 0.5555836518798476}534{"step": 2790, "records_seen": 89280, "elapsed_seconds": 16224.754898071289, "gradient_norm": 1.854884147644043, "learning_rate": 4.467179360606824e-06, "head_lr": 2.2335896803034114e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5292978405568647, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5292978405568647, "supervised_tokens": 88.875, "waves": 1.0, "detection_loss": 0.832324136630632, "grounding_loss": 0.2262715444830974, "detection_grounding_macro_loss": 0.5292978405568647}535{"step": 2800, "records_seen": 89600, "elapsed_seconds": 16281.912228822708, "gradient_norm": 1.5173200368881226, "learning_rate": 4.412391599993796e-06, "head_lr": 2.2061957999968979e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4431106180371245, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4431106180371245, "supervised_tokens": 106.1875, "waves": 1.0, "detection_loss": 0.5872296717104746, "grounding_loss": 0.2989915643637744, "detection_grounding_macro_loss": 0.4431106180371245}536{"step": 2810, "records_seen": 89920, "elapsed_seconds": 16334.752942800522, "gradient_norm": 1.8361176252365112, "learning_rate": 4.358124719022381e-06, "head_lr": 2.1790623595111903e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5642450649756938, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5642450649756938, "supervised_tokens": 110.65625, "waves": 1.0, "detection_loss": 0.6929837386397755, "grounding_loss": 0.418341234823068, "detection_grounding_macro_loss": 0.5556624867314217}537{"step": 2820, "records_seen": 90240, "elapsed_seconds": 16386.49912261963, "gradient_norm": 1.6393862962722778, "learning_rate": 4.304383008553816e-06, "head_lr": 2.152191504276908e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4956308415030861, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4956308415030861, "supervised_tokens": 129.4375, "waves": 1.0, "detection_loss": 0.6169608605750374, "grounding_loss": 0.35812348655487464, "detection_grounding_macro_loss": 0.487542173564956}538{"step": 2830, "records_seen": 90560, "elapsed_seconds": 16440.898027658463, "gradient_norm": 1.573603630065918, "learning_rate": 4.251170717924307e-06, "head_lr": 2.1255853589621533e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5327767041862899, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5327767041862899, "supervised_tokens": 115.875, "waves": 1.0, "detection_loss": 0.6669389954651705, "grounding_loss": 0.3091728853881553, "detection_grounding_macro_loss": 0.4880559404266629}539{"step": 2840, "records_seen": 90880, "elapsed_seconds": 16491.372393369675, "gradient_norm": 1.4741544723510742, "learning_rate": 4.198492054609041e-06, "head_lr": 2.09924602730452e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.44144476628876106, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.44144476628876106, "supervised_tokens": 86.53125, "waves": 1.0, "detection_loss": 0.6545898678092011, "grounding_loss": 0.2756652428839743, "detection_grounding_macro_loss": 0.46512755534658773}540{"step": 2850, "records_seen": 91200, "elapsed_seconds": 16544.729294776917, "gradient_norm": 1.8191142082214355, "learning_rate": 4.146351183889496e-06, "head_lr": 2.0731755919447473e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.689710873819422, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.689710873819422, "supervised_tokens": 98.125, "waves": 1.0, "detection_loss": 1.0449385751477058, "grounding_loss": 0.28711947898070017, "detection_grounding_macro_loss": 0.666029027064203}541{"step": 2860, "records_seen": 91520, "elapsed_seconds": 16598.340760946274, "gradient_norm": 1.9046311378479004, "learning_rate": 4.094752228524102e-06, "head_lr": 2.0473761142620504e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.516862898055706, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.516862898055706, "supervised_tokens": 91.78125, "waves": 1.0, "detection_loss": 0.6272883825651278, "grounding_loss": 0.35547180531116634, "detection_grounding_macro_loss": 0.4913800939381471}542{"step": 2870, "records_seen": 91840, "elapsed_seconds": 16650.462400197983, "gradient_norm": 1.7892447710037231, "learning_rate": 4.043699268422253e-06, "head_lr": 2.021849634211126e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5082990679049999, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5082990679049999, "supervised_tokens": 101.53125, "waves": 1.0, "detection_loss": 0.7880324395373464, "grounding_loss": 0.2907286677465082, "detection_grounding_macro_loss": 0.5393805536419273}543{"step": 2880, "records_seen": 92160, "elapsed_seconds": 16703.611969470978, "gradient_norm": 1.6733806133270264, "learning_rate": 3.9931963403217e-06, "head_lr": 1.9965981701608494e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5925489406451447, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5925489406451447, "supervised_tokens": 104.09375, "waves": 1.0, "detection_loss": 0.8029568693944498, "grounding_loss": 0.3821410118958397, "detection_grounding_macro_loss": 0.5925489406451447}544{"step": 2890, "records_seen": 92480, "elapsed_seconds": 16756.366561174393, "gradient_norm": 1.7829669713974, "learning_rate": 3.943247437469387e-06, "head_lr": 1.9716237187346935e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.7182712730279093, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7182712730279093, "supervised_tokens": 106.5, "waves": 1.0, "detection_loss": 1.020559478449286, "grounding_loss": 0.4159830676065326, "detection_grounding_macro_loss": 0.7182712730279093}545{"step": 2900, "records_seen": 92800, "elapsed_seconds": 16809.533203840256, "gradient_norm": 1.4469399452209473, "learning_rate": 3.893856509305695e-06, "head_lr": 1.946928254652847e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4486628439356366, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4486628439356366, "supervised_tokens": 138.28125, "waves": 1.0, "detection_loss": 0.5757781260146754, "grounding_loss": 0.26287897012781075, "detection_grounding_macro_loss": 0.4193285480712431}546{"step": 2910, "records_seen": 93120, "elapsed_seconds": 16860.62432217598, "gradient_norm": 1.6848362684249878, "learning_rate": 3.8450274611521585e-06, "head_lr": 1.9225137305760792e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5329493709141389, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5329493709141389, "supervised_tokens": 119.84375, "waves": 1.0, "detection_loss": 0.7276481890252658, "grounding_loss": 0.38151695682770675, "detection_grounding_macro_loss": 0.5545825729264863}547{"step": 2920, "records_seen": 93440, "elapsed_seconds": 16911.25654578209, "gradient_norm": 1.5800414085388184, "learning_rate": 3.7967641539026806e-06, "head_lr": 1.8983820769513402e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4153833056682288, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4153833056682288, "supervised_tokens": 120.4375, "waves": 1.0, "detection_loss": 0.5251260005679796, "grounding_loss": 0.31855151605080156, "detection_grounding_macro_loss": 0.4218387583093906}548{"step": 2930, "records_seen": 93760, "elapsed_seconds": 16963.10364151001, "gradient_norm": 2.0548017024993896, "learning_rate": 3.74907040371825e-06, "head_lr": 1.874535201859125e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.47563316342504436, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.47563316342504436, "supervised_tokens": 103.71875, "waves": 1.0, "detection_loss": 0.6640359065513621, "grounding_loss": 0.2872304202987266, "detection_grounding_macro_loss": 0.47563316342504436}549{"step": 2940, "records_seen": 94080, "elapsed_seconds": 17015.091328144073, "gradient_norm": 1.5210353136062622, "learning_rate": 3.7019499817252004e-06, "head_lr": 1.8509749908626e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4867841436456928, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4867841436456928, "supervised_tokens": 129.28125, "waves": 1.0, "detection_loss": 0.7061464435913983, "grounding_loss": 0.2381735370405598, "detection_grounding_macro_loss": 0.47215999031597905}550{"step": 2950, "records_seen": 94400, "elapsed_seconds": 17069.33596253395, "gradient_norm": 1.8339614868164062, "learning_rate": 3.655406613717023e-06, "head_lr": 1.8277033068585112e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5007660346167313, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5007660346167313, "supervised_tokens": 115.6875, "waves": 1.0, "detection_loss": 0.6014082980649879, "grounding_loss": 0.3086308043973325, "detection_grounding_macro_loss": 0.45501955123116017}551{"step": 2960, "records_seen": 94720, "elapsed_seconds": 17122.518037080765, "gradient_norm": 2.1342341899871826, "learning_rate": 3.609443979859778e-06, "head_lr": 1.8047219899298888e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5657148951577256, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5657148951577256, "supervised_tokens": 82.5625, "waves": 1.0, "detection_loss": 0.7958767905656714, "grounding_loss": 0.3355529997497797, "detection_grounding_macro_loss": 0.5657148951577256}552{"step": 2970, "records_seen": 95040, "elapsed_seconds": 17175.102278470993, "gradient_norm": 1.880021572113037, "learning_rate": 3.5640657144011025e-06, "head_lr": 1.782032857200551e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5114755911210978, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5114755911210978, "supervised_tokens": 118.59375, "waves": 1.0, "detection_loss": 0.5811462740261959, "grounding_loss": 0.4418049082159996, "detection_grounding_macro_loss": 0.5114755911210978}553{"step": 2980, "records_seen": 95360, "elapsed_seconds": 17228.897544145584, "gradient_norm": 1.5294742584228516, "learning_rate": 3.5192754053828487e-06, "head_lr": 1.759637702691424e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5994775266158499, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5994775266158499, "supervised_tokens": 123.75, "waves": 1.0, "detection_loss": 0.7918229638988123, "grounding_loss": 0.27890179781091246, "detection_grounding_macro_loss": 0.5353623808548624}554{"step": 2990, "records_seen": 95680, "elapsed_seconds": 17279.312748908997, "gradient_norm": 1.655634880065918, "learning_rate": 3.4750765943573806e-06, "head_lr": 1.73753829717869e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.663534054466453, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.663534054466453, "supervised_tokens": 103.4375, "waves": 1.0, "detection_loss": 0.8807207583820855, "grounding_loss": 0.30155621460706544, "detection_grounding_macro_loss": 0.5911384864945755}555{"step": 3000, "records_seen": 96000, "elapsed_seconds": 17331.708535909653, "gradient_norm": 2.0851943492889404, "learning_rate": 3.4314727761075482e-06, "head_lr": 1.7157363880537738e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5788579581874274, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5788579581874274, "supervised_tokens": 78.03125, "waves": 1.0, "detection_loss": 0.8665911655324245, "grounding_loss": 0.3550654635857629, "detection_grounding_macro_loss": 0.6108283145590937}556/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead557  warnings.warn(558/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead559  warnings.warn(560/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead561  warnings.warn(562/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead563  warnings.warn(564/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead565  warnings.warn(566/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead567  warnings.warn(568/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead569  warnings.warn(570/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead571  warnings.warn(572{"validation_step": 3000, "loss": 1.122324453958054, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.122324453958054, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.633136281888166, "grounding_loss": 0.5429178948487554, "detection_grounding_macro_loss": 1.0880270883684608}573{"step": 3010, "records_seen": 96320, "elapsed_seconds": 17525.063952684402, "gradient_norm": 1.881605863571167, "learning_rate": 3.3884673983703532e-06, "head_lr": 1.6942336991851764e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4864856610658421, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4864856610658421, "supervised_tokens": 95.8125, "waves": 1.0, "detection_loss": 0.5154081980469457, "grounding_loss": 0.46669655681561334, "detection_grounding_macro_loss": 0.4910523774312795}574{"step": 3020, "records_seen": 96640, "elapsed_seconds": 17580.38715338707, "gradient_norm": 1.4317913055419922, "learning_rate": 3.3460638615643316e-06, "head_lr": 1.6730319307821654e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4323180752749636, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4323180752749636, "supervised_tokens": 122.21875, "waves": 1.0, "detection_loss": 0.6445647417925872, "grounding_loss": 0.2870966718681685, "detection_grounding_macro_loss": 0.46583070683037786}575{"step": 3030, "records_seen": 96960, "elapsed_seconds": 17633.12707042694, "gradient_norm": 1.9288791418075562, "learning_rate": 3.3042655185206973e-06, "head_lr": 1.6521327592603484e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4963997732105696, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4963997732105696, "supervised_tokens": 85.0, "waves": 1.0, "detection_loss": 0.5861567551078224, "grounding_loss": 0.3946751937270164, "detection_grounding_macro_loss": 0.4904159744174194}576{"step": 3040, "records_seen": 97280, "elapsed_seconds": 17686.744015216827, "gradient_norm": 1.7173113822937012, "learning_rate": 3.263075674218227e-06, "head_lr": 1.6315378371091132e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5067404144997489, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5067404144997489, "supervised_tokens": 71.4375, "waves": 1.0, "detection_loss": 0.6294239068152628, "grounding_loss": 0.32743377188476713, "detection_grounding_macro_loss": 0.478428839350015}577{"step": 3050, "records_seen": 97600, "elapsed_seconds": 17740.03647708893, "gradient_norm": 1.5798614025115967, "learning_rate": 3.2224975855219354e-06, "head_lr": 1.6112487927609673e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5327305662315212, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5327305662315212, "supervised_tokens": 107.625, "waves": 1.0, "detection_loss": 0.8141207584617405, "grounding_loss": 0.2513403740013018, "detection_grounding_macro_loss": 0.5327305662315212}578{"step": 3060, "records_seen": 97920, "elapsed_seconds": 17793.44575238228, "gradient_norm": 1.625080943107605, "learning_rate": 3.1825344609255615e-06, "head_lr": 1.5912672304627806e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5315727206755412, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5315727206755412, "supervised_tokens": 121.59375, "waves": 1.0, "detection_loss": 0.703653099713847, "grounding_loss": 0.3594923416372353, "detection_grounding_macro_loss": 0.5315727206755412}579{"step": 3070, "records_seen": 98240, "elapsed_seconds": 17844.295268774033, "gradient_norm": 1.4467787742614746, "learning_rate": 3.143189460297872e-06, "head_lr": 1.5715947301489357e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.38082103710334536, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.38082103710334536, "supervised_tokens": 134.96875, "waves": 1.0, "detection_loss": 0.44453551985488476, "grounding_loss": 0.28769987000494157, "detection_grounding_macro_loss": 0.3661176949299132}580{"step": 3080, "records_seen": 98560, "elapsed_seconds": 17897.789343118668, "gradient_norm": 1.5557680130004883, "learning_rate": 3.104465694632804e-06, "head_lr": 1.5522328473164018e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5674107043964796, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5674107043964796, "supervised_tokens": 142.40625, "waves": 1.0, "detection_loss": 0.6350732177439071, "grounding_loss": 0.41855317503213885, "detection_grounding_macro_loss": 0.5268131963880229}581{"step": 3090, "records_seen": 98880, "elapsed_seconds": 17952.155151844025, "gradient_norm": 1.8075377941131592, "learning_rate": 3.066366225803495e-06, "head_lr": 1.5331831129017476e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6783830256049441, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6783830256049441, "supervised_tokens": 88.65625, "waves": 1.0, "detection_loss": 0.9863023819098089, "grounding_loss": 0.37046366930007935, "detection_grounding_macro_loss": 0.6783830256049441}582{"step": 3100, "records_seen": 99200, "elapsed_seconds": 18003.92463850975, "gradient_norm": 1.8166056871414185, "learning_rate": 3.028894066320171e-06, "head_lr": 1.5144470331600854e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5327157700521639, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5327157700521639, "supervised_tokens": 126.28125, "waves": 1.0, "detection_loss": 0.5357094225473702, "grounding_loss": 0.5300743119681582, "detection_grounding_macro_loss": 0.5328918672577643}583{"step": 3110, "records_seen": 99520, "elapsed_seconds": 18058.860979557037, "gradient_norm": 1.6093369722366333, "learning_rate": 2.9920521790919518e-06, "head_lr": 1.4960260895459756e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.49571375206141965, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49571375206141965, "supervised_tokens": 98.65625, "waves": 1.0, "detection_loss": 0.6003434896757383, "grounding_loss": 0.2959660711613568, "detection_grounding_macro_loss": 0.4481547804185475}584{"step": 3120, "records_seen": 99840, "elapsed_seconds": 18110.4629240036, "gradient_norm": 1.6246861219406128, "learning_rate": 2.9558434771925744e-06, "head_lr": 1.477921738596287e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4619192505910803, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4619192505910803, "supervised_tokens": 126.78125, "waves": 1.0, "detection_loss": 0.6937879307476275, "grounding_loss": 0.25732923868824453, "detection_grounding_macro_loss": 0.475558584717936}585{"step": 3130, "records_seen": 100160, "elapsed_seconds": 18163.055297136307, "gradient_norm": 1.6662334203720093, "learning_rate": 2.920270823630053e-06, "head_lr": 1.4601354118150265e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5315020771761994, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5315020771761994, "supervised_tokens": 109.9375, "waves": 1.0, "detection_loss": 0.8390933991556153, "grounding_loss": 0.22391075519678338, "detection_grounding_macro_loss": 0.5315020771761994}586{"step": 3140, "records_seen": 100480, "elapsed_seconds": 18216.045642852783, "gradient_norm": 1.9370899200439453, "learning_rate": 2.8853370311203133e-06, "head_lr": 1.4426685155601563e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5538930243953999, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5538930243953999, "supervised_tokens": 113.375, "waves": 1.0, "detection_loss": 0.6725360218827662, "grounding_loss": 0.380491720375403, "detection_grounding_macro_loss": 0.5265138711290847}587{"step": 3150, "records_seen": 100800, "elapsed_seconds": 18267.20052599907, "gradient_norm": 1.5831165313720703, "learning_rate": 2.851044861864781e-06, "head_lr": 1.4255224309323906e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.44346117101326854, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.44346117101326854, "supervised_tokens": 118.3125, "waves": 1.0, "detection_loss": 0.510142317158553, "grounding_loss": 0.3460041112624682, "detection_grounding_macro_loss": 0.4280732142105106}588{"step": 3160, "records_seen": 101120, "elapsed_seconds": 18323.340247154236, "gradient_norm": 1.7392481565475464, "learning_rate": 2.817397027331983e-06, "head_lr": 1.4086985136659912e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6208494537277147, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6208494537277147, "supervised_tokens": 140.5, "waves": 1.0, "detection_loss": 0.7582613423111892, "grounding_loss": 0.3585176664319905, "detection_grounding_macro_loss": 0.5583895043715899}589{"step": 3170, "records_seen": 101440, "elapsed_seconds": 18375.28396296501, "gradient_norm": 1.7692711353302002, "learning_rate": 2.784396188043146e-06, "head_lr": 1.3921980940215726e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5159254218349929, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5159254218349929, "supervised_tokens": 105.875, "waves": 1.0, "detection_loss": 0.7534046178417546, "grounding_loss": 0.33121938049640043, "detection_grounding_macro_loss": 0.5423119991690775}590{"step": 3180, "records_seen": 101760, "elapsed_seconds": 18427.451608181, "gradient_norm": 1.79482901096344, "learning_rate": 2.7520449533618376e-06, "head_lr": 1.3760224766809188e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6352717516510893, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6352717516510893, "supervised_tokens": 99.53125, "waves": 1.0, "detection_loss": 0.9057363857953953, "grounding_loss": 0.45021700197340625, "detection_grounding_macro_loss": 0.6779766938844007}591{"step": 3190, "records_seen": 102080, "elapsed_seconds": 18480.62703347206, "gradient_norm": 1.7164015769958496, "learning_rate": 2.720345881287635e-06, "head_lr": 1.3601729406438173e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.439938532934093, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.439938532934093, "supervised_tokens": 97.21875, "waves": 1.0, "detection_loss": 0.5656549480040934, "grounding_loss": 0.29745992918809255, "detection_grounding_macro_loss": 0.43155743859609297}592{"step": 3200, "records_seen": 102400, "elapsed_seconds": 18533.608822107315, "gradient_norm": 1.9306800365447998, "learning_rate": 2.6893014782538773e-06, "head_lr": 1.3446507391269384e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6095454092100923, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6095454092100923, "supervised_tokens": 128.65625, "waves": 1.0, "detection_loss": 0.7395066175805894, "grounding_loss": 0.4196021046685963, "detection_grounding_macro_loss": 0.5795543611245929}593{"step": 3210, "records_seen": 102720, "elapsed_seconds": 18585.54665160179, "gradient_norm": 1.5192943811416626, "learning_rate": 2.6589141989294735e-06, "head_lr": 1.3294570994647367e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4104621129499719, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4104621129499719, "supervised_tokens": 115.5625, "waves": 1.0, "detection_loss": 0.4907435743209337, "grounding_loss": 0.3072430911873068, "detection_grounding_macro_loss": 0.39899333275412024}594{"step": 3220, "records_seen": 103040, "elapsed_seconds": 18641.398361682892, "gradient_norm": 1.7513552904129028, "learning_rate": 2.6291864460248154e-06, "head_lr": 1.3145932230124077e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.49353042342605846, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49353042342605846, "supervised_tokens": 78.0, "waves": 1.0, "detection_loss": 0.8270757514983416, "grounding_loss": 0.2934032265826886, "detection_grounding_macro_loss": 0.5602394890405151}595{"step": 3230, "records_seen": 103360, "elapsed_seconds": 18692.98578763008, "gradient_norm": 2.046354293823242, "learning_rate": 2.600120570101798e-06, "head_lr": 1.3000602850508988e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6588436123274732, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6588436123274732, "supervised_tokens": 107.71875, "waves": 1.0, "detection_loss": 0.8778565636123804, "grounding_loss": 0.4106289342045784, "detection_grounding_macro_loss": 0.6442427489084794}596{"step": 3240, "records_seen": 103680, "elapsed_seconds": 18744.006551980972, "gradient_norm": 1.885042428970337, "learning_rate": 2.5717188693879598e-06, "head_lr": 1.2858594346939796e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5952248952817172, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5952248952817172, "supervised_tokens": 104.78125, "waves": 1.0, "detection_loss": 0.7799748569726944, "grounding_loss": 0.43221022320144314, "detection_grounding_macro_loss": 0.6060925400870688}597{"step": 3250, "records_seen": 104000, "elapsed_seconds": 18798.389640569687, "gradient_norm": 1.5237905979156494, "learning_rate": 2.54398358959476e-06, "head_lr": 1.2719917947973798e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.38569981791999197, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.38569981791999197, "supervised_tokens": 124.875, "waves": 1.0, "detection_loss": 0.4184219310991466, "grounding_loss": 0.34861475631695005, "detection_grounding_macro_loss": 0.3835183437080483}598{"step": 3260, "records_seen": 104320, "elapsed_seconds": 18853.631741046906, "gradient_norm": 1.9147000312805176, "learning_rate": 2.516916923740019e-06, "head_lr": 1.2584584618700092e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5432604161847223, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5432604161847223, "supervised_tokens": 90.78125, "waves": 1.0, "detection_loss": 0.8022597841071811, "grounding_loss": 0.4075940806062915, "detection_grounding_macro_loss": 0.6049269323567363}599{"step": 3270, "records_seen": 104640, "elapsed_seconds": 18906.192849874496, "gradient_norm": 1.5776492357254028, "learning_rate": 2.4905210119745087e-06, "head_lr": 1.245260505987254e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5033908001023519, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5033908001023519, "supervised_tokens": 126.59375, "waves": 1.0, "detection_loss": 0.6772267540355268, "grounding_loss": 0.3063767189780871, "detection_grounding_macro_loss": 0.4918017365068069}600{"step": 3280, "records_seen": 104960, "elapsed_seconds": 18958.054188489914, "gradient_norm": 1.7199597358703613, "learning_rate": 2.464797941412736e-06, "head_lr": 1.2323989707063678e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.45630408215674834, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.45630408215674834, "supervised_tokens": 129.59375, "waves": 1.0, "detection_loss": 0.5072111557049485, "grounding_loss": 0.3986093988021215, "detection_grounding_macro_loss": 0.45291027725353494}601{"step": 3290, "records_seen": 105280, "elapsed_seconds": 19011.69212245941, "gradient_norm": 1.8505523204803467, "learning_rate": 2.439749745967918e-06, "head_lr": 1.219874872983959e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5376403766127282, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5376403766127282, "supervised_tokens": 74.96875, "waves": 1.0, "detection_loss": 0.7825755302877523, "grounding_loss": 0.2600472024477009, "detection_grounding_macro_loss": 0.5213113663677266}602{"step": 3300, "records_seen": 105600, "elapsed_seconds": 19060.179696559906, "gradient_norm": 2.0207998752593994, "learning_rate": 2.4153784061911516e-06, "head_lr": 1.2076892030955758e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.729274633033242, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.729274633033242, "supervised_tokens": 145.5625, "waves": 1.0, "detection_loss": 0.9418611019779443, "grounding_loss": 0.32342773777517403, "detection_grounding_macro_loss": 0.6326444198765592}603/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead604  warnings.warn(605/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead606  warnings.warn(607/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead608  warnings.warn(609/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead610  warnings.warn(611/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead612  warnings.warn(613/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead614  warnings.warn(615/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead616  warnings.warn(617/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead618  warnings.warn(619{"validation_step": 3300, "loss": 1.1239397874671178, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1239397874671178, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6358873465411692, "grounding_loss": 0.5432449847459793, "detection_grounding_macro_loss": 1.0895661656435742}620{"step": 3310, "records_seen": 105920, "elapsed_seconds": 19255.91290283203, "gradient_norm": 1.7107354402542114, "learning_rate": 2.3916858491148248e-06, "head_lr": 1.1958429245574123e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6337677401024848, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6337677401024848, "supervised_tokens": 119.125, "waves": 1.0, "detection_loss": 0.8072215651995257, "grounding_loss": 0.38025830342219424, "detection_grounding_macro_loss": 0.59373993431086}621{"step": 3320, "records_seen": 106240, "elapsed_seconds": 19307.620282888412, "gradient_norm": 1.9768171310424805, "learning_rate": 2.3686739481002364e-06, "head_lr": 1.184336974050118e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.7086828899581548, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.7086828899581548, "supervised_tokens": 87.03125, "waves": 1.0, "detection_loss": 1.113268884443267, "grounding_loss": 0.2501520962083608, "detection_grounding_macro_loss": 0.6817104903258139}622{"step": 3330, "records_seen": 106560, "elapsed_seconds": 19359.08557486534, "gradient_norm": 1.909664273262024, "learning_rate": 2.346344522689476e-06, "head_lr": 1.1731722613447379e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.668707057113501, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.668707057113501, "supervised_tokens": 120.65625, "waves": 1.0, "detection_loss": 0.8966110707302792, "grounding_loss": 0.3356165756735943, "detection_grounding_macro_loss": 0.6161138232019367}623{"step": 3340, "records_seen": 106880, "elapsed_seconds": 19411.40511918068, "gradient_norm": 1.980181097984314, "learning_rate": 2.3246993384615532e-06, "head_lr": 1.1623496692307765e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6343119775933701, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6343119775933701, "supervised_tokens": 82.71875, "waves": 1.0, "detection_loss": 0.8931705660535953, "grounding_loss": 0.20288099682632796, "detection_grounding_macro_loss": 0.5480257814399616}624{"step": 3350, "records_seen": 107200, "elapsed_seconds": 19464.018806934357, "gradient_norm": 1.7118752002716064, "learning_rate": 2.30374010689279e-06, "head_lr": 1.151870053446395e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6095512823667377, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6095512823667377, "supervised_tokens": 128.75, "waves": 1.0, "detection_loss": 0.7853666203362601, "grounding_loss": 0.47280601950155365, "detection_grounding_macro_loss": 0.629086319918907}625{"step": 3360, "records_seen": 107520, "elapsed_seconds": 19517.844104528427, "gradient_norm": 1.4798529148101807, "learning_rate": 2.2834684852214993e-06, "head_lr": 1.1417342426107495e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5232820992614506, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5232820992614506, "supervised_tokens": 131.3125, "waves": 1.0, "detection_loss": 0.6484972741136084, "grounding_loss": 0.28423494727096776, "detection_grounding_macro_loss": 0.46636611069228806}626{"step": 3370, "records_seen": 107840, "elapsed_seconds": 19574.330279111862, "gradient_norm": 1.7279242277145386, "learning_rate": 2.263886076316946e-06, "head_lr": 1.1319430381584729e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5906550585059449, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5906550585059449, "supervised_tokens": 129.96875, "waves": 1.0, "detection_loss": 0.7890535510248609, "grounding_loss": 0.3355712824101959, "detection_grounding_macro_loss": 0.5623124167175284}627{"step": 3380, "records_seen": 108160, "elapsed_seconds": 19626.181078910828, "gradient_norm": 1.6714497804641724, "learning_rate": 2.244994428552607e-06, "head_lr": 1.1224972142763034e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5323436377889266, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5323436377889266, "supervised_tokens": 83.84375, "waves": 1.0, "detection_loss": 0.7650133188872132, "grounding_loss": 0.2996739566906399, "detection_grounding_macro_loss": 0.5323436377889266}628{"step": 3390, "records_seen": 108480, "elapsed_seconds": 19677.300104141235, "gradient_norm": 1.6518473625183105, "learning_rate": 2.226795035683746e-06, "head_lr": 1.113397517841873e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.42659490113814513, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.42659490113814513, "supervised_tokens": 108.65625, "waves": 1.0, "detection_loss": 0.49909346064511434, "grounding_loss": 0.3626255839261135, "detection_grounding_macro_loss": 0.4308595222856139}629{"step": 3400, "records_seen": 108800, "elapsed_seconds": 19730.62628865242, "gradient_norm": 1.6233214139938354, "learning_rate": 2.2092893367293013e-06, "head_lr": 1.1046446683646505e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.49010136452746167, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.49010136452746167, "supervised_tokens": 111.375, "waves": 1.0, "detection_loss": 0.6030061985883448, "grounding_loss": 0.3449380064491834, "detection_grounding_macro_loss": 0.4739721025187641}630{"step": 3410, "records_seen": 109120, "elapsed_seconds": 19782.00643014908, "gradient_norm": 1.7299237251281738, "learning_rate": 2.1924787158580983e-06, "head_lr": 1.0962393579290491e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5704136450767692, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5704136450767692, "supervised_tokens": 105.21875, "waves": 1.0, "detection_loss": 0.5918708205738638, "grounding_loss": 0.5537247308012512, "detection_grounding_macro_loss": 0.5727977756875575}631{"step": 3420, "records_seen": 109440, "elapsed_seconds": 19833.822680711746, "gradient_norm": 1.7232317924499512, "learning_rate": 2.176364502279412e-06, "head_lr": 1.0881822511397058e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.536650685177392, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.536650685177392, "supervised_tokens": 102.53125, "waves": 1.0, "detection_loss": 0.6876722272761125, "grounding_loss": 0.36549293746550876, "detection_grounding_macro_loss": 0.5265825823708107}632{"step": 3430, "records_seen": 109760, "elapsed_seconds": 19884.101554632187, "gradient_norm": 2.175614356994629, "learning_rate": 2.16094797013786e-06, "head_lr": 1.08047398506893e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6230626555383765, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6230626555383765, "supervised_tokens": 93.34375, "waves": 1.0, "detection_loss": 0.8827523082494736, "grounding_loss": 0.39392472667564377, "detection_grounding_macro_loss": 0.6383385174625587}633{"step": 3440, "records_seen": 110080, "elapsed_seconds": 19938.506802797318, "gradient_norm": 1.7987736463546753, "learning_rate": 2.146230338412661e-06, "head_lr": 1.0731151692063304e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4995030130452278, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4995030130452278, "supervised_tokens": 101.9375, "waves": 1.0, "detection_loss": 0.5522096551548679, "grounding_loss": 0.39888124174500594, "detection_grounding_macro_loss": 0.4755454484499369}634{"step": 3450, "records_seen": 110400, "elapsed_seconds": 19991.94386458397, "gradient_norm": 1.8226670026779175, "learning_rate": 2.1322127708212485e-06, "head_lr": 1.0661063854106242e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6094232503364765, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6094232503364765, "supervised_tokens": 92.5, "waves": 1.0, "detection_loss": 0.8781431232960636, "grounding_loss": 0.2639262708170073, "detection_grounding_macro_loss": 0.5710346970565354}635{"step": 3460, "records_seen": 110720, "elapsed_seconds": 20043.85125398636, "gradient_norm": 1.6039738655090332, "learning_rate": 2.118896375727257e-06, "head_lr": 1.0594481878636282e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5586600164095898, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5586600164095898, "supervised_tokens": 100.6875, "waves": 1.0, "detection_loss": 0.8361287256991545, "grounding_loss": 0.28119130712002516, "detection_grounding_macro_loss": 0.5586600164095898}636{"step": 3470, "records_seen": 111040, "elapsed_seconds": 20094.115243434906, "gradient_norm": 1.6158809661865234, "learning_rate": 2.1062822060528804e-06, "head_lr": 1.05314110302644e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.39672201209904756, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.39672201209904756, "supervised_tokens": 107.3125, "waves": 1.0, "detection_loss": 0.5094274294747533, "grounding_loss": 0.2689892057399144, "detection_grounding_macro_loss": 0.3892083176073339}637{"step": 3480, "records_seen": 111360, "elapsed_seconds": 20141.06724333763, "gradient_norm": 1.828302264213562, "learning_rate": 2.0943712591956246e-06, "head_lr": 1.0471856295978122e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6428891008254141, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6428891008254141, "supervised_tokens": 143.75, "waves": 1.0, "detection_loss": 0.8067843860222234, "grounding_loss": 0.43216659128665924, "detection_grounding_macro_loss": 0.6194754886544414}638{"step": 3490, "records_seen": 111680, "elapsed_seconds": 20193.061103105545, "gradient_norm": 1.8974746465682983, "learning_rate": 2.0831644769494385e-06, "head_lr": 1.0415822384747191e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5299418166396208, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5299418166396208, "supervised_tokens": 88.6875, "waves": 1.0, "detection_loss": 0.7631496742367745, "grounding_loss": 0.42393824500455096, "detection_grounding_macro_loss": 0.5935439596206628}639{"step": 3500, "records_seen": 112000, "elapsed_seconds": 20245.486268520355, "gradient_norm": 2.034576177597046, "learning_rate": 2.07266274543025e-06, "head_lr": 1.0363313727151247e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5704302028752863, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5704302028752863, "supervised_tokens": 113.25, "waves": 1.0, "detection_loss": 0.7095059797167778, "grounding_loss": 0.33863724147280055, "detection_grounding_macro_loss": 0.5240716105947891}640{"step": 3510, "records_seen": 112320, "elapsed_seconds": 20293.81389284134, "gradient_norm": 1.565675973892212, "learning_rate": 2.0628668950058946e-06, "head_lr": 1.0314334475029472e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.39303751557599753, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.39303751557599753, "supervised_tokens": 124.125, "waves": 1.0, "detection_loss": 0.4596038459635833, "grounding_loss": 0.31759567447006704, "detection_grounding_macro_loss": 0.3885997602168252}641{"step": 3520, "records_seen": 112640, "elapsed_seconds": 20347.403079271317, "gradient_norm": 1.6491392850875854, "learning_rate": 2.05377770023047e-06, "head_lr": 1.0268888501152349e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.47171818824385525, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.47171818824385525, "supervised_tokens": 118.59375, "waves": 1.0, "detection_loss": 0.5726076374865241, "grounding_loss": 0.34200318207470964, "detection_grounding_macro_loss": 0.45730540978061684}642{"step": 3530, "records_seen": 112960, "elapsed_seconds": 20398.919103384018, "gradient_norm": 1.4883389472961426, "learning_rate": 2.0453958797830814e-06, "head_lr": 1.0226979398915406e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5097817098229598, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5097817098229598, "supervised_tokens": 153.625, "waves": 1.0, "detection_loss": 0.5672722880974106, "grounding_loss": 0.4139640793655417, "detection_grounding_macro_loss": 0.49061818373147614}643{"step": 3540, "records_seen": 113280, "elapsed_seconds": 20450.69775867462, "gradient_norm": 1.8299418687820435, "learning_rate": 2.0377220964110203e-06, "head_lr": 1.01886104820551e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5529785877436808, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5529785877436808, "supervised_tokens": 85.1875, "waves": 1.0, "detection_loss": 0.8027325399962137, "grounding_loss": 0.30322463549114786, "detection_grounding_macro_loss": 0.5529785877436808}644{"step": 3550, "records_seen": 113600, "elapsed_seconds": 20505.158492803574, "gradient_norm": 1.7653274536132812, "learning_rate": 2.030756956877362e-06, "head_lr": 1.015378478438681e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5551904048770666, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5551904048770666, "supervised_tokens": 142.6875, "waves": 1.0, "detection_loss": 0.7185337858067619, "grounding_loss": 0.345177486538887, "detection_grounding_macro_loss": 0.5318556361728244}645{"step": 3560, "records_seen": 113920, "elapsed_seconds": 20556.745532512665, "gradient_norm": 1.5905953645706177, "learning_rate": 2.0245010119129904e-06, "head_lr": 1.0122505059564951e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.37663519999586015, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.37663519999586015, "supervised_tokens": 77.375, "waves": 1.0, "detection_loss": 0.4335577043802914, "grounding_loss": 0.3376882233117756, "detection_grounding_macro_loss": 0.3856229638460335}646{"step": 3570, "records_seen": 114240, "elapsed_seconds": 20609.491645097733, "gradient_norm": 1.7216311693191528, "learning_rate": 2.018954756173046e-06, "head_lr": 1.009477378086523e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5201071584115837, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5201071584115837, "supervised_tokens": 85.59375, "waves": 1.0, "detection_loss": 0.7132642464930541, "grounding_loss": 0.32695007033011336, "detection_grounding_macro_loss": 0.5201071584115837}647{"step": 3580, "records_seen": 114560, "elapsed_seconds": 20664.880935430527, "gradient_norm": 1.5496927499771118, "learning_rate": 2.01411862819782e-06, "head_lr": 1.0070593140989099e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.39962933357094244, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.39962933357094244, "supervised_tokens": 92.3125, "waves": 1.0, "detection_loss": 0.6276681498690907, "grounding_loss": 0.22226580978349375, "detection_grounding_macro_loss": 0.4249669798262922}648{"step": 3590, "records_seen": 114880, "elapsed_seconds": 20718.63494157791, "gradient_norm": 1.662767767906189, "learning_rate": 2.0099930103780766e-06, "head_lr": 1.0049965051890381e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5578314181370843, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5578314181370843, "supervised_tokens": 108.9375, "waves": 1.0, "detection_loss": 0.6957214146759725, "grounding_loss": 0.41994142159819603, "detection_grounding_macro_loss": 0.5578314181370843}649{"step": 3600, "records_seen": 115200, "elapsed_seconds": 20770.89947628975, "gradient_norm": 1.8549381494522095, "learning_rate": 2.0065782289248158e-06, "head_lr": 1.0032891144624079e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.6254315024707466, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.6254315024707466, "supervised_tokens": 120.40625, "waves": 1.0, "detection_loss": 0.7232439333901686, "grounding_loss": 0.5145774140954018, "detection_grounding_macro_loss": 0.6189106737427852}650/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead651  warnings.warn(652/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead653  warnings.warn(654/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead655  warnings.warn(656/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead657  warnings.warn(658/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead659  warnings.warn(660/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead661  warnings.warn(662/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead663  warnings.warn(664/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead665  warnings.warn(666{"validation_step": 3600, "loss": 1.1266198598418735, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1266198598418735, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.641883857473507, "grounding_loss": 0.5421632682425636, "detection_grounding_macro_loss": 1.0920235628580353}667{"step": 3610, "records_seen": 115520, "elapsed_seconds": 20962.549855709076, "gradient_norm": 1.80453622341156, "learning_rate": 2.0038745538434838e-06, "head_lr": 1.0019372769217418e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5583856466073485, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5583856466073485, "supervised_tokens": 107.65625, "waves": 1.0, "detection_loss": 0.637576527803586, "grounding_loss": 0.384165707975626, "detection_grounding_macro_loss": 0.5108711178896059}668{"step": 3620, "records_seen": 115840, "elapsed_seconds": 21015.76571249962, "gradient_norm": 2.183336019515991, "learning_rate": 2.0018821989126215e-06, "head_lr": 1.0009410994563107e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5705160690665707, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5705160690665707, "supervised_tokens": 75.28125, "waves": 1.0, "detection_loss": 0.7841489056835209, "grounding_loss": 0.4043571961422761, "detection_grounding_macro_loss": 0.5942530509128985}669{"step": 3630, "records_seen": 116160, "elapsed_seconds": 21069.74217224121, "gradient_norm": 1.8390766382217407, "learning_rate": 2.000601321666959e-06, "head_lr": 1.0003006608334794e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.5249945758478134, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.5249945758478134, "supervised_tokens": 93.0625, "waves": 1.0, "detection_loss": 0.8043695738933121, "grounding_loss": 0.3338432613956301, "detection_grounding_macro_loss": 0.5691064176444711}670{"step": 3640, "records_seen": 116480, "elapsed_seconds": 21124.5970454216, "gradient_norm": 1.7119742631912231, "learning_rate": 2.000032023384964e-06, "head_lr": 1.0000160116924819e-07, "peak_gpu_gib": 19.782602787017822, "loss": 0.4549675283487886, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.4549675283487886, "supervised_tokens": 167.9375, "waves": 1.0, "detection_loss": 0.4078671551413006, "grounding_loss": 0.5155251510441303, "detection_grounding_macro_loss": 0.4616961530927155}671/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead672  warnings.warn(673/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead674  warnings.warn(675/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead676  warnings.warn(677/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead678  warnings.warn(679/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead680  warnings.warn(681/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead682  warnings.warn(683/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead684  warnings.warn(685/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead686  warnings.warn(687{"validation_step": 3642, "loss": 1.1256048969833603, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1256048969833603, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6414700296196356, "grounding_loss": 0.5404664465359279, "detection_grounding_macro_loss": 1.0909682380777816}688wandb: updating run metadata689wandb: uploading output.log; uploading wandb-summary.json; uploading config.yaml690wandb: 691wandb: Run history:692wandb:                       optimizer_step โ–โ–โ–โ–โ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–„โ–„โ–„โ–„โ–…โ–…โ–…โ–…โ–†โ–†โ–†โ–†โ–†โ–‡โ–‡โ–‡โ–‡โ–‡โ–ˆโ–ˆ693wandb: train/detection_grounding_macro_loss โ–ˆโ–‡โ–ƒโ–„โ–ƒโ–…โ–…โ–ƒโ–„โ–†โ–‚โ–ƒโ–„โ–‚โ–…โ–…โ–ƒโ–ƒโ–‚โ–ƒโ–‚โ–ƒโ–‚โ–‚โ–„โ–„โ–„โ–ƒโ–„โ–…โ–‚โ–„โ–ƒโ–ƒโ–‚โ–โ–„โ–โ–ƒโ–ƒ694wandb:                 train/detection_loss โ–‚โ–„โ–…โ–„โ–…โ–„โ–†โ–ƒโ–‚โ–‚โ–‚โ–„โ–ƒโ–ƒโ–„โ–†โ–‚โ–‚โ–ƒโ–…โ–†โ–ƒโ–„โ–ƒโ–ƒโ–โ–„โ–ƒโ–ƒโ–„โ–†โ–ƒโ–ˆโ–†โ–…โ–†โ–„โ–…โ–…โ–ƒ695wandb:                    train/direct_loss โ–ˆโ–…โ–…โ–†โ–„โ–‚โ–†โ–‚โ–…โ–…โ–‚โ–…โ–„โ–…โ–…โ–ƒโ–†โ–„โ–‚โ–‚โ–‡โ–…โ–โ–‡โ–…โ–†โ–‚โ–†โ–…โ–„โ–„โ–โ–„โ–†โ–„โ–…โ–‚โ–„โ–‚โ–ƒ696wandb:                train/elapsed_seconds โ–โ–โ–โ–‚โ–‚โ–‚โ–‚โ–‚โ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–„โ–„โ–„โ–„โ–„โ–„โ–„โ–„โ–„โ–„โ–…โ–†โ–†โ–†โ–†โ–‡โ–‡โ–‡โ–‡โ–‡โ–‡โ–‡โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ697wandb:                  train/gradient_norm โ–„โ–„โ–ˆโ–โ–โ–‚โ–โ–‚โ–„โ–‚โ–„โ–‚โ–ƒโ–‚โ–ƒโ–‚โ–‚โ–‚โ–„โ–โ–‚โ–ƒโ–„โ–„โ–‚โ–ƒโ–‚โ–ƒโ–„โ–‚โ–ƒโ–โ–ƒโ–„โ–ƒโ–„โ–ƒโ–‚โ–ƒโ–‚698wandb:                 train/grounding_loss โ–ˆโ–„โ–ƒโ–†โ–ƒโ–„โ–ƒโ–ƒโ–ƒโ–‚โ–„โ–‚โ–ƒโ–‚โ–‚โ–‚โ–ƒโ–ƒโ–‚โ–‚โ–‚โ–โ–ƒโ–ƒโ–ƒโ–‚โ–ƒโ–ƒโ–โ–‚โ–‚โ–„โ–ƒโ–„โ–ƒโ–โ–ƒโ–‚โ–‚โ–ƒ699wandb:                        train/head_lr โ–‚โ–‚โ–ˆโ–ˆโ–ˆโ–ˆโ–‡โ–‡โ–‡โ–‡โ–‡โ–‡โ–†โ–†โ–†โ–…โ–…โ–…โ–…โ–…โ–…โ–…โ–„โ–„โ–„โ–ƒโ–ƒโ–ƒโ–ƒโ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–โ–โ–โ–โ–700wandb:                  train/learning_rate โ–†โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‡โ–‡โ–‡โ–‡โ–‡โ–‡โ–†โ–…โ–…โ–…โ–…โ–…โ–…โ–„โ–„โ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–‚โ–‚โ–‚โ–โ–โ–โ–โ–โ–โ–701wandb:                           train/loss โ–ˆโ–ƒโ–ƒโ–‚โ–„โ–…โ–‚โ–ƒโ–…โ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–โ–ƒโ–„โ–„โ–‚โ–„โ–‚โ–‚โ–ƒโ–ƒโ–‚โ–‚โ–‚โ–ƒโ–‚โ–„โ–ƒโ–ƒโ–„โ–‚โ–‚โ–‚โ–‚โ–ƒโ–ƒโ–„702wandb:                                  +16 ...703wandb: 704wandb: Run summary:705wandb:                       optimizer_step 3642706wandb: train/detection_grounding_macro_loss 0.6226707wandb:                 train/detection_loss 0.6226708wandb:                    train/direct_loss 0.6226709wandb:                train/elapsed_seconds 21130.91437710wandb:                  train/gradient_norm 7.9568711wandb:                 train/grounding_loss 0.33361712wandb:                        train/head_lr 0.0713wandb:                  train/learning_rate 0.0714wandb:                           train/loss 0.6226715wandb:                                  +16 ...716wandb: 717wandb: ๐Ÿš€ View run standard_lora_4gpu_20260908T185018Z_lJdQkO_stage1 at: https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale/runs/4e0c9a3a72c04a7b718wandb: โญ๏ธ View project at: https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale719wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)720wandb: Find logs at: /dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/runs/standard_lora_4gpu_20260908T185018Z_lJdQkO/stage1/wandb/run-20260908_145233-4e0c9a3a72c04a7b/logs721/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead722  warnings.warn(723/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead724  warnings.warn(725/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead726  warnings.warn(727/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead728  warnings.warn(729Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.730  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes731  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).732  - If you are not the owner of the model architecture class, please contact the model code owner to update it.733Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.734  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes735  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).736  - If you are not the owner of the model architecture class, please contact the model code owner to update it.737Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.738  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes739  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).740  - If you are not the owner of the model architecture class, please contact the model code owner to update it.741Qwen2ForCausalLM has generative capabilities, as `prepare_inputs_for_generation` is explicitly defined. However, it doesn't directly inherit from `GenerationMixin`. From ๐Ÿ‘‰v4.50๐Ÿ‘ˆ onwards, `PreTrainedModel` will NOT inherit from `GenerationMixin`, and this model will lose the ability to call `generate` and other related functions.742  - If you're using `trust_remote_code=True`, you can get rid of this warning by loading the model with an auto class. See https://huggingface.co/docs/transformers/en/model_doc/auto#auto-classes743  - If you are the owner of the model architecture code, please modify your model class such that it inherits from `GenerationMixin` (after `PreTrainedModel`, otherwise you'll get an exception).744  - If you are not the owner of the model architecture class, please contact the model code owner to update it.745
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:   0%|          | 0/2 [00:00<?, ?it/s]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:06<00:06,  6.23s/it]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:06<00:06,  6.21s/it]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:06<00:06,  6.19s/it]
Loading checkpoint shards:  50%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ     | 1/2 [00:06<00:06,  6.21s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.59s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.58s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.58s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.57s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.84s/it]746
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.83s/it]
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.81s/it]747 748
Loading checkpoint shards: 100%|โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ| 2/2 [00:09<00:00,  4.83s/it]749trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.9022750trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.9022751trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.9022752trainable params: 59,867,136 || all params: 3,147,331,584 || trainable%: 1.9022753{754  "schema": "refwave-simple-lora-head-v1",755  "method": "standard",756  "base_model": "/dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/LocateAnything-3B",757  "stage": "stage2",758  "train_sha256": "ee4f3df5196a5f3f7b49c372fa8915da332df32d67024f876d0f0727422193e4",759  "validation_sha256": "a18e9ff2975574c64d89717901876a4909ad60c892d471cd52f2b4937920bf72",760  "records": 39918,761  "validation_records": 747,762  "epochs": 1,763  "steps": 1248,764  "world_size": 4,765  "batch_size": 8,766  "accumulation": 1,767  "global_batch": 32,768  "batch_tokens": 16384,769  "max_sequence": 32768,770  "lora_rank": 32,771  "learning_rate": 1e-05,772  "head_lr": 5e-07,773  "warmup_ratio": 0.03,774  "seed": 20260908,775  "image_token_limit": 6000,776  "trainable": {777    "rank": 32,778    "tensors": 504,779    "parameters": 59867136,780    "head_parameters": 312690688,781    "trainable_parameters": 372557824,782    "scope": "Qwen LoRA + complete LM head; frozen input embeddings, vision and projector"783  },784  "synchronized_trainable_initialization": true,785  "init_checkpoint_sha256": "b60aa3980fb4758af923f924e0ea3750d4c56c945fae2f2595d6fde6508033cd",786  "tuning": "lora",787  "vision_lr": 1e-06,788  "projector_lr": 1e-05,789  "optimizer_sharding": false790}791W&B tracking mode: online792/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead793  warnings.warn(794/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead795  warnings.warn(796/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead797  warnings.warn(798wandb: [wandb.login()] Loaded credentials for https://api.wandb.ai from /u/mfiaz/.netrc.799/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead800  warnings.warn(801/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead802  warnings.warn(803/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead804  warnings.warn(805wandb: Currently logged in as: shubhamrpatle (shubhamrpatle-mbzuai) to https://api.wandb.ai. Use `wandb login --relogin` to force relogin806wandb: setting up run fa7d4d62be7f59b5807wandb: Tracking run with wandb version 0.27.2808wandb: Run data is saved locally in /dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/runs/standard_lora_4gpu_20260908T185018Z_lJdQkO/stage2/wandb/run-20260908_204844-fa7d4d62be7f59b5809wandb: Run `wandb offline` to turn off syncing.810wandb: Syncing run standard_lora_4gpu_20260908T185018Z_lJdQkO_stage2811wandb: โญ๏ธ View project at https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale812wandb: ๐Ÿš€ View run at https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale/runs/fa7d4d62be7f59b5813/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead814  warnings.warn(815/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead816  warnings.warn(817{"step": 1, "records_seen": 32, "elapsed_seconds": 8.776795864105225, "gradient_norm": 1.6377675533294678, "learning_rate": 2.702702702702703e-07, "head_lr": 1.3513513513513514e-08, "peak_gpu_gib": 16.011971950531006, "loss": 1.426052556336714, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.426052556336714, "supervised_tokens": 507.0, "waves": 1.0, "detection_loss": 2.145985822928579, "grounding_loss": 0.37384239747167997, "detection_grounding_macro_loss": 1.2599141102001294}818{"step": 10, "records_seen": 320, "elapsed_seconds": 81.61682486534119, "gradient_norm": 1.4349390268325806, "learning_rate": 2.432432432432433e-06, "head_lr": 1.2162162162162163e-07, "peak_gpu_gib": 19.62620210647583, "loss": 1.2001647937577218, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2001647937577218, "supervised_tokens": 491.96875, "waves": 1.0, "detection_loss": 1.9775481224060059, "grounding_loss": 0.3191303546229998, "detection_grounding_macro_loss": 1.1483392385145028}819{"step": 20, "records_seen": 640, "elapsed_seconds": 179.28230047225952, "gradient_norm": 1.1348942518234253, "learning_rate": 5.135135135135135e-06, "head_lr": 2.567567567567567e-07, "peak_gpu_gib": 19.733869552612305, "loss": 1.3633187480736524, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3633187480736524, "supervised_tokens": 774.34375, "waves": 1.0, "detection_loss": 2.0062123196465627, "grounding_loss": 0.13597647507082333, "detection_grounding_macro_loss": 1.071094397358693}820{"step": 30, "records_seen": 960, "elapsed_seconds": 269.79269886016846, "gradient_norm": 1.3350390195846558, "learning_rate": 7.837837837837838e-06, "head_lr": 3.918918918918919e-07, "peak_gpu_gib": 19.73678731918335, "loss": 1.1833987799721513, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1833987799721513, "supervised_tokens": 1010.65625, "waves": 1.0, "detection_loss": 1.9073265029324427, "grounding_loss": 0.252634564737491, "detection_grounding_macro_loss": 1.0799805338349668}821{"step": 40, "records_seen": 1280, "elapsed_seconds": 359.2742326259613, "gradient_norm": 1.4687213897705078, "learning_rate": 9.999939430638673e-06, "head_lr": 4.999969715319335e-07, "peak_gpu_gib": 21.530005931854248, "loss": 1.0624786177998402, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0624786177998402, "supervised_tokens": 527.53125, "waves": 1.0, "detection_loss": 1.8848867639899254, "grounding_loss": 0.24007047160975503, "detection_grounding_macro_loss": 1.0624786177998402}822{"step": 50, "records_seen": 1600, "elapsed_seconds": 447.79870533943176, "gradient_norm": 1.271756649017334, "learning_rate": 9.997819674190822e-06, "head_lr": 4.998909837095411e-07, "peak_gpu_gib": 21.530005931854248, "loss": 1.2368733312468976, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2368733312468976, "supervised_tokens": 1129.75, "waves": 1.0, "detection_loss": 2.068110225831761, "grounding_loss": 0.2948048507173856, "detection_grounding_macro_loss": 1.1814575382745733}823{"step": 60, "records_seen": 1920, "elapsed_seconds": 548.2826659679413, "gradient_norm": 1.752354383468628, "learning_rate": 9.992673079989292e-06, "head_lr": 4.996336539994645e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.5449427161365747, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5449427161365747, "supervised_tokens": 780.03125, "waves": 1.0, "detection_loss": 2.044415587470645, "grounding_loss": 0.5914035981351679, "detection_grounding_macro_loss": 1.3179095928029063}824{"step": 70, "records_seen": 2240, "elapsed_seconds": 638.6924557685852, "gradient_norm": 1.5058701038360596, "learning_rate": 9.984503111468979e-06, "head_lr": 4.992251555734489e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.1127863926730583, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1127863926730583, "supervised_tokens": 425.9375, "waves": 1.0, "detection_loss": 1.9431533575057984, "grounding_loss": 0.3801096589971113, "detection_grounding_macro_loss": 1.1616315082514548}825{"step": 80, "records_seen": 2560, "elapsed_seconds": 745.8089981079102, "gradient_norm": 1.4876848459243774, "learning_rate": 9.973315266664702e-06, "head_lr": 4.98665763333235e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.34911097609438, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.34911097609438, "supervised_tokens": 381.15625, "waves": 1.0, "detection_loss": 2.113875006635984, "grounding_loss": 0.36584293682660374, "detection_grounding_macro_loss": 1.2398589717312938}826{"step": 90, "records_seen": 2880, "elapsed_seconds": 830.1605360507965, "gradient_norm": 1.4813661575317383, "learning_rate": 9.959117074511251e-06, "head_lr": 4.979558537255626e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.1797288867668883, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1797288867668883, "supervised_tokens": 392.15625, "waves": 1.0, "detection_loss": 1.9048410095274448, "grounding_loss": 0.45461676400633166, "detection_grounding_macro_loss": 1.1797288867668883}827{"step": 100, "records_seen": 3200, "elapsed_seconds": 924.15829205513, "gradient_norm": 1.4628880023956299, "learning_rate": 9.94191808977675e-06, "head_lr": 4.970959044888375e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.1493996432858182, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1493996432858182, "supervised_tokens": 584.9375, "waves": 1.0, "detection_loss": 1.8887338563799858, "grounding_loss": 0.41006543019165065, "detection_grounding_macro_loss": 1.1493996432858182}828{"step": 110, "records_seen": 3520, "elapsed_seconds": 1010.468700170517, "gradient_norm": 1.251939058303833, "learning_rate": 9.921729886632708e-06, "head_lr": 4.960864943316353e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.5278332176967524, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5278332176967524, "supervised_tokens": 892.1875, "waves": 1.0, "detection_loss": 2.0093598392876713, "grounding_loss": 0.46847465019673107, "detection_grounding_macro_loss": 1.2389172447422012}829{"step": 120, "records_seen": 3840, "elapsed_seconds": 1122.2028198242188, "gradient_norm": 1.5466127395629883, "learning_rate": 9.898566050865097e-06, "head_lr": 4.949283025432548e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.3713023082157463, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3713023082157463, "supervised_tokens": 421.0625, "waves": 1.0, "detection_loss": 2.0239520982692114, "grounding_loss": 0.41742953813760464, "detection_grounding_macro_loss": 1.220690818203408}830{"step": 130, "records_seen": 4160, "elapsed_seconds": 1218.6816132068634, "gradient_norm": 1.5758085250854492, "learning_rate": 9.872442170731724e-06, "head_lr": 4.936221085365861e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.4400132128503174, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4400132128503174, "supervised_tokens": 458.625, "waves": 1.0, "detection_loss": 1.9048321609911711, "grounding_loss": 0.2521425676014688, "detection_grounding_macro_loss": 1.07848736429632}831{"step": 140, "records_seen": 4480, "elapsed_seconds": 1303.7305736541748, "gradient_norm": 1.4393653869628906, "learning_rate": 9.843375826471984e-06, "head_lr": 4.921687913235991e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.2237247975757555, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2237247975757555, "supervised_tokens": 273.15625, "waves": 1.0, "detection_loss": 1.9588042961226568, "grounding_loss": 0.27862258515831073, "detection_grounding_macro_loss": 1.1187134406404837}832{"step": 150, "records_seen": 4800, "elapsed_seconds": 1386.4434111118317, "gradient_norm": 1.5345556735992432, "learning_rate": 9.811386578476146e-06, "head_lr": 4.905693289238073e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.2743896746542305, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2743896746542305, "supervised_tokens": 542.03125, "waves": 1.0, "detection_loss": 2.1035098766579345, "grounding_loss": 0.33472011238336563, "detection_grounding_macro_loss": 1.21911499452065}833{"step": 160, "records_seen": 5120, "elapsed_seconds": 1463.648839712143, "gradient_norm": 1.7763111591339111, "learning_rate": 9.776495954122042e-06, "head_lr": 4.888247977061021e-07, "peak_gpu_gib": 22.92335844039917, "loss": 1.1872701509510364, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1872701509510364, "supervised_tokens": 410.3125, "waves": 1.0, "detection_loss": 1.7669048940434176, "grounding_loss": 0.5303507754463378, "detection_grounding_macro_loss": 1.1486278347448777}834{"step": 170, "records_seen": 5440, "elapsed_seconds": 1548.6345446109772, "gradient_norm": 1.5293998718261719, "learning_rate": 9.738727433288077e-06, "head_lr": 4.869363716644038e-07, "peak_gpu_gib": 23.705328464508057, "loss": 0.9962430049250273, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.9962430049250273, "supervised_tokens": 260.90625, "waves": 1.0, "detection_loss": 1.9814665487834386, "grounding_loss": 0.22995802636848517, "detection_grounding_macro_loss": 1.105712287575962}835{"step": 180, "records_seen": 5760, "elapsed_seconds": 1650.6142036914825, "gradient_norm": 1.5993905067443848, "learning_rate": 9.698106432552293e-06, "head_lr": 4.849053216276146e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.4376622859854251, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4376622859854251, "supervised_tokens": 738.875, "waves": 1.0, "detection_loss": 2.139369293263084, "grounding_loss": 0.4120905061180775, "detection_grounding_macro_loss": 1.2757298996905808}836{"step": 190, "records_seen": 6080, "elapsed_seconds": 1737.5261886119843, "gradient_norm": 1.2243117094039917, "learning_rate": 9.654660288088108e-06, "head_lr": 4.827330144044054e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.437821429929656, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.437821429929656, "supervised_tokens": 640.4375, "waves": 1.0, "detection_loss": 2.0626846126147678, "grounding_loss": 0.24490080843989662, "detection_grounding_macro_loss": 1.1537927105273322}837{"step": 200, "records_seen": 6400, "elapsed_seconds": 1823.2934312820435, "gradient_norm": 1.3421677350997925, "learning_rate": 9.608418237268257e-06, "head_lr": 4.804209118634128e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.37226967245806, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.37226967245806, "supervised_tokens": 357.21875, "waves": 1.0, "detection_loss": 1.8345216940278593, "grounding_loss": 0.19095895066857338, "detection_grounding_macro_loss": 1.0127403223482163}838{"step": 210, "records_seen": 6720, "elapsed_seconds": 1927.4718804359436, "gradient_norm": 1.677152156829834, "learning_rate": 9.55941139898931e-06, "head_lr": 4.779705699494655e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.0641218359795062, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0641218359795062, "supervised_tokens": 272.40625, "waves": 1.0, "detection_loss": 1.891301566362381, "grounding_loss": 0.3342573679946166, "detection_grounding_macro_loss": 1.1127794671784987}839{"step": 220, "records_seen": 7040, "elapsed_seconds": 2018.205292224884, "gradient_norm": 1.491274356842041, "learning_rate": 9.507672752730003e-06, "head_lr": 4.7538363763650007e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.2676797599997371, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2676797599997371, "supervised_tokens": 459.09375, "waves": 1.0, "detection_loss": 2.099644923911375, "grounding_loss": 0.3247859075665474, "detection_grounding_macro_loss": 1.2122154157389613}840{"step": 230, "records_seen": 7360, "elapsed_seconds": 2114.3499069213867, "gradient_norm": 1.4830100536346436, "learning_rate": 9.45323711635747e-06, "head_lr": 4.7266185581787347e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.3908473401960748, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3908473401960748, "supervised_tokens": 636.59375, "waves": 1.0, "detection_loss": 2.063971315559588, "grounding_loss": 0.4070507608186325, "detection_grounding_macro_loss": 1.2355110381891101}841{"step": 240, "records_seen": 7680, "elapsed_seconds": 2203.4616186618805, "gradient_norm": 1.535933494567871, "learning_rate": 9.396141122696338e-06, "head_lr": 4.698070561348168e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.441262545704376, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.441262545704376, "supervised_tokens": 622.6875, "waves": 1.0, "detection_loss": 2.0238239680017744, "grounding_loss": 0.3290998304093426, "detection_grounding_macro_loss": 1.1764618992055584}842{"step": 250, "records_seen": 8000, "elapsed_seconds": 2285.1372282505035, "gradient_norm": 1.473137378692627, "learning_rate": 9.336423194876412e-06, "head_lr": 4.668211597438205e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.0510590192861855, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0510590192861855, "supervised_tokens": 307.21875, "waves": 1.0, "detection_loss": 1.721650614457972, "grounding_loss": 0.2910552114248276, "detection_grounding_macro_loss": 1.0063529129413997}843{"step": 260, "records_seen": 8320, "elapsed_seconds": 2381.2536725997925, "gradient_norm": 1.6894150972366333, "learning_rate": 9.274123520475586e-06, "head_lr": 4.6370617602377925e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.1825387155517433, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1825387155517433, "supervised_tokens": 460.3125, "waves": 1.0, "detection_loss": 1.9548186138272285, "grounding_loss": 0.410258817276258, "detection_grounding_macro_loss": 1.1825387155517433}844{"step": 270, "records_seen": 8640, "elapsed_seconds": 2475.497398853302, "gradient_norm": 1.465378761291504, "learning_rate": 9.209284024475335e-06, "head_lr": 4.604642012237667e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.3867531795985997, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3867531795985997, "supervised_tokens": 456.84375, "waves": 1.0, "detection_loss": 2.008191931247711, "grounding_loss": 0.3510219268500805, "detection_grounding_macro_loss": 1.1796069290488957}845{"step": 280, "records_seen": 8960, "elapsed_seconds": 2567.551795721054, "gradient_norm": 1.614486813545227, "learning_rate": 9.141948341047023e-06, "head_lr": 4.5709741705235107e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.1595140015124343, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1595140015124343, "supervised_tokens": 809.5625, "waves": 1.0, "detection_loss": 1.6485017776489257, "grounding_loss": 0.3445343746182819, "detection_grounding_macro_loss": 0.9965180761336038}846{"step": 290, "records_seen": 9280, "elapsed_seconds": 2657.9620406627655, "gradient_norm": 1.4571141004562378, "learning_rate": 9.07216178418799e-06, "head_lr": 4.536080892093994e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.3967424165807643, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3967424165807643, "supervised_tokens": 806.09375, "waves": 1.0, "detection_loss": 1.7843859791755676, "grounding_loss": 0.40609775661626674, "detection_grounding_macro_loss": 1.0952418678959173}847{"step": 300, "records_seen": 9600, "elapsed_seconds": 2744.7017912864685, "gradient_norm": 1.3479982614517212, "learning_rate": 8.999971317227208e-06, "head_lr": 4.499985658613603e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.5414734605764693, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5414734605764693, "supervised_tokens": 676.9375, "waves": 1.0, "detection_loss": 2.077687447721308, "grounding_loss": 0.36180268885782424, "detection_grounding_macro_loss": 1.219745068289566}848/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead849  warnings.warn(850/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead851  warnings.warn(852/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead853  warnings.warn(854/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead855  warnings.warn(856/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead857  warnings.warn(858/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead859  warnings.warn(860/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead861  warnings.warn(862/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead863  warnings.warn(864{"validation_step": 300, "loss": 1.1073472436182101, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1073472436182101, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6017170094589348, "grounding_loss": 0.5465906806503024, "detection_grounding_macro_loss": 1.0741538450546186}865{"step": 310, "records_seen": 9920, "elapsed_seconds": 3015.9679033756256, "gradient_norm": 1.411186695098877, "learning_rate": 8.925425521220983e-06, "head_lr": 4.462712760610491e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.6118691652081907, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.6118691652081907, "supervised_tokens": 1151.15625, "waves": 1.0, "detection_loss": 2.101937656817229, "grounding_loss": 0.3594719088739819, "detection_grounding_macro_loss": 1.2307047828456055}866{"step": 320, "records_seen": 10240, "elapsed_seconds": 3115.08722114563, "gradient_norm": 1.3241944313049316, "learning_rate": 8.848574562260017e-06, "head_lr": 4.4242872811300083e-07, "peak_gpu_gib": 23.705328464508057, "loss": 1.377150432178297, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.377150432178297, "supervised_tokens": 557.40625, "waves": 1.0, "detection_loss": 2.07553358823061, "grounding_loss": 0.21317850542444225, "detection_grounding_macro_loss": 1.1443560468275262}867{"step": 330, "records_seen": 10560, "elapsed_seconds": 3217.3480348587036, "gradient_norm": 1.3399890661239624, "learning_rate": 8.769470157709799e-06, "head_lr": 4.384735078854899e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.4108685250704696, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4108685250704696, "supervised_tokens": 430.96875, "waves": 1.0, "detection_loss": 2.0529203982580277, "grounding_loss": 0.18513313080331334, "detection_grounding_macro_loss": 1.1190267645306706}868{"step": 340, "records_seen": 10880, "elapsed_seconds": 3294.1827025413513, "gradient_norm": 1.3188598155975342, "learning_rate": 8.688165541407053e-06, "head_lr": 4.344082770703526e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.179135123764425, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.179135123764425, "supervised_tokens": 441.40625, "waves": 1.0, "detection_loss": 2.0406361678067375, "grounding_loss": 0.20276727384980403, "detection_grounding_macro_loss": 1.1217017208282707}869{"step": 350, "records_seen": 11200, "elapsed_seconds": 3382.2524523735046, "gradient_norm": 1.591460943222046, "learning_rate": 8.60471542783569e-06, "head_lr": 4.302357713917845e-07, "peak_gpu_gib": 25.681715965270996, "loss": 0.913421563222073, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.913421563222073, "supervised_tokens": 347.03125, "waves": 1.0, "detection_loss": 1.6584907557283128, "grounding_loss": 0.33392330238388646, "detection_grounding_macro_loss": 0.9962070290560996}870{"step": 360, "records_seen": 11520, "elapsed_seconds": 3483.7790760993958, "gradient_norm": 1.5961250066757202, "learning_rate": 8.519175975306314e-06, "head_lr": 4.259587987653156e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.0936023472459055, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0936023472459055, "supervised_tokens": 320.09375, "waves": 1.0, "detection_loss": 1.9718693653742472, "grounding_loss": 0.31866086066207466, "detection_grounding_macro_loss": 1.1452651130181608}871{"step": 370, "records_seen": 11840, "elapsed_seconds": 3563.104251384735, "gradient_norm": 1.5716739892959595, "learning_rate": 8.431604748164116e-06, "head_lr": 4.215802374082058e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.439251821809762, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.439251821809762, "supervised_tokens": 455.5, "waves": 1.0, "detection_loss": 1.8391480666139852, "grounding_loss": 0.41729475175452535, "detection_grounding_macro_loss": 1.1282214091842553}872{"step": 380, "records_seen": 12160, "elapsed_seconds": 3649.918340444565, "gradient_norm": 1.3221817016601562, "learning_rate": 8.342060678050572e-06, "head_lr": 4.171030339025285e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.2243293843930587, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2243293843930587, "supervised_tokens": 543.96875, "waves": 1.0, "detection_loss": 1.9175005382613133, "grounding_loss": 0.21123308258560988, "detection_grounding_macro_loss": 1.0643668104234616}873{"step": 390, "records_seen": 12480, "elapsed_seconds": 3734.4364125728607, "gradient_norm": 1.488150715827942, "learning_rate": 8.250604024244993e-06, "head_lr": 4.1253020121224956e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.3892528631258756, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3892528631258756, "supervised_tokens": 383.15625, "waves": 1.0, "detection_loss": 1.8366555246439846, "grounding_loss": 0.40496700778603556, "detection_grounding_macro_loss": 1.12081126621501}874{"step": 400, "records_seen": 12800, "elapsed_seconds": 3820.4121644496918, "gradient_norm": 1.5421123504638672, "learning_rate": 8.15729633311264e-06, "head_lr": 4.07864816655632e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.2250005500457632, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2250005500457632, "supervised_tokens": 339.0625, "waves": 1.0, "detection_loss": 1.847829619520589, "grounding_loss": 0.31471191004409493, "detection_grounding_macro_loss": 1.0812707647823419}875{"step": 410, "records_seen": 13120, "elapsed_seconds": 3902.1517083644867, "gradient_norm": 1.457060694694519, "learning_rate": 8.062200396686704e-06, "head_lr": 4.031100198343351e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.5337842754088342, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5337842754088342, "supervised_tokens": 734.875, "waves": 1.0, "detection_loss": 2.114221853869302, "grounding_loss": 0.42567617107521405, "detection_grounding_macro_loss": 1.269949012472258}876{"step": 420, "records_seen": 13440, "elapsed_seconds": 3991.35892701149, "gradient_norm": 1.4768739938735962, "learning_rate": 7.965380210411954e-06, "head_lr": 3.982690105205977e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.4472760375695088, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4472760375695088, "supervised_tokens": 914.34375, "waves": 1.0, "detection_loss": 2.1522768877054514, "grounding_loss": 0.41689017967851566, "detection_grounding_macro_loss": 1.2845835336919835}877{"step": 430, "records_seen": 13760, "elapsed_seconds": 4079.3483130931854, "gradient_norm": 1.5779674053192139, "learning_rate": 7.86690093007862e-06, "head_lr": 3.933450465039309e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.3916791973169893, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3916791973169893, "supervised_tokens": 689.625, "waves": 1.0, "detection_loss": 1.9316984642119635, "grounding_loss": 0.3607333241538568, "detection_grounding_macro_loss": 1.1462158941829101}878{"step": 440, "records_seen": 14080, "elapsed_seconds": 4174.601387023926, "gradient_norm": 1.6085224151611328, "learning_rate": 7.766828827975327e-06, "head_lr": 3.8834144139876636e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.4043594374958275, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4043594374958275, "supervised_tokens": 1017.5, "waves": 1.0, "detection_loss": 1.993218857049942, "grounding_loss": 0.42292707157230325, "detection_grounding_macro_loss": 1.2080729643111225}879{"step": 450, "records_seen": 14400, "elapsed_seconds": 4250.5213413238525, "gradient_norm": 1.2468832731246948, "learning_rate": 7.66523124829075e-06, "head_lr": 3.8326156241453745e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.0703957838930478, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0703957838930478, "supervised_tokens": 591.65625, "waves": 1.0, "detection_loss": 1.936434455215931, "grounding_loss": 0.20435711257016465, "detection_grounding_macro_loss": 1.0703957838930478}880{"step": 460, "records_seen": 14720, "elapsed_seconds": 4350.738319158554, "gradient_norm": 1.4906535148620605, "learning_rate": 7.562176561793868e-06, "head_lr": 3.7810882808969335e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.465327101722437, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.465327101722437, "supervised_tokens": 521.875, "waves": 1.0, "detection_loss": 1.9075105786323547, "grounding_loss": 0.33530266073042486, "detection_grounding_macro_loss": 1.1214066196813899}881{"step": 470, "records_seen": 15040, "elapsed_seconds": 4453.121688842773, "gradient_norm": 1.4456909894943237, "learning_rate": 7.457734119823416e-06, "head_lr": 3.7288670599117075e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.4925342563074082, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4925342563074082, "supervised_tokens": 475.40625, "waves": 1.0, "detection_loss": 2.0817240306309293, "grounding_loss": 0.3677174144170501, "detection_grounding_macro_loss": 1.2247207225239896}882{"step": 480, "records_seen": 15360, "elapsed_seconds": 4552.560523271561, "gradient_norm": 1.5509144067764282, "learning_rate": 7.351974207617428e-06, "head_lr": 3.6759871038087134e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.3145948820747435, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3145948820747435, "supervised_tokens": 769.71875, "waves": 1.0, "detection_loss": 1.7590814388933635, "grounding_loss": 0.4660296372391961, "detection_grounding_macro_loss": 1.11255553806628}883{"step": 490, "records_seen": 15680, "elapsed_seconds": 4644.169975042343, "gradient_norm": 1.683659553527832, "learning_rate": 7.244967997014324e-06, "head_lr": 3.6224839985071613e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.3670169153642746, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3670169153642746, "supervised_tokens": 890.34375, "waves": 1.0, "detection_loss": 1.959497830271721, "grounding_loss": 0.37954872385186417, "detection_grounding_macro_loss": 1.1695232770617925}884{"step": 500, "records_seen": 16000, "elapsed_seconds": 4732.919767141342, "gradient_norm": 1.6017661094665527, "learning_rate": 7.1367874985573445e-06, "head_lr": 3.568393749278672e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.2665167401873987, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2665167401873987, "supervised_tokens": 295.90625, "waves": 1.0, "detection_loss": 1.856443750858307, "grounding_loss": 0.283305055735885, "detection_grounding_macro_loss": 1.069874403297096}885{"step": 510, "records_seen": 16320, "elapsed_seconds": 4821.123128890991, "gradient_norm": 1.840946078300476, "learning_rate": 7.027505513034573e-06, "head_lr": 3.513752756517286e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.0258454391732812, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0258454391732812, "supervised_tokens": 237.78125, "waves": 1.0, "detection_loss": 1.814073566879545, "grounding_loss": 0.41277911762396496, "detection_grounding_macro_loss": 1.113426342251755}886{"step": 520, "records_seen": 16640, "elapsed_seconds": 4909.14504981041, "gradient_norm": 1.2810124158859253, "learning_rate": 6.917195582487162e-06, "head_lr": 3.458597791243581e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.6495838111732155, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.6495838111732155, "supervised_tokens": 1052.65625, "waves": 1.0, "detection_loss": 2.0736973683039346, "grounding_loss": 0.37724313978105783, "detection_grounding_macro_loss": 1.2254702540424962}887{"step": 530, "records_seen": 16960, "elapsed_seconds": 4995.191153049469, "gradient_norm": 1.4036279916763306, "learning_rate": 6.805931940718727e-06, "head_lr": 3.4029659703593626e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.4056432874176608, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4056432874176608, "supervised_tokens": 602.625, "waves": 1.0, "detection_loss": 2.0333218827843664, "grounding_loss": 0.3595122951398177, "detection_grounding_macro_loss": 1.196417088962092}888{"step": 540, "records_seen": 17280, "elapsed_seconds": 5090.073473930359, "gradient_norm": 1.3931784629821777, "learning_rate": 6.693789463339208e-06, "head_lr": 3.3468947316696034e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.2578700671589758, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2578700671589758, "supervised_tokens": 533.125, "waves": 1.0, "detection_loss": 1.7802114671184903, "grounding_loss": 0.260672849054448, "detection_grounding_macro_loss": 1.020442158086469}889{"step": 550, "records_seen": 17600, "elapsed_seconds": 5186.911737918854, "gradient_norm": 1.5730959177017212, "learning_rate": 6.580843617376837e-06, "head_lr": 3.2904218086884183e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.4654623223468661, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4654623223468661, "supervised_tokens": 478.15625, "waves": 1.0, "detection_loss": 1.8953464964161748, "grounding_loss": 0.36686943305863273, "detection_grounding_macro_loss": 1.1311079647374038}890{"step": 560, "records_seen": 17920, "elapsed_seconds": 5286.556321620941, "gradient_norm": 1.802736520767212, "learning_rate": 6.467170410492089e-06, "head_lr": 3.2335852052460444e-07, "peak_gpu_gib": 25.681715965270996, "loss": 0.922025756282153, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.922025756282153, "supervised_tokens": 894.3125, "waves": 1.0, "detection_loss": 1.7098008011068617, "grounding_loss": 0.3093118325296018, "detection_grounding_macro_loss": 1.0095563168182318}891{"step": 570, "records_seen": 18240, "elapsed_seconds": 5378.557268381119, "gradient_norm": 1.710984706878662, "learning_rate": 6.352846339827827e-06, "head_lr": 3.176423169913913e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.1059783597174828, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1059783597174828, "supervised_tokens": 444.75, "waves": 1.0, "detection_loss": 1.5524997755885124, "grounding_loss": 0.3617759999324335, "detection_grounding_macro_loss": 0.9571378877604729}892{"step": 580, "records_seen": 18560, "elapsed_seconds": 5468.290054798126, "gradient_norm": 1.5063748359680176, "learning_rate": 6.2379483405300356e-06, "head_lr": 3.1189741702650174e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.5183463005973294, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5183463005973294, "supervised_tokens": 508.5625, "waves": 1.0, "detection_loss": 2.0708577876741234, "grounding_loss": 0.3028210290283823, "detection_grounding_macro_loss": 1.1868394083512528}893{"step": 590, "records_seen": 18880, "elapsed_seconds": 5551.7540118694305, "gradient_norm": 1.5484527349472046, "learning_rate": 6.122553733973795e-06, "head_lr": 3.061276866986897e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.1512338840402663, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1512338840402663, "supervised_tokens": 781.125, "waves": 1.0, "detection_loss": 2.121495254834493, "grounding_loss": 0.2951209098100662, "detection_grounding_macro_loss": 1.2083080823222796}894{"step": 600, "records_seen": 19200, "elapsed_seconds": 5644.422732830048, "gradient_norm": 1.4438694715499878, "learning_rate": 6.006740175729365e-06, "head_lr": 3.0033700878646824e-07, "peak_gpu_gib": 25.681715965270996, "loss": 1.5210001185270983, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5210001185270983, "supervised_tokens": 1128.25, "waves": 1.0, "detection_loss": 2.0113057548349556, "grounding_loss": 0.4423277186498126, "detection_grounding_macro_loss": 1.2268167367423841}895/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead896  warnings.warn(897/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead898  warnings.warn(899/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead900  warnings.warn(901/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead902  warnings.warn(903/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead904  warnings.warn(905/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead906  warnings.warn(907/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead908  warnings.warn(909/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead910  warnings.warn(911{"validation_step": 600, "loss": 1.1077201205024925, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1077201205024925, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6046872378033128, "grounding_loss": 0.544017418878419, "detection_grounding_macro_loss": 1.0743523283408658}912{"step": 610, "records_seen": 19520, "elapsed_seconds": 5890.47816824913, "gradient_norm": 1.351336121559143, "learning_rate": 5.890585603303328e-06, "head_lr": 2.9452928016516636e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.5024704816751182, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5024704816751182, "supervised_tokens": 548.71875, "waves": 1.0, "detection_loss": 1.9907513364501621, "grounding_loss": 0.2546416305833393, "detection_grounding_macro_loss": 1.1226964835167508}913{"step": 620, "records_seen": 19840, "elapsed_seconds": 5975.76282119751, "gradient_norm": 1.4534567594528198, "learning_rate": 5.774168183690039e-06, "head_lr": 2.887084091845019e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.3796392558142543, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3796392558142543, "supervised_tokens": 440.46875, "waves": 1.0, "detection_loss": 1.9584445016724723, "grounding_loss": 0.2746474228122018, "detection_grounding_macro_loss": 1.116545962242337}914{"step": 630, "records_seen": 20160, "elapsed_seconds": 6074.399399280548, "gradient_norm": 1.7068716287612915, "learning_rate": 5.657566260768623e-06, "head_lr": 2.828783130384311e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2056901729665697, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2056901729665697, "supervised_tokens": 569.4375, "waves": 1.0, "detection_loss": 1.8887038651634664, "grounding_loss": 0.43160798847675325, "detection_grounding_macro_loss": 1.1601559268201098}915{"step": 640, "records_seen": 20480, "elapsed_seconds": 6165.758774518967, "gradient_norm": 1.5336917638778687, "learning_rate": 5.5408583025809345e-06, "head_lr": 2.770429151290467e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.3731851395147885, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3731851395147885, "supervised_tokens": 351.1875, "waves": 1.0, "detection_loss": 1.8810450226068496, "grounding_loss": 0.5267520010280199, "detection_grounding_macro_loss": 1.2038985118174348}916{"step": 650, "records_seen": 20800, "elapsed_seconds": 6252.482875347137, "gradient_norm": 1.625543236732483, "learning_rate": 5.424122848525967e-06, "head_lr": 2.7120614242629833e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.0813259955320973, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0813259955320973, "supervised_tokens": 330.03125, "waves": 1.0, "detection_loss": 1.9555067930902754, "grounding_loss": 0.4014075974312921, "detection_grounding_macro_loss": 1.1784571952607839}917{"step": 660, "records_seen": 21120, "elapsed_seconds": 6344.542241096497, "gradient_norm": 1.4303293228149414, "learning_rate": 5.307438456506243e-06, "head_lr": 2.6537192282531215e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.3363272174225358, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3363272174225358, "supervised_tokens": 731.4375, "waves": 1.0, "detection_loss": 2.0060849457979204, "grounding_loss": 0.22006433679689508, "detection_grounding_macro_loss": 1.1130746412974077}918{"step": 670, "records_seen": 21440, "elapsed_seconds": 6425.496515750885, "gradient_norm": 1.7144300937652588, "learning_rate": 5.190883650061747e-06, "head_lr": 2.5954418250308737e-07, "peak_gpu_gib": 26.304283142089844, "loss": 0.9579395230909, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 0.9579395230909, "supervised_tokens": 267.84375, "waves": 1.0, "detection_loss": 1.73591006398201, "grounding_loss": 0.27149492818697957, "detection_grounding_macro_loss": 1.0037024960844947}919{"step": 680, "records_seen": 21760, "elapsed_seconds": 6508.047918558121, "gradient_norm": 1.6987581253051758, "learning_rate": 5.074536865526989e-06, "head_lr": 2.537268432763494e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2064995183609426, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2064995183609426, "supervised_tokens": 465.5, "waves": 1.0, "detection_loss": 2.0036348402500153, "grounding_loss": 0.40936419647186995, "detection_grounding_macro_loss": 1.2064995183609426}920{"step": 690, "records_seen": 22080, "elapsed_seconds": 6597.573613405228, "gradient_norm": 1.175805687904358, "learning_rate": 4.958476399246741e-06, "head_lr": 2.47923819962337e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.3882116216700524, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3882116216700524, "supervised_tokens": 614.34375, "waves": 1.0, "detection_loss": 2.010414394594374, "grounding_loss": 0.20036996426907452, "detection_grounding_macro_loss": 1.1053921794317243}921{"step": 700, "records_seen": 22400, "elapsed_seconds": 6689.172871828079, "gradient_norm": 1.5054681301116943, "learning_rate": 4.842780354886001e-06, "head_lr": 2.4213901774430003e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2224171900888905, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2224171900888905, "supervised_tokens": 450.03125, "waves": 1.0, "detection_loss": 2.0358553283354817, "grounding_loss": 0.30052063340942065, "detection_grounding_macro_loss": 1.1681879808724511}922{"step": 710, "records_seen": 22720, "elapsed_seconds": 6775.550484418869, "gradient_norm": 1.6136798858642578, "learning_rate": 4.727526590869606e-06, "head_lr": 2.3637632954348024e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.5497906559612602, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5497906559612602, "supervised_tokens": 473.59375, "waves": 1.0, "detection_loss": 2.272596986670243, "grounding_loss": 0.49338140338659286, "detection_grounding_macro_loss": 1.3829891950284179}923{"step": 720, "records_seen": 23040, "elapsed_seconds": 6863.585356712341, "gradient_norm": 1.612791657447815, "learning_rate": 4.612792667986877e-06, "head_lr": 2.3063963339934385e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.1962533795740455, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1962533795740455, "supervised_tokens": 542.0625, "waves": 1.0, "detection_loss": 1.9576836061828278, "grounding_loss": 0.3332991227507591, "detection_grounding_macro_loss": 1.1454913644667934}924{"step": 730, "records_seen": 23360, "elapsed_seconds": 6947.0133764743805, "gradient_norm": 1.4023692607879639, "learning_rate": 4.4986557971965865e-06, "head_lr": 2.2493278985982928e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.47105028713122, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.47105028713122, "supervised_tokens": 585.46875, "waves": 1.0, "detection_loss": 2.134516328573227, "grounding_loss": 0.36527355139454204, "detection_grounding_macro_loss": 1.2498949399838846}925{"step": 740, "records_seen": 23680, "elapsed_seconds": 7045.366122245789, "gradient_norm": 1.5677331686019897, "learning_rate": 4.385192787667313e-06, "head_lr": 2.192596393833656e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.304836132554442, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.304836132554442, "supervised_tokens": 490.90625, "waves": 1.0, "detection_loss": 1.804365331218356, "grounding_loss": 0.3511894805596967, "detection_grounding_macro_loss": 1.0777774058890264}926{"step": 750, "records_seen": 24000, "elapsed_seconds": 7136.008407592773, "gradient_norm": 1.3719826936721802, "learning_rate": 4.272479995088202e-06, "head_lr": 2.1362399975441008e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.5280273100361228, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5280273100361228, "supervised_tokens": 643.25, "waves": 1.0, "detection_loss": 1.8056704550981522, "grounding_loss": 0.3249070147673289, "detection_grounding_macro_loss": 1.0652887349327405}927{"step": 760, "records_seen": 24320, "elapsed_seconds": 7235.798012018204, "gradient_norm": 1.3085474967956543, "learning_rate": 4.160593270284906e-06, "head_lr": 2.0802966351424528e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.6012480647768825, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.6012480647768825, "supervised_tokens": 743.375, "waves": 1.0, "detection_loss": 1.9686617529392243, "grounding_loss": 0.28905632133994785, "detection_grounding_macro_loss": 1.1288590371395861}928{"step": 770, "records_seen": 24640, "elapsed_seconds": 7330.800405263901, "gradient_norm": 1.6082870960235596, "learning_rate": 4.049607908175248e-06, "head_lr": 2.0248039540876238e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.218310077674687, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.218310077674687, "supervised_tokens": 319.40625, "waves": 1.0, "detection_loss": 1.9675202825490166, "grounding_loss": 0.36920517881711323, "detection_grounding_macro_loss": 1.1683627306830648}929{"step": 780, "records_seen": 24960, "elapsed_seconds": 7416.380357980728, "gradient_norm": 1.6991201639175415, "learning_rate": 3.939598597099023e-06, "head_lr": 1.969799298549511e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.0643689678900614, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0643689678900614, "supervised_tokens": 284.875, "waves": 1.0, "detection_loss": 1.7901063300669193, "grounding_loss": 0.3386316057132035, "detection_grounding_macro_loss": 1.0643689678900614}930{"step": 790, "records_seen": 25280, "elapsed_seconds": 7510.249130964279, "gradient_norm": 1.427179217338562, "learning_rate": 3.830639368555962e-06, "head_lr": 1.9153196842779805e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2150423900288274, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2150423900288274, "supervised_tokens": 491.6875, "waves": 1.0, "detection_loss": 1.8954554137430693, "grounding_loss": 0.22059258613878158, "detection_grounding_macro_loss": 1.0580239999409253}931{"step": 800, "records_seen": 25600, "elapsed_seconds": 7604.407128095627, "gradient_norm": 1.2983521223068237, "learning_rate": 3.722803547385756e-06, "head_lr": 1.861401773692878e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2315234732828344, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2315234732828344, "supervised_tokens": 490.0625, "waves": 1.0, "detection_loss": 1.9388518960852372, "grounding_loss": 0.19773577841778417, "detection_grounding_macro_loss": 1.0682938372515107}932{"step": 810, "records_seen": 25920, "elapsed_seconds": 7695.878808498383, "gradient_norm": 1.667921543121338, "learning_rate": 3.616163702423615e-06, "head_lr": 1.8080818512118074e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.3683340107090771, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3683340107090771, "supervised_tokens": 506.78125, "waves": 1.0, "detection_loss": 2.02629018771021, "grounding_loss": 0.40670575201511383, "detection_grounding_macro_loss": 1.216497969862662}933{"step": 820, "records_seen": 26240, "elapsed_seconds": 7795.989032506943, "gradient_norm": 1.5328656435012817, "learning_rate": 3.510791597664569e-06, "head_lr": 1.7553957988322842e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2074708426371217, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2074708426371217, "supervised_tokens": 383.46875, "waves": 1.0, "detection_loss": 2.2394811590512593, "grounding_loss": 0.29687350462464723, "detection_grounding_macro_loss": 1.2681773318379532}934{"step": 830, "records_seen": 26560, "elapsed_seconds": 7883.976423501968, "gradient_norm": 1.5332063436508179, "learning_rate": 3.406758143969433e-06, "head_lr": 1.7033790719847163e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.4119987919111736, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4119987919111736, "supervised_tokens": 359.53125, "waves": 1.0, "detection_loss": 2.126338431709691, "grounding_loss": 0.3679639337441096, "detection_grounding_macro_loss": 1.2471511827269004}935{"step": 840, "records_seen": 26880, "elapsed_seconds": 7983.847622394562, "gradient_norm": 1.53697669506073, "learning_rate": 3.3041333513448536e-06, "head_lr": 1.6520666756724265e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2500378935812222, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2500378935812222, "supervised_tokens": 351.3125, "waves": 1.0, "detection_loss": 1.9736568695969052, "grounding_loss": 0.3196706387039155, "detection_grounding_macro_loss": 1.1466637541504103}936{"step": 850, "records_seen": 27200, "elapsed_seconds": 8070.778815746307, "gradient_norm": 1.4798164367675781, "learning_rate": 3.2029862818296164e-06, "head_lr": 1.601493140914808e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.4662093315273523, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4662093315273523, "supervised_tokens": 722.375, "waves": 1.0, "detection_loss": 1.966190297495235, "grounding_loss": 0.36625120639801023, "detection_grounding_macro_loss": 1.1662207519466226}937{"step": 860, "records_seen": 27520, "elapsed_seconds": 8155.691386461258, "gradient_norm": 1.4517751932144165, "learning_rate": 3.1033850030188797e-06, "head_lr": 1.5516925015094396e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.43508172291331, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.43508172291331, "supervised_tokens": 628.78125, "waves": 1.0, "detection_loss": 1.9538773054426366, "grounding_loss": 0.29373144134879114, "detection_grounding_macro_loss": 1.123804373395714}938{"step": 870, "records_seen": 27840, "elapsed_seconds": 8236.07496881485, "gradient_norm": 1.5090348720550537, "learning_rate": 3.0053965422576344e-06, "head_lr": 1.502698271128817e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.45494444668293, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.45494444668293, "supervised_tokens": 385.90625, "waves": 1.0, "detection_loss": 2.247432313467327, "grounding_loss": 0.29669294907496524, "detection_grounding_macro_loss": 1.272062631271146}939{"step": 880, "records_seen": 28160, "elapsed_seconds": 8315.190553665161, "gradient_norm": 1.5250667333602905, "learning_rate": 2.90908684153421e-06, "head_lr": 1.4545434207671048e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2897257895674556, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2897257895674556, "supervised_tokens": 536.0, "waves": 1.0, "detection_loss": 1.900668928027153, "grounding_loss": 0.27148722546796006, "detection_grounding_macro_loss": 1.0860780767475564}940{"step": 890, "records_seen": 28480, "elapsed_seconds": 8406.885415315628, "gradient_norm": 1.6281527280807495, "learning_rate": 2.8145207131041698e-06, "head_lr": 1.4072603565520847e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.5253934104694054, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5253934104694054, "supervised_tokens": 874.90625, "waves": 1.0, "detection_loss": 2.071767733855681, "grounding_loss": 0.32336989901959895, "detection_grounding_macro_loss": 1.19756881643764}941{"step": 900, "records_seen": 28800, "elapsed_seconds": 8494.43049955368, "gradient_norm": 1.2167465686798096, "learning_rate": 2.7217617958744695e-06, "head_lr": 1.3608808979372347e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.1849857764318585, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1849857764318585, "supervised_tokens": 550.03125, "waves": 1.0, "detection_loss": 2.07478065701092, "grounding_loss": 0.17655157844225566, "detection_grounding_macro_loss": 1.1256661177265879}942/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead943  warnings.warn(944/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead945  warnings.warn(946/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead947  warnings.warn(948/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead949  warnings.warn(950/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead951  warnings.warn(952/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead953  warnings.warn(954/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead955  warnings.warn(956/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead957  warnings.warn(958{"validation_step": 900, "loss": 1.1092656364278255, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1092656364278255, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6084073001993304, "grounding_loss": 0.5430963778070041, "detection_grounding_macro_loss": 1.0757518390031673}959{"step": 910, "records_seen": 29120, "elapsed_seconds": 8721.202590942383, "gradient_norm": 1.6279257535934448, "learning_rate": 2.6308725125772304e-06, "head_lr": 1.315436256288615e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.1393084630907424, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1393084630907424, "supervised_tokens": 431.8125, "waves": 1.0, "detection_loss": 1.5854166788714272, "grounding_loss": 0.2876473238730713, "detection_grounding_macro_loss": 0.9365320013722492}960{"step": 920, "records_seen": 29440, "elapsed_seconds": 8816.34403181076, "gradient_norm": 1.5059194564819336, "learning_rate": 2.5419140277619514e-06, "head_lr": 1.2709570138809754e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.416412405204028, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.416412405204028, "supervised_tokens": 401.4375, "waves": 1.0, "detection_loss": 2.0268187493085863, "grounding_loss": 0.3990684983630975, "detection_grounding_macro_loss": 1.212943623835842}961{"step": 930, "records_seen": 29760, "elapsed_seconds": 8907.648521900177, "gradient_norm": 1.7387686967849731, "learning_rate": 2.4549462066344104e-06, "head_lr": 1.2274731033172052e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.0621722039611825, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0621722039611825, "supervised_tokens": 335.40625, "waves": 1.0, "detection_loss": 1.7054985130534452, "grounding_loss": 0.33306905365661804, "detection_grounding_macro_loss": 1.0192837833550317}962{"step": 940, "records_seen": 30080, "elapsed_seconds": 9014.87066602707, "gradient_norm": 1.3319677114486694, "learning_rate": 2.3700275747699776e-06, "head_lr": 1.1850137873849886e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2205005884534827, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2205005884534827, "supervised_tokens": 403.1875, "waves": 1.0, "detection_loss": 2.0041077915165157, "grounding_loss": 0.21300561308672578, "detection_grounding_macro_loss": 1.1085567023016207}963{"step": 950, "records_seen": 30400, "elapsed_seconds": 9105.831940412521, "gradient_norm": 1.716751217842102, "learning_rate": 2.2872152787284386e-06, "head_lr": 1.1436076393642192e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.1655170655177187, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1655170655177187, "supervised_tokens": 294.6875, "waves": 1.0, "detection_loss": 2.015223644177119, "grounding_loss": 0.4157759667006009, "detection_grounding_macro_loss": 1.21549980543886}964{"step": 960, "records_seen": 30720, "elapsed_seconds": 9191.99558877945, "gradient_norm": 1.3490896224975586, "learning_rate": 2.2065650475968386e-06, "head_lr": 1.1032825237984191e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.5313877779990435, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5313877779990435, "supervised_tokens": 1085.09375, "waves": 1.0, "detection_loss": 1.8778168118000031, "grounding_loss": 0.29414122870990206, "detection_grounding_macro_loss": 1.0859790202549526}965{"step": 970, "records_seen": 31040, "elapsed_seconds": 9284.005653142929, "gradient_norm": 1.5163516998291016, "learning_rate": 2.128131155486214e-06, "head_lr": 1.0640655777431068e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2043182871725975, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2043182871725975, "supervised_tokens": 682.0625, "waves": 1.0, "detection_loss": 1.7969343447022967, "grounding_loss": 0.44238335606298407, "detection_grounding_macro_loss": 1.1196588503826403}966{"step": 980, "records_seen": 31360, "elapsed_seconds": 9374.062463998795, "gradient_norm": 1.415270447731018, "learning_rate": 2.051966385007477e-06, "head_lr": 1.0259831925037386e-07, "peak_gpu_gib": 26.304283142089844, "loss": 1.2527698238845915, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2527698238845915, "supervised_tokens": 737.65625, "waves": 1.0, "detection_loss": 2.0252549052238464, "grounding_loss": 0.2595747193055494, "detection_grounding_macro_loss": 1.142414812264698}967{"step": 990, "records_seen": 31680, "elapsed_seconds": 9471.205087184906, "gradient_norm": 1.7701075077056885, "learning_rate": 1.978121991750999e-06, "head_lr": 9.890609958754994e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.0860920906998217, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0860920906998217, "supervised_tokens": 300.34375, "waves": 1.0, "detection_loss": 1.633134169711007, "grounding_loss": 0.38275227482829777, "detection_grounding_macro_loss": 1.0079432222696525}968{"step": 1000, "records_seen": 32000, "elapsed_seconds": 9565.669588327408, "gradient_norm": 1.545108437538147, "learning_rate": 1.9066476697938216e-06, "head_lr": 9.533238348969106e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.1309864087836559, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1309864087836559, "supervised_tokens": 859.59375, "waves": 1.0, "detection_loss": 1.8702485421124626, "grounding_loss": 0.29315599101100814, "detection_grounding_macro_loss": 1.0817022665617353}969{"step": 1010, "records_seen": 32320, "elapsed_seconds": 9647.91195011139, "gradient_norm": 1.5181052684783936, "learning_rate": 1.8375915182576934e-06, "head_lr": 9.187957591288466e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.3711203089915216, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3711203089915216, "supervised_tokens": 421.75, "waves": 1.0, "detection_loss": 2.0381461570137427, "grounding_loss": 0.3962363772667371, "detection_grounding_macro_loss": 1.2171912671402398}970{"step": 1020, "records_seen": 32640, "elapsed_seconds": 9734.055211782455, "gradient_norm": 1.4054203033447266, "learning_rate": 1.7710000089404493e-06, "head_lr": 8.855000044702245e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.6098391411360353, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.6098391411360353, "supervised_tokens": 621.15625, "waves": 1.0, "detection_loss": 1.9836223936080932, "grounding_loss": 0.2748989537358284, "detection_grounding_macro_loss": 1.1292606736719608}971{"step": 1030, "records_seen": 32960, "elapsed_seconds": 9826.632216215134, "gradient_norm": 1.5838711261749268, "learning_rate": 1.7069179550425026e-06, "head_lr": 8.534589775212511e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.2736646900884807, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2736646900884807, "supervised_tokens": 507.375, "waves": 1.0, "detection_loss": 1.796183305978775, "grounding_loss": 0.4028003302713235, "detection_grounding_macro_loss": 1.0994918181250493}972{"step": 1040, "records_seen": 33280, "elapsed_seconds": 9920.888817071915, "gradient_norm": 1.6130170822143555, "learning_rate": 1.6453884810094984e-06, "head_lr": 8.22694240504749e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.3935268576404667, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3935268576404667, "supervised_tokens": 547.09375, "waves": 1.0, "detection_loss": 1.9095417034058344, "grounding_loss": 0.4084076066338556, "detection_grounding_macro_loss": 1.158974655019845}973{"step": 1050, "records_seen": 33600, "elapsed_seconds": 10010.569228172302, "gradient_norm": 1.7808738946914673, "learning_rate": 1.5864529935114375e-06, "head_lr": 7.932264967557187e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.2812218043254688, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2812218043254688, "supervised_tokens": 608.125, "waves": 1.0, "detection_loss": 1.875836664124539, "grounding_loss": 0.4121693169268278, "detection_grounding_macro_loss": 1.1440029905256834}974{"step": 1060, "records_seen": 33920, "elapsed_seconds": 10113.566853523254, "gradient_norm": 1.1638743877410889, "learning_rate": 1.5301511535777785e-06, "head_lr": 7.650755767888892e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.1848948919689946, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1848948919689946, "supervised_tokens": 1186.46875, "waves": 1.0, "detection_loss": 1.818583419919014, "grounding_loss": 0.12874734538562885, "detection_grounding_macro_loss": 0.9736653826523215}975{"step": 1070, "records_seen": 34240, "elapsed_seconds": 10194.842452764511, "gradient_norm": 1.5247225761413574, "learning_rate": 1.4765208499072787e-06, "head_lr": 7.382604249536392e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.491091577685438, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.491091577685438, "supervised_tokens": 805.90625, "waves": 1.0, "detection_loss": 1.854976156949997, "grounding_loss": 0.1915037945977279, "detection_grounding_macro_loss": 1.0232399757738624}976{"step": 1080, "records_seen": 34560, "elapsed_seconds": 10277.024828195572, "gradient_norm": 1.488142490386963, "learning_rate": 1.4255981733705498e-06, "head_lr": 7.127990866852749e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.1772950956131467, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1772950956131467, "supervised_tokens": 488.8125, "waves": 1.0, "detection_loss": 1.892207283112738, "grounding_loss": 0.25812228311367236, "detection_grounding_macro_loss": 1.0751647831132052}977{"step": 1090, "records_seen": 34880, "elapsed_seconds": 10360.371773958206, "gradient_norm": 1.5025273561477661, "learning_rate": 1.3774173927224625e-06, "head_lr": 6.887086963612312e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.295951577834785, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.295951577834785, "supervised_tokens": 632.375, "waves": 1.0, "detection_loss": 1.7654454030773856, "grounding_loss": 0.2630651623010635, "detection_grounding_macro_loss": 1.0142552826892246}978{"step": 1100, "records_seen": 35200, "elapsed_seconds": 10449.240037202835, "gradient_norm": 1.5522024631500244, "learning_rate": 1.3320109315407546e-06, "head_lr": 6.660054657703771e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.352504185196608, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.352504185196608, "supervised_tokens": 847.53125, "waves": 1.0, "detection_loss": 2.0293275274728475, "grounding_loss": 0.36330083879287345, "detection_grounding_macro_loss": 1.1963141831328605}979{"step": 1110, "records_seen": 35520, "elapsed_seconds": 10548.903951644897, "gradient_norm": 1.2924150228500366, "learning_rate": 1.289409346406375e-06, "head_lr": 6.447046732031874e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.0763200692897144, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0763200692897144, "supervised_tokens": 933.84375, "waves": 1.0, "detection_loss": 2.0198075771331787, "grounding_loss": 0.24383109178077522, "detection_grounding_macro_loss": 1.131819334456977}980{"step": 1120, "records_seen": 35840, "elapsed_seconds": 10638.683242797852, "gradient_norm": 1.2397348880767822, "learning_rate": 1.2496413063402264e-06, "head_lr": 6.248206531701132e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.2148762350176554, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2148762350176554, "supervised_tokens": 437.9375, "waves": 1.0, "detection_loss": 2.155588384936838, "grounding_loss": 0.14873579844258225, "detection_grounding_macro_loss": 1.15216209168971}981{"step": 1130, "records_seen": 36160, "elapsed_seconds": 10734.159244775772, "gradient_norm": 1.3885751962661743, "learning_rate": 1.2127335735101542e-06, "head_lr": 6.06366786755077e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.4768923000719774, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4768923000719774, "supervised_tokens": 783.3125, "waves": 1.0, "detection_loss": 1.9233024296435444, "grounding_loss": 0.4947900150145301, "detection_grounding_macro_loss": 1.2090462223290372}982{"step": 1140, "records_seen": 36480, "elapsed_seconds": 10814.161434173584, "gradient_norm": 1.5205278396606445, "learning_rate": 1.17871098522117e-06, "head_lr": 5.89355492610585e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.2908304282464087, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2908304282464087, "supervised_tokens": 600.5, "waves": 1.0, "detection_loss": 1.6587537554177372, "grounding_loss": 0.48139910846948625, "detection_grounding_macro_loss": 1.0700764319436118}983{"step": 1150, "records_seen": 36800, "elapsed_seconds": 10897.2194480896, "gradient_norm": 1.6045104265213013, "learning_rate": 1.147596437201019e-06, "head_lr": 5.7379821860050944e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.5933989314362407, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.5933989314362407, "supervised_tokens": 475.0625, "waves": 1.0, "detection_loss": 2.2286291463034495, "grounding_loss": 0.38068670305338775, "detection_grounding_macro_loss": 1.3046579246784187}984{"step": 1160, "records_seen": 37120, "elapsed_seconds": 10983.332487344742, "gradient_norm": 1.6121147871017456, "learning_rate": 1.1194108681923507e-06, "head_lr": 5.5970543409617525e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.3861207305453718, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3861207305453718, "supervised_tokens": 436.3125, "waves": 1.0, "detection_loss": 2.080308434036043, "grounding_loss": 0.4935936832002231, "detection_grounding_macro_loss": 1.2869510586181332}985{"step": 1170, "records_seen": 37440, "elapsed_seconds": 11065.81792640686, "gradient_norm": 1.6877046823501587, "learning_rate": 1.0941732458618478e-06, "head_lr": 5.470866229309239e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.2392212244437104, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.2392212244437104, "supervised_tokens": 408.40625, "waves": 1.0, "detection_loss": 1.7964064538478852, "grounding_loss": 0.31057917543675256, "detection_grounding_macro_loss": 1.0534928146423188}986{"step": 1180, "records_seen": 37760, "elapsed_seconds": 11144.793893575668, "gradient_norm": 1.7247551679611206, "learning_rate": 1.0719005540358127e-06, "head_lr": 5.3595027701790636e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.1012538463540906, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1012538463540906, "supervised_tokens": 439.0, "waves": 1.0, "detection_loss": 2.02616218158177, "grounding_loss": 0.3818806967325625, "detection_grounding_macro_loss": 1.2040214391571662}987{"step": 1190, "records_seen": 38080, "elapsed_seconds": 11238.168888568878, "gradient_norm": 1.4584319591522217, "learning_rate": 1.0526077812707857e-06, "head_lr": 5.263038906353928e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.4675516392321697, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.4675516392321697, "supervised_tokens": 1088.75, "waves": 1.0, "detection_loss": 1.9904171625773113, "grounding_loss": 0.46935382193689956, "detection_grounding_macro_loss": 1.2298854922571054}988{"step": 1200, "records_seen": 38400, "elapsed_seconds": 11330.678287506104, "gradient_norm": 1.8357263803482056, "learning_rate": 1.0363079107668967e-06, "head_lr": 5.181539553834483e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.0606519773136824, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0606519773136824, "supervised_tokens": 461.125, "waves": 1.0, "detection_loss": 1.712966626510024, "grounding_loss": 0.4083373281173408, "detection_grounding_macro_loss": 1.0606519773136824}989/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead990  warnings.warn(991/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead992  warnings.warn(993/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead994  warnings.warn(995/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead996  warnings.warn(997/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead998  warnings.warn(999/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1000  warnings.warn(1001/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1002  warnings.warn(1003/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1004  warnings.warn(1005{"validation_step": 1200, "loss": 1.1037772887085915, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1037772887085915, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.5998848528725562, "grounding_loss": 0.5410495659283229, "detection_grounding_macro_loss": 1.0704672094004395}1006{"step": 1210, "records_seen": 38720, "elapsed_seconds": 11556.73000574112, "gradient_norm": 1.5140602588653564, "learning_rate": 1.0230119116307337e-06, "head_lr": 5.1150595581536676e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.0829605648763163, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.0829605648763163, "supervised_tokens": 258.875, "waves": 1.0, "detection_loss": 1.838232720270753, "grounding_loss": 0.3276884094818797, "detection_grounding_macro_loss": 1.0829605648763163}1007{"step": 1220, "records_seen": 39040, "elapsed_seconds": 11637.56585264206, "gradient_norm": 1.3675578832626343, "learning_rate": 1.0127287314936108e-06, "head_lr": 5.063643657468053e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.3630990117089823, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3630990117089823, "supervised_tokens": 664.21875, "waves": 1.0, "detection_loss": 1.7418638542294502, "grounding_loss": 0.22680448414757848, "detection_grounding_macro_loss": 0.9843341691885144}1008{"step": 1230, "records_seen": 39360, "elapsed_seconds": 11733.456425905228, "gradient_norm": 1.5297515392303467, "learning_rate": 1.0054652904901968e-06, "head_lr": 5.0273264524509834e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.3640903916652363, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.3640903916652363, "supervised_tokens": 651.03125, "waves": 1.0, "detection_loss": 1.9149661745343889, "grounding_loss": 0.31241844255139933, "detection_grounding_macro_loss": 1.113692308542894}1009{"step": 1240, "records_seen": 39680, "elapsed_seconds": 11824.26021194458, "gradient_norm": 1.679861068725586, "learning_rate": 1.0012264766015687e-06, "head_lr": 5.0061323830078426e-08, "peak_gpu_gib": 26.304283142089844, "loss": 1.1270903375698254, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1270903375698254, "supervised_tokens": 277.6875, "waves": 1.0, "detection_loss": 1.8670303155394161, "grounding_loss": 0.2884916958709558, "detection_grounding_macro_loss": 1.0777610057051858}1010/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1011  warnings.warn(1012/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1013  warnings.warn(1014/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1015  warnings.warn(1016/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1017  warnings.warn(1018/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1019  warnings.warn(1020/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1021  warnings.warn(1022/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1023  warnings.warn(1024/dccstor/urban/mustansar/envs/refwave/lib/python3.10/site-packages/transformers/models/auto/image_processing_auto.py:647: FutureWarning: The image_processor_class argument is deprecated and will be removed in v4.42. Please use `slow_image_processor_class`, or `fast_image_processor_class` instead1025  warnings.warn(1026{"validation_step": 1248, "loss": 1.1051623993970883, "round1_loss": 0.0, "round2_loss": 0.0, "direct_loss": 1.1051623993970883, "supervised_tokens": 1495.4779116465863, "waves": 1.0, "detection_loss": 1.6025433617219405, "grounding_loss": 0.5409902792743274, "detection_grounding_macro_loss": 1.0717668204981339}1027wandb: updating run metadata1028wandb: uploading output.log; uploading wandb-summary.json; uploading config.yaml1029wandb: 1030wandb: Run history:1031wandb:                       optimizer_step โ–โ–โ–โ–โ–โ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–„โ–„โ–„โ–„โ–…โ–…โ–…โ–…โ–…โ–†โ–†โ–†โ–†โ–†โ–†โ–‡โ–‡โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆ1032wandb: train/detection_grounding_macro_loss โ–…โ–ƒโ–…โ–‡โ–†โ–…โ–…โ–ƒโ–‚โ–ƒโ–…โ–„โ–ƒโ–‡โ–‡โ–ˆโ–„โ–„โ–†โ–‚โ–…โ–†โ–†โ–…โ–„โ–‚โ–„โ–†โ–„โ–…โ–„โ–…โ–†โ–…โ–ƒโ–†โ–โ–†โ–„โ–†1033wandb:                 train/detection_loss โ–„โ–„โ–†โ–…โ–ƒโ–„โ–„โ–„โ–„โ–†โ–ƒโ–„โ–…โ–โ–‡โ–…โ–ˆโ–„โ–‚โ–…โ–…โ–โ–†โ–…โ–†โ–‚โ–ƒโ–ƒโ–ƒโ–…โ–…โ–ƒโ–‡โ–„โ–ƒโ–‡โ–‚โ–‚โ–…โ–‚1034wandb:                    train/direct_loss โ–„โ–‡โ–ƒโ–†โ–‡โ–…โ–…โ–†โ–ƒโ–…โ–…โ–„โ–†โ–‡โ–…โ–„โ–…โ–‡โ–…โ–†โ–โ–โ–‡โ–‡โ–…โ–†โ–„โ–ˆโ–…โ–ƒโ–†โ–ƒโ–ƒโ–…โ–…โ–‚โ–…โ–„โ–†โ–ƒ1035wandb:                train/elapsed_seconds โ–โ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–„โ–„โ–„โ–„โ–„โ–…โ–…โ–…โ–…โ–…โ–…โ–…โ–…โ–…โ–…โ–…โ–†โ–†โ–‡โ–‡โ–ˆโ–ˆโ–ˆโ–ˆ1036wandb:                  train/gradient_norm โ–„โ–…โ–ƒโ–…โ–‚โ–ƒโ–โ–…โ–โ–„โ–„โ–…โ–ˆโ–ƒโ–†โ–…โ–„โ–‡โ–‡โ–†โ–ƒโ–‚โ–…โ–†โ–‡โ–ˆโ–‡โ–ƒโ–‡โ–ƒโ–†โ–‚โ–„โ–…โ–‡โ–„โ–†โ–„โ–„โ–„1037wandb:                 train/grounding_loss โ–ƒโ–‚โ–…โ–ˆโ–„โ–ƒโ–โ–‚โ–‚โ–…โ–โ–ƒโ–…โ–ƒโ–ƒโ–„โ–‚โ–‚โ–ƒโ–…โ–„โ–โ–ƒโ–…โ–…โ–…โ–…โ–†โ–„โ–…โ–ˆโ–ƒโ–ƒโ–‚โ–ƒโ–†โ–‚โ–ƒโ–„โ–†1038wandb:                        train/head_lr โ–โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‡โ–‡โ–‡โ–‡โ–‡โ–‡โ–†โ–†โ–†โ–†โ–†โ–…โ–…โ–…โ–„โ–„โ–„โ–„โ–ƒโ–ƒโ–ƒโ–ƒโ–‚โ–‚โ–‚โ–‚โ–‚โ–‚โ–โ–โ–โ–1039wandb:                  train/learning_rate โ–‚โ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–ˆโ–‡โ–‡โ–‡โ–‡โ–‡โ–‡โ–‡โ–†โ–†โ–†โ–†โ–…โ–…โ–…โ–…โ–…โ–„โ–„โ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–ƒโ–‚โ–‚โ–‚โ–‚โ–‚โ–โ–โ–1040wandb:                           train/loss โ–…โ–…โ–‡โ–‡โ–†โ–„โ–‚โ–‚โ–…โ–‚โ–„โ–ˆโ–…โ–†โ–โ–†โ–ˆโ–„โ–„โ–„โ–…โ–‚โ–„โ–†โ–‚โ–ƒโ–†โ–†โ–‚โ–„โ–ˆโ–…โ–…โ–†โ–„โ–…โ–…โ–…โ–†โ–…1041wandb:                                  +16 ...1042wandb: 1043wandb: Run summary:1044wandb:                       optimizer_step 12481045wandb: train/detection_grounding_macro_loss 1.185121046wandb:                 train/detection_loss 1.900371047wandb:                    train/direct_loss 1.491651048wandb:                train/elapsed_seconds 11893.776831049wandb:                  train/gradient_norm 2.4241050wandb:                 train/grounding_loss 0.469861051wandb:                        train/head_lr 0.01052wandb:                  train/learning_rate 0.01053wandb:                           train/loss 1.491651054wandb:                                  +16 ...1055wandb: 1056wandb: ๐Ÿš€ View run standard_lora_4gpu_20260908T185018Z_lJdQkO_stage2 at: https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale/runs/fa7d4d62be7f59b51057wandb: โญ๏ธ View project at: https://wandb.ai/shubhamrpatle-mbzuai/refwave-fullscale1058wandb: Synced 5 W&B file(s), 0 media file(s), 0 artifact file(s) and 0 other file(s)1059wandb: Find logs at: /dccstor/pretrain5/mustansar/sub_pivr/refwave-assets/runs/standard_lora_4gpu_20260908T185018Z_lJdQkO/stage2/wandb/run-20260908_204844-fa7d4d62be7f59b5/logs1060