happy8825/experiment_bitsandbytes
/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350 · happy8825/valid_ecva_clean results Model: /hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350 Dataset: happy8825/valid_ecva_clean Generated: 2026-01-08 11:28:44Z Metrics Metric Value Total samples 924 With GT 0 Parsed answers 0 Top-1 accuracy 0 Recall@5 0 MRR 0 The uploaded JSON contains full per-sample predictions produced via t3_infer_with_vllm.bash.… See the full description on the dataset page: https://huggingface.co/datasets/happy8825/experiment_bitsandbytes.
06
1---2pretty_name: "/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350 · happy8825/valid_ecva_clean results"3language:4- en5tags:6- video-retrieval7- evaluation8- vllm9---10 11# /hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350 · happy8825/valid_ecva_clean results12 13- **Model**: `/hub_data4/seohyun/saves/ecva_instruct_1223/full/sft/checkpoint-350`14- **Dataset**: `happy8825/valid_ecva_clean`15- **Generated**: `2026-01-08 11:28:44Z`16 17## Metrics18| Metric | Value |19| --- | --- |20| Total samples | 924 |21| With GT | 0 |22| Parsed answers | 0 |23| Top-1 accuracy | 0 |24| Recall@5 | 0 |25| MRR | 0 |26 27The uploaded JSON contains full per-sample predictions produced via `t3_infer_with_vllm.bash`.28 29### EVQA/ECVA Metrics30| Metric | Value |31| --- | --- |32| EVQA total | 924 |33| EVQA with GT label | 924 |34| EVQA accuracy | 0.731602 |35 36## Run Summary37 38```39Saved 924 results to /home/seohyun/vid_understanding/video_retrieval/video_retrieval/output_ecva/experiment_bitsandbytes.json40Metrics: {41 "total": 924,42 "with_gt": 0,43 "with_parsed_answer": 0,44 "top1_acc": 0.0,45 "recall_at_5": 0.0,46 "mrr": 0.0,47 "num_shards": 1,48 "shard_index": 0,49 "evqa_total": 924,50 "evqa_with_gt_label": 924,51 "evqa_acc": 0.731601731601731652}53Pushed experiment_bitsandbytes.jsonl and README to https://huggingface.co/datasets/happy8825/experiment_bitsandbytes54```55 56 