CoolFace
Modelpublic

Mexico-Oxford-TEAM/frame-qwen3vl-8b-r61-ep2

sourceHugging Facecc-by-nc-sa-4.0updated 17d agoView on Hugging Face
0likes28downloads
Model Card

FRAME Qwen3-VL-8B - rung 61, epoch 2

The checkpoint submitted to the challenge's final test phase — model B, run alongside `r61-ep4`.

Fine-tuned Qwen3-VL-8B-Instruct for the **ORENA SAVE FOCUS Challenge — FRAME track** (MICCAI 2026), by team Mexico-Oxford_TEAM.

📦 Code, training configs, every experiment and every negative result: [`RodMed0709/ORENA-Challenge-MEXICO`](https://github.com/RodMed0709/ORENA-Challenge-MEXICO)

🔗 This checkpoint is half of a pair

Our submitted system runs two checkpoints, not one. This repository is model B; its partner is [`frame-qwen3vl-8b-r61-ep4`](https://huggingface.co/Mexico-Oxford-TEAM/frame-qwen3vl-8b-r61-ep4) (model A). Both are epochs of the same rung 61 training run — same data, same recipe, different epoch.

The rule. Model A answers. If A's answer parses entirely as legal foreign-object class names, model B is asked the same question, and if B returns a strictly shorter list, B's answer wins; ties and anything that is not a class list keep A. Numbers, yes/no and multiple-choice answers never parse as class lists, so they fall through untouched.

python
ans_a = ask(model_a, question, frame)
ans_b = ask(model_b, question, frame)          # only needed when ans_a is a class list
final = ans_b if is_class_list(ans_a) and is_class_list(ans_b) \
                 and len(set_of(ans_b)) < len(set_of(ans_a)) else ans_a

Why it works, measured on `fo_class` questions:

model A alonewith the pair
out-of-centre probe (another hospital)0.36180.3838 (+0.0221, 15/15 leave-one-video-out folds)
in-domain control0.84080.8367 (−0.0041, within noise)

The mechanism is not "two heads are better". This checkpoint's dominant error is naming a class that is not in the frame, and a differently-trained sibling disagrees about which extra class often enough to delete some of them.

🔴 The direction is the whole lever. Taking the union of the two answers instead of the shorter list costs −0.1562. Intersect, never combine.

Either checkpoint also works alone — that is what our submission 03 did, and it scored 0.5809.

Score

Not published. This is the entry for the challenge's final test phase, whose leaderboard was not public at upload time. It is the same recipe as r42, retrained on the full released corpus — all 130 videos and 20,667 rows, against r42's 122 videos and 19,384 rows.

If you want a checkpoint with a verified leaderboard number, use `frame-qwen3vl-8b-r42-ep4` (0.5809).

What is in this repository

rootthe full merged model — load it directly, no extra step
adapter/the LoRA adapter alone (~100 MB), if you would rather apply it to the base yourself

The root holds the standard Hugging Face file set for this architecture — the same one Qwen/Qwen3-VL-8B-Instruct ships: config.json (architecture and dtype), the tokenizer files (tokenizer.json, vocab.json, merges.txt, added_tokens.json, special_tokens_map.json, tokenizer_config.json), chat_template.jinja (the prompt format — without it the model is prompted wrong and answers badly), preprocessor_config.json and video_preprocessor_config.json (how images and video frames are resized and normalised), generation_config.json, and model.safetensors.index.json (which tensor lives in which shard).

python
from transformers import AutoProcessor, AutoModelForImageTextToText
model = AutoModelForImageTextToText.from_pretrained(
    "Mexico-Oxford-TEAM/frame-qwen3vl-8b-r61-ep2", dtype="bfloat16", device_map="auto")
processor = AutoProcessor.from_pretrained("Mexico-Oxford-TEAM/frame-qwen3vl-8b-r61-ep2")

Requirements

These are the versions this checkpoint was merged and evaluated on — the exact set our submission container ran with:

transformers==4.57.6
torch==2.5.1          # CUDA 12.4 build
qwen-vl-utils==0.0.14
accelerate>=1.0.0,<2.0
numpy>=1.23.0,<2.0    # numpy 2.x is untested against this torch build
pillow>=9.0,<14.0
⚠️ Stay on the `transformers` 4.57 line. The config here is written in the 4.57 layout, and 5.x writes an incompatible one for this architecture. Mixing the two produces a load-time failure, not a silent degradation, so you will know immediately.

Training recipe

[ms-swift](https://github.com/modelscope/ms-swift) 4.4, tuner_type: lora. Every value is read from adapter/args.json in this repository, not from notes:

base modelQwen/Qwen3-VL-8B-Instruct (Apache-2.0)
LoRA rank / alpha / dropout8 / 32 / 0.1
target modulesall-linear plus the vision merger and all three `deepstack` mergers
vision towernot frozen (freeze_vit: false, freeze_aligner: false)
learning rate2e-4, cosine, warmup 0.03, weight decay 0.1
batch1 x grad-accum 16
epochs5 total; this is epoch 2
precision / seedbfloat16 / 42

Intended use and limits

Research artifact from a benchmark challenge on laparoscopic surgical video. It answers short questions about foreign objects — clips, drains, specimen bags, needles, threads — in a frame.

🔴 Not a medical device. Not for clinical decision-making.

Known weakness, measured rather than assumed: it cannot enumerate reliably past about two objects. On external laparoscopic frames with large metallic instruments, accuracy falls to 0.04 at three objects. That is a perception limit, not a formatting one — it reproduces across two unrelated answer formats and two different datasets.

Data

Trained only on the challenge's own corpus (ORENA SAVE FOCUS FRAME: `heico-focus-vqa` + `lapchole-focus-vqa`), built on the Heidelberg Colorectal dataset (doi:10.1038/s41597-021-00882-2). No external surgical corpus is in this checkpoint — several were tried and each closed as a measured negative.

License

CC BY-NC-SA 4.0. The backbone is Apache-2.0, but the fine-tuning data is CC BY-NC-SA and the stricter term governs the derivative. NonCommercial, ShareAlike, attribution required.

Links

Code and experiments`RodMed0709/ORENA-Challenge-MEXICO`
ChallengeFRAME track
Training data`heico-focus-vqa` · `lapchole-focus-vqa` — 🔒 access must be requested from the organizers

Team

`Mexico-Oxford_TEAM`

📌 TODO — team members. Names, affiliations, ORCIDs and roles — each member fills in their own. (placeholder: to be completed by the team)