Mexico-Oxford-TEAM/frame-qwen3vl-8b-r42-ep2
FRAME Qwen3-VL-8B - rung 42, epoch 2
Model B of the checkpoint pair that scored 0.58128 on the official FRAME leaderboard. Designed to be run alongside `r42-ep4`, not on its own.
Fine-tuned Qwen3-VL-8B-Instruct for the **ORENA SAVE FOCUS Challenge โ FRAME track** (MICCAI 2026), by team Mexico-Oxford_TEAM.
๐ฆ Code, training configs, every experiment and every negative result: [`RodMed0709/ORENA-Challenge-MEXICO`](https://github.com/RodMed0709/ORENA-Challenge-MEXICO)
๐ This checkpoint is half of a pair
Our submitted system runs two checkpoints, not one. This repository is model B; its partner is [`frame-qwen3vl-8b-r42-ep4`](https://huggingface.co/Mexico-Oxford-TEAM/frame-qwen3vl-8b-r42-ep4) (model A). Both are epochs of the same rung 42 training run โ same data, same recipe, different epoch.
The rule. Model A answers. If A's answer parses entirely as legal foreign-object class names, model B is asked the same question, and if B returns a strictly shorter list, B's answer wins; ties and anything that is not a class list keep A. Numbers, yes/no and multiple-choice answers never parse as class lists, so they fall through untouched.
ans_a = ask(model_a, question, frame)
ans_b = ask(model_b, question, frame) # only needed when ans_a is a class list
final = ans_b if is_class_list(ans_a) and is_class_list(ans_b) \
and len(set_of(ans_b)) < len(set_of(ans_a)) else ans_aWhy it works, measured on `fo_class` questions:
The mechanism is not "two heads are better". This checkpoint's dominant error is naming a class that is not in the frame, and a differently-trained sibling disagrees about which extra class often enough to delete some of them.
๐ด The direction is the whole lever. Taking the union of the two answers instead of the shorter list costs โ0.1562. Intersect, never combine.
Either checkpoint also works alone โ that is what our submission 03 did, and it scored 0.5809.
Score on the official FRAME leaderboard
This is model B, so it has no solo score. The pair โ r42-ep4 as A, this as B โ scored 0.58128, against official baselines of 0.5189 (fine-tuned) and 0.3883 (proprietary).
What is in this repository
The root holds the standard Hugging Face file set for this architecture โ the same one Qwen/Qwen3-VL-8B-Instruct ships: config.json (architecture and dtype), the tokenizer files (tokenizer.json, vocab.json, merges.txt, added_tokens.json, special_tokens_map.json, tokenizer_config.json), chat_template.jinja (the prompt format โ without it the model is prompted wrong and answers badly), preprocessor_config.json and video_preprocessor_config.json (how images and video frames are resized and normalised), generation_config.json, and model.safetensors.index.json (which tensor lives in which shard).
from transformers import AutoProcessor, AutoModelForImageTextToText
model = AutoModelForImageTextToText.from_pretrained(
"Mexico-Oxford-TEAM/frame-qwen3vl-8b-r42-ep2", dtype="bfloat16", device_map="auto")
processor = AutoProcessor.from_pretrained("Mexico-Oxford-TEAM/frame-qwen3vl-8b-r42-ep2")Requirements
These are the versions this checkpoint was merged and evaluated on โ the exact set our submission container ran with:
transformers==4.57.6
torch==2.5.1 # CUDA 12.4 build
qwen-vl-utils==0.0.14
accelerate>=1.0.0,<2.0
numpy>=1.23.0,<2.0 # numpy 2.x is untested against this torch build
pillow>=9.0,<14.0โ ๏ธ Stay on the `transformers` 4.57 line. The config here is written in the 4.57 layout, and 5.x writes an incompatible one for this architecture. Mixing the two produces a load-time failure, not a silent degradation, so you will know immediately.
Training recipe
[ms-swift](https://github.com/modelscope/ms-swift) 4.4, tuner_type: lora. Every value is read from adapter/args.json in this repository, not from notes:
Intended use and limits
Research artifact from a benchmark challenge on laparoscopic surgical video. It answers short questions about foreign objects โ clips, drains, specimen bags, needles, threads โ in a frame.
๐ด Not a medical device. Not for clinical decision-making.
Known weakness, measured rather than assumed: it cannot enumerate reliably past about two objects. On external laparoscopic frames with large metallic instruments, accuracy falls to 0.04 at three objects. That is a perception limit, not a formatting one โ it reproduces across two unrelated answer formats and two different datasets.
Data
Trained only on the challenge's own corpus (ORENA SAVE FOCUS FRAME: `heico-focus-vqa` + `lapchole-focus-vqa`), built on the Heidelberg Colorectal dataset (doi:10.1038/s41597-021-00882-2). No external surgical corpus is in this checkpoint โ several were tried and each closed as a measured negative.
License
CC BY-NC-SA 4.0. The backbone is Apache-2.0, but the fine-tuning data is CC BY-NC-SA and the stricter term governs the derivative. NonCommercial, ShareAlike, attribution required.
Links
Team
`Mexico-Oxford_TEAM`
๐ TODO โ team members. Names, affiliations, ORCIDs and roles โ each member fills in their own. (placeholder: to be completed by the team)
