maveryn/trace-qwen2.5-vl-7b
TRACE Qwen2.5-VL 7B
TRACE Qwen2.5-VL 7B is a GRPO checkpoint derived from `Qwen/Qwen2.5-VL-7B-Instruct` and trained on 64,000 grounded visual-reasoning examples spanning 1,000 tasks from `maveryn/trace`. This repository contains the merged step-500 inference checkpoint.
Paper · Project page · GitHub · Collection · Training configuration · Evaluation suite · Run artifacts
Training and provenance
The released training profile reads prompt_answer, scores answer_gt, and does not use the advisory trace_supervision_mode column. The consumed fields, embedded image bytes, and row order in the current dataset release were verified identical to the original training input in the public equivalence receipt.
The canonical output ends with {"answer": ...}. Exact hashes, source revision, and run provenance are recorded in `trace_training_provenance.json` and `.trace_model_revision.json`. The repository does not include optimizer, scheduler, trainer, or FSDP state for continuation.
Evaluation
trace_eval_v1 evaluates 24 external benchmarks and 32,805 examples per model and decoding seed. Scores below are the unweighted macro mean of the 24 benchmark percentages, reported as mean ± sample standard deviation across seeds 42, 43, and 44.
Per-benchmark scores, model revisions, aggregation, and benchmark provenance are available in the public `results.json` and evaluation documentation.
Usage
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
model_id = "maveryn/trace-qwen2.5-vl-7b"
revision = "4d0f1ae8ee25022058090dbdbff61957ece7331d"
image_url = "https://raw.githubusercontent.com/maveryn/trace/main/docs/assets/paper-domain-montage/trace-paper-domain-montage.png"
processor = AutoProcessor.from_pretrained(model_id, revision=revision)
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
model_id,
revision=revision,
dtype="auto",
device_map="auto",
)
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": image_url},
{"type": "text", "text": "How many visual domains are shown with example images? Respond with only a JSON object using the key \"answer\"."},
],
}
]
inputs = processor.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
output_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids = [
output[len(prompt) :] for prompt, output in zip(inputs.input_ids, output_ids)
]
print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0])Verified reference answer
For the published montage used above, the expected output is:
{"answer": 11}The value is backed by the committed montage manifest (layout.panel_count). This is reference ground truth, not a claim about a particular decoding run; publish qualitative model outputs only with their recorded inference settings.
The pinned revision is the checkpoint used by the canonical evaluation; later repository heads may update documentation without changing the model weights.
Intended use and limitations
This checkpoint is intended for research on multimodal reasoning and verifiable-reward post-training. It can produce incorrect answers, unsupported reasoning, or unreliable grounding. It was trained on synthetic tasks and has not been validated for safety-critical, medical, legal, or autonomous decision-making uses. Results depend on the recorded prompts, decoding, parsers, and scorers.
License
This checkpoint is released under Apache-2.0, consistent with the upstream base model.
