TencentARC/TimeLens-8B
TimeLens-8B
π **Paper** | π» **Code** | π **Project Page** | π€ **Model & Data**
β¨ Model Description
TimeLens-8B is an MLLM with state-of-the-art video temporal grounding performance among open-source models, finetuned from Qwen3-VL-8B-Instruct. It is trained with carefully crafted RLVR (reinforcement learning with verifiable rewards) recipe proposed in our paper, utilizing our high-quality VTG training dataset TimeLens-100K.
π Performance
TimeLens-8B achieves state-of-the-art video temporal grounding performance among open-source models:
<table> <thead> <tr> <th rowspan="2" align="center">Model</th> <th colspan="4" align="center">Charades-TimeLens</th> <th colspan="4" align="center">ActivityNet-TimeLens</th> <th colspan="4" align="center">QVHighlights-TimeLens</th> </tr> <tr> <th align="center">R1<br>@0.3</th> <th align="center">R1<br>@0.5</th> <th align="center">R1<br>@0.7</th> <th align="center">mIoU</th> <th align="center">R1<br>@0.3</th> <th align="center">R1<br>@0.5</th> <th align="center">R1<br>@0.7</th> <th align="center">mIoU</th> <th align="center">R1<br>@0.3</th> <th align="center">R1<br>@0.5</th> <th align="center">R1<br>@0.7</th> <th align="center">mIoU</th> </tr> </thead> <tbody> <tr> <td><a href="https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct">Qwen2.5-VL-7B-Instruct</a></td> <td align="center">59.7</td> <td align="center">37.8</td> <td align="center">16.6</td> <td align="center">39.3</td> <td align="center">44.1</td> <td align="center">31.0</td> <td align="center">16.1</td> <td align="center">31.4</td> <td align="center">41.5</td> <td align="center">27.8</td> <td align="center">15.2</td> <td align="center">31.6</td> </tr> <tr> <td><a href="https://huggingface.co/TencentARC/TimeLens-7B"><b>TimeLens-7B</b>π</a></td> <td align="center"><b>70.5</b></td> <td align="center"><b>55.6</b></td> <td align="center"><b>28.4</b></td> <td align="center"><b>48.8</b></td> <td align="center"><b>62.8</b></td> <td align="center"><b>51.0</b></td> <td align="center"><b>32.6</b></td> <td align="center"><b>46.2</b></td> <td align="center"><b>74.1</b></td> <td align="center"><b>62.7</b></td> <td align="center"><b>43.1</b></td> <td align="center"><b>56.0</b></td> </tr> <tr> <td><a href="https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct">Qwen3-VL-8B-Instruct</a></td> <td align="center">69.2</td> <td align="center">53.4</td> <td align="center">27.5</td> <td align="center">48.3</td> <td align="center">62.1</td> <td align="center">51.2</td> <td align="center">34.4</td> <td align="center">46.8</td> <td align="center">74.2</td> <td align="center">64.6</td> <td align="center">49.3</td> <td align="center">59.4</td> </tr> <tr> <td><a href="https://huggingface.co/TencentARC/TimeLens-8B"><b>TimeLens-8B</b>π</a></td> <td align="center"><b>76.6</b></td> <td align="center"><b>63.0</b></td> <td align="center"><b>35.2</b></td> <td align="center"><b>55.2</b></td> <td align="center"><b>68.9</b></td> <td align="center"><b>58.4</b></td> <td align="center"><b>40.6</b></td> <td align="center"><b>53.2</b></td> <td align="center"><b>80.2</b></td> <td align="center"><b>71.6</b></td> <td align="center"><b>55.5</b></td> <td align="center"><b>65.5</b></td> </tr> </tbody> </table>
For detailed comparison with other models, please refer to the π Leaderboard.
π Usage
Install the following packages:
pip install transformers==4.57.1 accelerate==1.6.0 torch==2.6.0 torchvision==0.21.0
pip install qwen-vl-utils[decord]==0.0.14
# use Flash-Attention 2 to speed up generation
pip install flash-attn==2.7.4.post1 --no-build-isolation --no-cache-dirUsing π€Transformers for Inference:
import requests
import os
import torch
from transformers import AutoModelForImageTextToText, AutoProcessor
from qwen_vl_utils import process_vision_info
def download_video(url):
save_path = os.path.basename(url)
if not os.path.exists(save_path):
print(f"Downloading video from {url}...")
response = requests.get(url, stream=True)
response.raise_for_status()
with open(save_path, 'wb') as f:
for chunk in response.iter_content(chunk_size=8192):
f.write(chunk)
return save_path
# Load model and processor
model = AutoModelForImageTextToText.from_pretrained(
"TencentARC/TimeLens-8B",
dtype=torch.bfloat16,
attn_implementation="flash_attention_2",
device_map="auto",
)
processor = AutoProcessor.from_pretrained(
"TencentARC/TimeLens-8B",
padding_side="left",
do_resize=False,
)
# Prepare input
query = "A man drinks water with a glass"
video_path = download_video("https://huggingface.co/datasets/JungleGym/TimeLens-Assets/resolve/main/2Y8XQ.mp4")
GROUNDER_PROMPT = "Please find the visual event described by the sentence '{}', determining its starting and ending times. The format should be: 'The event happens in <start time> - <end time> seconds'."
messages = [{
'role': 'user',
'content': [
{
'type': 'video',
'video': video_path,
'min_pixels': 64 * 32 * 32,
'total_pixels': 14336 * 32 * 32,
'fps': 2,
},
{
'type': 'text',
'text': GROUNDER_PROMPT.format(query)
}
]
}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
images, videos, video_kwargs = process_vision_info(
messages,
image_patch_size=16,
return_video_kwargs=True,
return_video_metadata=True,
)
videos, video_metadatas = zip(*videos)
videos, video_metadatas = list(videos), list(video_metadatas)
inputs = processor(
text=[text],
images=images,
videos=videos,
video_metadata=video_metadatas,
padding=True,
return_tensors='pt',
**video_kwargs,
).to("cuda")
output_ids = model.generate(
**inputs,
do_sample=False,
max_new_tokens=512,
)
generated_ids_trimmed = [
out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, output_ids)
]
answer = processor.batch_decode(
generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)[0]
print(f"Answer: {answer}")Citation
If you find our work helpful for your research and applications, please cite our paper:
@article{zhang2025timelens,
title={TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs},
author={Zhang, Jun and Wang, Teng and Ge, Yuying and Ge, Yixiao and Li, Xinhao and Shan, Ying and Wang, Limin},
journal={arXiv preprint arXiv:2512.14698},
year={2025}
}