Qwen2.5-VL
Qwen2.5-VL-7B-Instructe621-2024-webp-3.4M-recap-qwen2.5-vl-3B
RareConcepts/e621-2024-webp-3.4M
a derived, soft-filtered mirror of the original e621 2024 dump, kept in .webp with machine-generated captions for compact storage and high throughput training
tl;dr
based on: boxingscorpionbagel/e621-2024 (via the processed mirror NebulaeWis/e621-2024-webp-4Mpixel)
this repo: image shards in .webp (≈3.4M images implied by name)
captions: caption field
soft filter: captions produced by Qwen2.5-VL-3B-Instruct (permissive; may include some… See the full description on the dataset page: https://huggingface.co/datasets/RareConcepts/e621-2024-webp-3.4M-recap-qwen2.5-vl-3B.Qwen2.5-7B-Instruct-vllm-20251128_042753e621-2024-metadata-4M-recap-qwen2.5-vl-7B
Caption Dataset
This dataset contains 4,068,421 captioned items exported from CaptionFlow.
Dataset Structure
Data Fields
caption_count
captions: List of captions/outputs
chunk_id
contributor_id
dataset
file_size
filename
image_format
image_height
image_width
item_index
item_key
job_id
metadata
processing_time_ms
quality_scores
shard
timestamp
url
Qwen2.5-7B-Instruct-vllm-retriever-20251202_093826ContextRL_Multimodal_Qwen2.5_VL
ContextRL-Multimodal-Qwen2.5-VL
The multimodal training set for ContextRL, used to train
ContextRL-Qwen2.5-VL-7B, from
the paper Context-Aware RL for Agentic and Multimodal LLMs. It is formatted for the
Qwen2.5-VL chat template.
Setup
Training and evaluation code, data construction pipelines, and detailed configurations are
available in the repository:
👉 https://github.com/xupy2003/ContextAwareRL
