vkehfdl1/banana-vidorev3-synthetic-arms
Banana ViDoRe v3 Synthetic Arms Domain-separated ViDoRe v3 synthetic training arms for finance and industrial adaptation. The Hub dataset uses finance and industrial as dataset configs/subsets. Within each config, splits separate vlm_in_batch, vlm_ocr_bm25, banana_fullpipe, and hybrid_vlm_ocr_bm25_banana_fullpipe. Generated at: 2026-06-29T11:49:45.670912+00:00 Total JSONL rows across configs/splits: 151691. Images are stored once per subset under… See the full description on the dataset page: https://huggingface.co/datasets/vkehfdl1/banana-vidorev3-synthetic-arms.
Banana ViDoRe v3 Synthetic Arms
Domain-separated ViDoRe v3 synthetic training arms for finance and industrial adaptation. The Hub dataset uses finance and industrial as dataset configs/subsets. Within each config, splits separate vlm_in_batch, vlm_ocr_bm25, banana_fullpipe, and hybrid_vlm_ocr_bm25_banana_fullpipe.
Generated at: 2026-06-29T11:49:45.670912+00:00
Total JSONL rows across configs/splits: 151691.
Images are stored once per subset under vkehfdl1/banana-vidorev3-synthetic-arms/<subset>/images/; all split JSONL files use paths relative to their subset directory.
Splits
Training
Use the repo file path with the training CLI; it downloads the matching JSONL plus subset images into the HF cache.
python -m banana.training.scripts.train \
--model colqwen \
--model-name vidore/colqwen2-v1.0 \
--data-path hf://datasets/vkehfdl1/banana-vidorev3-synthetic-arms/industrial/hybrid_vlm_ocr_bm25_banana_fullpipe.jsonl \
--output-dir output/training/industrial-hybrid \
--use-hard-negatives