datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DRR_dataDeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/DeepVision-103K.DeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/Devilishcode/DeepVision-103K.DeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/blsmash044/DeepVision-103K.DeepVision-103K
🔭 DeepVision-103K
A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoning
Training on DeepVision-103K yields top performance on both multimodal mathematical reasoning and general multimodal benchmarks:
Average Performance on multimodal math and general multimodal benchmarks.
Training on DeepVision-103K elicits more efficient reasoning.
Benchmark
Qwen3-VL-8B-Instruct (Acc / Tokens)
Qwen3-VL-8B-DeepVision (Acc /… See the full description on the dataset page: https://huggingface.co/datasets/JamesGoGo/DeepVision-103K.deepvision_datasetsdeepvision-vlm-predictions
VLM Evaluation Predictions — Zeroshot-DeepVision-24k
Model evaluation predictions and metrics from the Bengali Math VQA pipeline.
Folder structure
Baseline evaluation protocol
Base model output is constrained via a system-level instruction to output only the
Bengali MCQ option letter (ক/খ/গ/ঘ), ensuring format parity with the fine-tuned model.
deepvision-zero-shot-20kdeepvisiontools-demo-datasets
Description
Those are 3 datasets for deepvisiontools library demo. deepvisiontools homepage : https://forge.inrae.fr/ue-apc/librairies/python/deepvisiontools
Datasets
The dice dataset was downloaded from Kaggle : https://www.kaggle.com/datasets/nellbyler/d6-dice
The VegannSubDataset is a small portion from : https://zenodo.org/records/7636408
The coco_6cls_subset was obtained from : https://universe.roboflow.com/nan-ixwz3/coco-y1tdb
deepvisiontools_tutorialsDeepVision_ZS_Predictions
VLM Evaluation Predictions
Prediction outputs from the Zeroshot-DeepVision-24k Bengali Math VQA evaluation pipeline.
Structure
{model_short}/
baseline-testing/ ← predictions BEFORE fine-tuning
post-finetune-testing/ ← predictions AFTER fine-tuning
Generated by the reusable Kaggle evaluation notebook.
deepvision-atlas-data
