datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
doc_calibration_datasetViet-Doc-VQA-II-flash2
Dataset Overview
This dataset is a continuation of the ongoing work from Viet Document VAQ dataset was collected from 64,765 pages of Vietnamese 🇻🇳 textbooks( Sách bài tập, chuyên đề, sách giáo án của Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 388,277 detailed… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-II-flash2.mdpbench-doc-ocr-sft
mdpbench-doc-ocr-sft
Qwen 4B (Qwen3-VL) MDPBench 한국어/일본어 문서 OCR/파싱 SFT 학습 데이터(공개 가능분).
재현 코드: https://github.com/sionic-ai/qwen-mdpbench-ocr
Configs
config
rows
내용
상태
ko
2,407
한국 지자체 소식지(신문형) 이미지 + Markdown GT
공개
jp
—
일본어 (추후)
예정
스키마
image (Image, bytes 임베드) — 문서 페이지
markdown (string) — GT 전사(Markdown, 표/읽기순서 보존)
source (string) — snvision / seocho / gongju
lang (string) — ko
doc_type (string) — newspaper
출처 & 라이선스… See the full description on the dataset page: https://huggingface.co/datasets/sionic-ai/mdpbench-doc-ocr-sft.Viet-Doc-VQA-flash2
Dataset Overview
The Document VAQ dataset was collected from 51,856 pages of Vietnamese 🇻🇳 textbooks( Sách Bộ GDĐT, Cánh Diều, Chân trời sáng tạo, Kết nối tri thức), spanning all subjects from grades 1 to 12. Each page has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 310,952 detailed descriptions and query-based questions and answers generated by the Gemini 1.5 Flash model, currently… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-Doc-VQA-flash2.Viet-Doc-VQA-verIII
Cite
@misc{doan2024vintern1befficientmultimodallarge,
title={Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese},
author={Khang T. Doan and Bao G. Huynh and Dung T. Hoang and Thuc D. Pham and Nhat H. Pham and Quan T. M. Nguyen and Bang Q. Vo and Suong N. Hoang},
year={2024},
eprint={2408.12480},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2408.12480},
}
