w4
Datasets
All datasets matching “w4”train-vlm
train-vlm
A fine-tuning corpus for Nepali (Devanagari) government-document OCR and
layout analysis, built to adapt a vision-language model (dots.ocr) to Nepali
paperwork without erasing what it already knows.
About half the corpus by tokens is the target task: synthetic and real-layout
Nepali pages. The other half is replay data from general VQA, English
document QA, and handwritten math, mixed in on purpose. The split follows
Fine-tuning VLMs Without Forgetting. The failure… See the full description on the dataset page: https://huggingface.co/datasets/W4ashabii/train-vlm.filteredraw-mergedfiltered-embedmmu_vipers_w4
mmu_vipers_w4 HATS Catalog Collection
This is the collection of HATS catalogs representing mmu_vipers_w4.
This dataset is part of the Multimodal Universe,
a large-scale collection of multimodal astronomical data. For full details, see the paper:
The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data.
Access the catalog
We recommend the use of the LSDB Python framework to access HATS catalogs.
LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/LSDB/mmu_vipers_w4.Document_type
