CoolFace
20 results

w4

W4ashabii /train-vlmgated train-vlm A fine-tuning corpus for Nepali (Devanagari) government-document OCR and layout analysis, built to adapt a vision-language model (dots.ocr) to Nepali paperwork without erasing what it already knows. About half the corpus by tokens is the target task: synthetic and real-layout Nepali pages. The other half is replay data from general VQA, English document QA, and handwritten math, mixed in on purpose. The split follows Fine-tuning VLMs Without Forgetting. The failure… See the full description on the dataset page: https://huggingface.co/datasets/W4ashabii/train-vlm.tabularimage-to-text1M<n<10M0 likes415 downloads14h agoHugging Facew4nn4b3M4ST3R /filteredtabular100K<n<1M0 likes306 downloads2mo agoHugging Facew4nn4b3M4ST3R /raw-mergedtabular100K<n<1M0 likes288 downloads2mo agoHugging Facew4nn4b3M4ST3R /filtered-embedtabular100K<n<1M0 likes212 downloads2mo agoHugging FaceLSDB /mmu_vipers_w4 mmu_vipers_w4 HATS Catalog Collection This is the collection of HATS catalogs representing mmu_vipers_w4. This dataset is part of the Multimodal Universe, a large-scale collection of multimodal astronomical data. For full details, see the paper: The Multimodal Universe: Enabling Large-Scale Machine Learning with 100TBs of Astronomical Scientific Data. Access the catalog We recommend the use of the LSDB Python framework to access HATS catalogs. LSDB can be… See the full description on the dataset page: https://huggingface.co/datasets/LSDB/mmu_vipers_w4.tabular10K<n<100K0 likes188 downloads4mo agoHugging FaceW4ashabii /Document_typegated0 likes178 downloads29d agoHugging Face