CoolFace
Datasetpublic

OpenLLM-Ro/ro_sft_finepdfs

Dataset Description FinePDFs is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Here we provide the Romanian split of FinePDFs training set, prepared for OCR: pairs of images (pages) and extracted text. This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al.… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/ro_sft_finepdfs.

sourceHugging Faceodc-byupdated 4mo agoView on Hugging Face
1likes162downloads
settings

This repository belongs to OpenLLM-Ro on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namero_sft_finepdfs
visibilitypublic
licenceodc-by
gatedno
ownerOpenLLM-Ro
Account settings
OpenLLM-Ro/ro_sft_finepdfs · CoolFace