CoolFace
Datasetpublicgated

VillanovaAI/finepdfs-sample-75k-meta

finepdfs-sample-75k-meta Overview This dataset is a metadata-enriched multilingual sample of the original FinePDFs dataset. FinePDFs is a large-scale collection of document-level texts extracted from PDF files, sourced primarily from Common Crawl. The dataset emphasizes high-quality document extraction, structural coherence, and large-scale coverage of technical, scientific, educational, and administrative content commonly distributed in PDF form. This release… See the full description on the dataset page: https://huggingface.co/datasets/VillanovaAI/finepdfs-sample-75k-meta.

sourceHugging Faceodc-byupdated 9mo agoView on Hugging Face
0likes7downloads
settings

This repository belongs to VillanovaAI on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namefinepdfs-sample-75k-meta
visibilitypublic
licenceodc-by
gatedyes
ownerVillanovaAI
Account settings
VillanovaAI/finepdfs-sample-75k-meta · CoolFace