CoolFace
Datasetpublic

HumynLabs/French_Documents_Dataset_PDF

French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes1.2kdownloads
Dataset Card

French Documents Dataset (PDF)

This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction.

Contact

For queries or collaborations related to this dataset, contact:

  • anoushka@kgen.io
  • abhishek.vadapalli@kgen.io

Supported Tasks

  • Task Categories:
  • Document Classification
  • OCR and Text Recognition
  • Layout and Structure Analysis
  • Language Modeling for French
  • Supported Tasks:
  • Automatic extraction of French text from PDFs
  • Classification of documents by topic (literary, educational, official, news)
  • OCR for printed French text with diacritics and typographic variation
  • Benchmarking AI models for multilingual and French-language document understanding

Languages

  • Primary Language: French
  • Secondary Presence: English and other European languages (common in academic or bilingual contexts)

Dataset Creation

Curation Rationale

The dataset was curated to improve AI systems' ability to process and understand French-language documents with varied fonts, layouts, and linguistic structures. It is intended for research in multilingual OCR and document intelligence.

Source Data

  • Contributors: Open-access French repositories, public libraries, and volunteer data curators.
  • Collection Process: All documents were sourced from legally accessible, open-licensed PDF archives and public-domain resources.

Other Known Limitations

  • Bias: Primarily formal and academic content; limited representation of informal or regional French
  • Format Variation: Some PDFs may include embedded scanned pages or mixed formats
  • Geographical Bias: Mostly European French; limited coverage of African and Canadian French variants

Intended Uses

✅ Direct Use

  • Training and evaluation of OCR systems for French text
  • Research in document classification and layout analysis
  • Digitization of French-language archives and educational material

❌ Out-of-Scope Use

  • Identification of individuals or private data from PDFs
  • Commercial use of copyrighted works without permission
  • Use in profiling, surveillance, or handwriting analysis

License

CC BY 4.0