HumynLabs/French_Documents_Dataset_PDF
French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.
01.2k
1---2license: cc-by-4.03pretty_name: 'French Documents Dataset (PDF)'4language:5- fr6tags:7- pdf8- french9- document-understanding10- text-recognition11- ocr12- ai-research13- computer-vision14- text-extraction15size_categories:16- 1K<n<10K17---18 19# French Documents Dataset (PDF)20*This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction.*21 22## Contact23For queries or collaborations related to this dataset, contact: 24 - anoushka@kgen.io 25 - abhishek.vadapalli@kgen.io 26 27## Supported Tasks28 29- **Task Categories**: 30 - Document Classification 31 - OCR and Text Recognition 32 - Layout and Structure Analysis 33 - Language Modeling for French 34 35- **Supported Tasks**: 36 - Automatic extraction of French text from PDFs 37 - Classification of documents by topic (literary, educational, official, news) 38 - OCR for printed French text with diacritics and typographic variation 39 - Benchmarking AI models for multilingual and French-language document understanding 40 41## Languages42 43- **Primary Language**: French 44- **Secondary Presence**: English and other European languages (common in academic or bilingual contexts) 45 46## Dataset Creation47 48### Curation Rationale49The dataset was curated to improve AI systems' ability to process and understand French-language documents with varied fonts, layouts, and linguistic structures. It is intended for research in multilingual OCR and document intelligence.50 51### Source Data52- **Contributors**: Open-access French repositories, public libraries, and volunteer data curators. 53- **Collection Process**: All documents were sourced from legally accessible, open-licensed PDF archives and public-domain resources. 54 55### Other Known Limitations56- **Bias**: Primarily formal and academic content; limited representation of informal or regional French 57- **Format Variation**: Some PDFs may include embedded scanned pages or mixed formats 58- **Geographical Bias**: Mostly European French; limited coverage of African and Canadian French variants 59 60## Intended Uses61 62### ✅ Direct Use63- Training and evaluation of OCR systems for French text 64- Research in document classification and layout analysis 65- Digitization of French-language archives and educational material 66 67### ❌ Out-of-Scope Use68- Identification of individuals or private data from PDFs 69- Commercial use of copyrighted works without permission 70- Use in profiling, surveillance, or handwriting analysis 71 72## License73 74CC BY 4.075 