CoolFace
Datasetpublic

HumynLabs/French_Documents_Dataset_PDF

French Documents Dataset (PDF) This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction. Contact For queries or collaborations related to this dataset, contact: anoushka@kgen.io abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/French_Documents_Dataset_PDF.

sourceHugging Facecc-by-4.0updated 11mo agoView on Hugging Face
0likes1.2kdownloads
README.md75 linesDownload Raw Back to root
1---2license: cc-by-4.03pretty_name: 'French Documents Dataset (PDF)'4language:5- fr6tags:7- pdf8- french9- document-understanding10- text-recognition11- ocr12- ai-research13- computer-vision14- text-extraction15size_categories:16- 1K<n<10K17---18 19# French Documents Dataset (PDF)20*This dataset contains a curated collection of French-language documents in PDF format. It includes educational materials, books, news articles, government publications, and public-domain literature written in French. The dataset supports AI research in OCR, document understanding, and multilingual text extraction.*21 22## Contact23For queries or collaborations related to this dataset, contact:  24  - anoushka@kgen.io  25  - abhishek.vadapalli@kgen.io  26 27## Supported Tasks28 29- **Task Categories**:  30  - Document Classification  31  - OCR and Text Recognition  32  - Layout and Structure Analysis  33  - Language Modeling for French  34 35- **Supported Tasks**:  36  - Automatic extraction of French text from PDFs  37  - Classification of documents by topic (literary, educational, official, news)  38  - OCR for printed French text with diacritics and typographic variation  39  - Benchmarking AI models for multilingual and French-language document understanding  40 41## Languages42 43- **Primary Language**: French  44- **Secondary Presence**: English and other European languages (common in academic or bilingual contexts)  45 46## Dataset Creation47 48### Curation Rationale49The dataset was curated to improve AI systems' ability to process and understand French-language documents with varied fonts, layouts, and linguistic structures. It is intended for research in multilingual OCR and document intelligence.50 51### Source Data52- **Contributors**: Open-access French repositories, public libraries, and volunteer data curators.  53- **Collection Process**: All documents were sourced from legally accessible, open-licensed PDF archives and public-domain resources.  54 55### Other Known Limitations56- **Bias**: Primarily formal and academic content; limited representation of informal or regional French  57- **Format Variation**: Some PDFs may include embedded scanned pages or mixed formats  58- **Geographical Bias**: Mostly European French; limited coverage of African and Canadian French variants  59 60## Intended Uses61 62### ✅ Direct Use63- Training and evaluation of OCR systems for French text  64- Research in document classification and layout analysis  65- Digitization of French-language archives and educational material  66 67### ❌ Out-of-Scope Use68- Identification of individuals or private data from PDFs  69- Commercial use of copyrighted works without permission  70- Use in profiling, surveillance, or handwriting analysis  71 72## License73 74CC BY 4.075