CoolFace
Datasetpublic

1ou2/fr_wiki_paragraphs

Dataset Card for French Wikipedia Text Corpus Dataset Description The French Wikipedia Text Corpus is a comprehensive dataset derived from French Wikipedia articles. It is specifically designed for training language models (LLMs). The dataset contains the text of paragraphs from Wikipedia articles, with sections, footnotes, and titles removed to provide a clean and continuous text stream. Dataset Details Features text: A single attribute containing the full text… See the full description on the dataset page: https://huggingface.co/datasets/1ou2/fr_wiki_paragraphs.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes129downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
1ou2/fr_wiki_paragraphs · CoolFace