CoolFace
Datasetpublic

1ou2/fr_wiki_paragraphs

Dataset Card for French Wikipedia Text Corpus Dataset Description The French Wikipedia Text Corpus is a comprehensive dataset derived from French Wikipedia articles. It is specifically designed for training language models (LLMs). The dataset contains the text of paragraphs from Wikipedia articles, with sections, footnotes, and titles removed to provide a clean and continuous text stream. Dataset Details Features text: A single attribute containing the full text… See the full description on the dataset page: https://huggingface.co/datasets/1ou2/fr_wiki_paragraphs.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes129downloads
5 commits on main
78d338c1y ago

added description in card

1ou2
a2404ed1y ago

Created card

1ou2
0652d681y ago

Update README.md

1ou2
624f79b1y ago

Upload dataset

1ou2
b621fb41y ago

initial commit

1ou2