CoolFace
Datasetpublic

OpenLLM-France/wikimedia

Dataset Card This dataset is a curated collection of Wikimedia pages in markdown format, compiled from various Wikimedia projects across multiple languages. Covered Wikimedia Projects: wikipedia wikibooks wikinews wikiquote wikisource wikiversity wikivoyage wiktionary Supported Languages: ar (Arabic) br (Breton) ca (Catalan) co (Corsican) de (German) en (English) es (Spanish) eu (Basque) fr (French) frp (Arpitan) it (Italian) nl (Dutch) oc (Occitan) pcd (Picard) pt… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-France/wikimedia.

sourceHugging Facecc-by-sa-4.0updated 1y agoView on Hugging Face
3likes856downloads
Dataset Card

Dataset Card

This dataset is a curated collection of Wikimedia pages in markdown format, compiled from various Wikimedia projects across multiple languages.

Covered Wikimedia Projects:

  • wikipedia
  • wikibooks
  • wikinews
  • wikiquote
  • wikisource
  • wikiversity
  • wikivoyage
  • wiktionary

Supported Languages:

  • ar (Arabic)
  • br (Breton)
  • ca (Catalan)
  • co (Corsican)
  • de (German)
  • en (English)
  • es (Spanish)
  • eu (Basque)
  • fr (French)
  • frp (Arpitan)
  • it (Italian)
  • nl (Dutch)
  • oc (Occitan)
  • pcd (Picard)
  • pt (Portuguese)

Data Source

The content was extracted from the Wikimedia dumps, specifically from the dump dated March 20, 2025 (20250320).

The extraction process follows the same methodology as the one used in the OpenLLM-France/wikipedia.

Dataset Size

The tables below provide a detailed overview of the dataset size, organized by language:

language# pages# words# characters
en (English)16.46 M6.93 B39.97 B
fr (French)9.66 M3.07 B18.00 B
de (German)4.56 M2.21 B14.83 B
es (Spanish)3.06 M1.56 B9.07 B
it (Italian)2.75 M1.48 B8.86 B
nl (Dutch)3.16 M734.36 M4.40 B
pt (Portuguese)1.76 M710.99 M4.06 B
ca (Catalan)1.44 M564.51 M3.33 B
ar (Arabic)1.46 M562.65 M3.22 B
eu (Basque)511.70 K124.81 M882.24 M
br (Breton)148.47 K37.92 M206.85 M
oc (Occitan)160.15 K35.94 M202.87 M
co (Corsican)17.85 K2.64 M15.59 M
pcd (Picard)6.04 K1.59 M8.92 M
frp (Arpitan)5.79 K873.34 K4.97 M
TOTAL45.16 M18.03 B107.07 B

The tables below provide a detailed overview of the dataset size, organized by language and Wikimedia project:

languagedomain# pages# words# characters
arwikipedia1.29 M448.80 M2.60 B
arwikibooks1.14 K1.12 M6.48 M
arwikinews9.55 K3.51 M21.15 M
arwikiquote4.05 K1.44 M8.49 M
arwikisource79.05 K105.56 M569.20 M
arwikiversity943459.03 K2.74 M
arwiktionary71.95 K1.76 M11.96 M
brwikipedia87.91 K19.30 M108.62 M
brwikiquote171194.98 K1.06 M
brwikisource8.28 K14.27 M73.10 M
brwiktionary52.12 K4.16 M24.06 M
cawikipedia808.82 K519.77 M3.04 B
cawikibooks2.91 K1.94 M11.81 M
cawikinews4.88 K2.65 M10.47 M
cawikiquote4.12 K1.62 M9.55 M
cawikisource4.65 K7.99 M43.16 M
cawiktionary619.56 K30.53 M211.58 M
cowikipedia8.32 K2.40 M14.02 M
cowiktionary9.54 K243.04 K1.57 M
dewikipedia3.04 M1.78 B11.79 B
dewikibooks10.75 K12.13 M82.44 M
dewikinews14.32 K4.23 M31.46 M
dewikiquote8.00 K2.29 M14.61 M
dewikisource264.20 K279.29 M1.82 B
dewikiversity47.34 K4.18 M34.88 M
dewikivoyage20.71 K21.79 M154.83 M
dewiktionary1.15 M113.33 M906.85 M
enwikipedia7.14 M5.11 B29.00 B
enwikibooks86.51 K111.09 M661.77 M
enwikinews22.41 K9.83 M58.96 M
enwikiquote58.10 K90.33 M522.59 M
enwikisource617.50 K1.05 B6.02 B
enwikiversity47.58 K35.78 M225.45 M
enwikivoyage34.05 K50.32 M306.01 M
enwiktionary8.46 M472.32 M3.18 B
eswikipedia2.03 M1.38 B7.97 B
eswikibooks8.81 K8.74 M53.48 M
eswikinews12.32 K4.47 M26.57 M
eswikiquote8.67 K5.01 M29.62 M
eswikisource51.76 K93.02 M541.36 M
eswikiversity3.14 K3.61 M23.07 M
eswikivoyage3.40 K8.74 M52.47 M
eswiktionary939.35 K56.36 M374.12 M
euwikipedia452.95 K117.47 M826.58 M
euwikibooks2.20 K564.08 K4.23 M
euwikiquote37090.20 K654.51 K
euwikisource1.24 K3.61 M26.83 M
euwiktionary54.94 K3.08 M23.94 M
frpwikipedia5.79 K873.34 K4.97 M
frwikipedia2.77 M1.89 B10.90 B
frwikibooks23.74 K26.79 M161.96 M
frwikinews24.26 K12.03 M63.53 M
frwikiquote9.80 K4.09 M24.63 M
frwikisource316.03 K844.05 M4.95 B
frwikiversity17.69 K10.26 M70.43 M
frwikivoyage9.45 K11.04 M67.16 M
frwiktionary6.48 M267.14 M1.76 B
itwikipedia1.97 M1.19 B6.96 B
itwikibooks18.24 K40.78 M301.79 M
itwikinews12.23 K4.36 M25.64 M
itwikiquote53.55 K38.63 M234.14 M
itwikisource101.73 K157.49 M1.00 B
itwikiversity5.44 K6.43 M42.38 M
itwikivoyage12.83 K18.00 M115.42 M
itwiktionary575.73 K25.09 M180.32 M
nlwikipedia2.20 M650.71 M3.83 B
nlwikibooks10.39 K5.95 M37.40 M
nlwikinews4.82 K1.94 M12.38 M
nlwikiquote1.26 K567.74 K3.50 M
nlwikisource14.20 K13.69 M84.50 M
nlwikivoyage4.22 K2.12 M13.47 M
nlwiktionary922.49 K59.39 M421.34 M
ocwikipedia87.87 K30.29 M165.44 M
ocwikibooks6722.36 K127.57 K
ocwiktionary72.21 K5.63 M37.30 M
pcdwikipedia6.04 K1.59 M8.92 M
ptwikipedia1.17 M617.13 M3.52 B
ptwikibooks11.74 K7.87 M49.26 M
ptwikinews34.54 K16.16 M91.11 M
ptwikiquote11.36 K3.13 M18.63 M
ptwikisource31.04 K26.91 M158.31 M
ptwikiversity4.43 K3.27 M20.79 M
ptwikivoyage4.04 K3.93 M23.77 M
ptwiktionary496.48 K32.59 M184.28 M

Example use (python)

Load the full dataset:

python
import datasets

ds = datasets.load_dataset("OpenLLM-France/wikimedia",
    streaming=True,
    split="train"
)

Load the dataset for a given language (French in this example):

python
ds = datasets.load_dataset("OpenLLM-France/wikimedia", "fr",
    streaming=True,
    split="train"
)