CoolFace
Datasetpublic

HPLT/HPLT2.0_cleaned

NB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.

sourceHugging Facecc0-1.0updated 3mo agoView on Hugging Face
45likes176kdownloads
Dataset Card

NB: HPLT2.0 is now superseded by a newer release:

[HPLT3.0](https://huggingface.co/datasets/HPLT/HPLT3.0)

We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0.

This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl.

For a detailed description of the dataset, please refer to our website and our pre-print.

The Cleaned variant of HPLT Datasets v2.0

This is the ``cleaned`` variant of the HPLT Datasets v2.0 converted to the Parquet format semi-automatically when being uploaded here. The original JSONL files (which take ~4x fewer disk space than this HF version) and the larger non-cleaned version can be found at https://hplt-project.org/datasets/v2.0.

Dataset Performance

Internal Evaluation

We conducted the FineWeb-style ablation studies within the HPLT project with the focus on one high-resource and one low-resource language: English and Norwegian.

We train 1.7B decoder-only LMs using 100B/30B tokens sampled from the English/Norwegian parts of our HPLT v2 dataset respectively. We replicate the FineWeb corpora comparison design and train the models with a fixed pretraining setup except for the pretraining corpus (English: four corpora; Norwegian: five corpora). Please find the general description of the training and evalutaion setups below and refer to more details in Section 6.2 and Appendix I in our paper.

English ResultsNorwegian Results
<img src="LLMevalenscoresnorm.png" height="700" /><img src= "norevalablationcamera_ready.jpg" height="700" />

English

  • Corpora: HPLT v1.2, FineWeb and HPLT v2 (ours; deduplicated and cleaned versions).
  • Pretraining framework and infrastructure: We trained our English models using Megatron-LM on LUMI with 16 nodes, each with 4 AMD MI250x GPUs with dual-GCD (graphics compute die) design, amounting to 8 logical devices. In total, we used 128 devices and a single 64-core CPU for approximately 84 hours, totalling 11,008 GPU hours per model.
  • Evaluation tasks: ARC (Easy and Challenge), Hellaswag, PIQA, and OpenbookQA. We consider only the 0-shot evaluation regime.
  • Evaluation framework: LightEval.
  • Results: See the plot above. Our models trained on the HPLT v2 datasets reach similar performance to the models trained on FineWeb data and considerably outperform the models trained on HPLT v1.2.

Norwegian

  • Corpora: HPLT v1.2, FineWeb-2, mC4, CulturaX, and HPLT v2 (ours).
  • Pretraining framework and infrastructure: We trained our Norwegian models using Megatron-DeepSpeed on LUMI with 32 nodes, each with 4 AMD MI250x GPUs. The full pretraining run of each model took approximately 15 hours (wall-clock time), or 1,920 GPU-hours.
  • Evaluation tasks: NorCommonsenseQA, NorOpenBookQA, NRK-Quiz-QA, NCB, NorIdiom, and NorQuAD. We discarded tasks that provided a low signal based on the monotonicity and non-random performance criteria defined in the FineWeb-2 evaluation design. The resulting tasks were NCB, NRK-Quiz-QA, NorCommonsenseQA, and NorQuAD. We aggregated the performance using the average normalized score. We consider only the 0-shot evaluation regime.
  • Evaluation framework: NorEval, a Norwegian language understanding and generation evaluation benchmark based upon LM Evaluation Harness.
  • Results: See the plot above. The Norwegian models trained on FineWeb, CulturaX, and mC4 perform on par with HPLT v2 and outperform those trained on HPLT v1.2. Performance gains start to level off after 16B tokens, with the FineWeb and HPLT v2 scores being more stable during pretraining. This suggests that CulturaX, FineWeb, and HPLT v2 are more effective corpora for Norwegian, and their mixtures potentially provide further benefits.
External Evaluation

The HuggingFace team has compared the utility of various multilingual corpora for training large language models in their FineWeb2 initiative.

They found that the HPLT v2 datasets are next to their FineWeb-2, on par with the CulturaX dataset as shown in this figure produced by HuggingFace:

<img src="https://huggingface.co/datasets/HuggingFaceFW/admin/resolve/main/multilingualdatasetscomparison.png" width="800" height="800" />

This is a massive improvement compared to the HPLT v1 datasets, as can be seen on the plot above. In fact, it’s even better: if one looks at the language-specific results, it becomes clear that on Arabic, Hindi, Russian, Thai and Turkish (5 out of 9 languages HuggingFace evaluated on), HPLT v2 is on par or better than FineWeb 2. The average score is lower mostly because of Chinese, we expect it to improve a lot in HPLT v3. Note that the source of the FineWeb 2 (and CulturaX) data is exclusively CommonCrawl, while the HPLT datasets are to a large extent composed of Internet Archive crawls. Thus, FineWeb-2 and HPLT v2 are complementary to each other and should be used together.

Languages

The ``cleaned`` version of HPLT Datasets v2.0 consists of subsets corresponding to 191 language codes. Below we provide a list of language codes. For each language code the amount of text is shown as measured in:

  • segments: the number of sequences of characters (possibly empty) separated by the newline symbol,
  • wcwords: the number of words as defined by the Unix ``wc`` utility, i.e. the number of non-whitespaces with a whitespace or the beginning of document before,
  • chars: the number of characters,
  • docs: the number of documents, each document corresponds to an individual web page from the sourcing web crawls.
langsegmentswcwordscharsdocsLanguage NameISO693-3 codeISO693-3 code macroISO693-1 direct codeISO693-1 through macro
0TOTAL3.00e+115.56e+123.74e+131.06e+10
1ace_Arab1.17e+028.36e+034.97e+041.60e+01Achineseace
2ace_Latn2.06e+058.20e+065.08e+071.29e+04Achineseace
3afr_Latn3.77e+071.00e+095.95e+091.46e+06Afrikaansafrafaf
4als_Latn9.51e+072.71e+091.61e+105.38e+06Tosk Albanianalssqisq
5amh_Ethi7.01e+061.96e+081.03e+092.96e+05Amharicamhamam
6ara_Arab2.20e+094.81e+102.80e+118.27e+07Arabicaraarar
7asm_Beng2.68e+067.34e+074.76e+081.76e+05Assameseasmasas
8ast_Latn7.43e+061.95e+081.24e+092.73e+05Asturianast
9awa_Deva1.32e+056.05e+062.88e+077.28e+03Awadhiawa
10ayr_Latn1.88e+053.07e+062.51e+079.22e+03Central Aymaraayraymay
11azb_Arab2.39e+063.96e+072.60e+086.61e+04South Azerbaijaniazbazeaz
12azj_Latn1.27e+082.57e+091.96e+106.48e+06North Azerbaijaniazjazeaz
13bak_Cyrl3.14e+067.53e+075.58e+081.71e+05Bashkirbakbaba
14bam_Latn9.17e+043.98e+062.07e+075.72e+03Bambarabambmbm
15ban_Latn6.01e+051.13e+077.72e+071.07e+04Balineseban
16bel_Cyrl4.88e+071.21e+098.54e+092.32e+06Belarusianbelbebe
17bem_Latn1.34e+054.52e+063.23e+076.14e+03Bemba (Zambia)bem
18ben_Beng1.76e+084.64e+093.02e+101.10e+07Bengalibenbnbn
19bho_Deva4.58e+051.35e+076.86e+072.86e+04Bhojpuribho
20bjn_Arab1.95e+045.48e+053.32e+061.11e+03Banjarbjnmsams
21bjn_Latn3.66e+058.05e+065.60e+071.88e+04Banjarbjnmsams
22bod_Tibt4.65e+055.78e+062.68e+082.74e+04Tibetanbodbobo
23bos_Latn2.68e+087.26e+094.61e+101.46e+07Bosnianboshbsbsbs
24bug_Latn3.86e+042.70e+061.93e+072.02e+03Buginesebug
25bul_Cyrl6.81e+081.53e+109.69e+102.81e+07Bulgarianbulbgbg
26cat_Latn3.83e+081.00e+106.02e+101.86e+07Catalancatcaca
27ceb_Latn2.86e+068.59e+075.16e+081.39e+05Cebuanoceb
28ces_Latn1.93e+094.21e+102.74e+117.53e+07Czechcescscs
29cjk_Latn3.67e+049.65e+057.43e+061.20e+03Chokwecjk
30ckb_Arab5.23e+061.43e+089.13e+082.74e+05Central Kurdishckbkurku
31crh_Latn1.38e+063.68e+072.81e+081.23e+05Crimean Tatarcrh
32cym_Latn1.56e+074.09e+082.40e+097.58e+05Welshcymcycy
33dan_Latn8.73e+082.12e+101.33e+113.38e+07Danishdandada
34deu_Latn1.11e+102.52e+111.78e+124.82e+08Germandeudede
35dik_Latn3.46e+042.30e+061.15e+072.32e+03Southwestern Dinkadikdin
36dyu_Latn2.46e+041.19e+065.55e+061.39e+03Dyuladyu
37dzo_Tibt4.00e+044.22e+057.38e+061.63e+03Dzongkhadzodzdz
38ell_Grek1.85e+094.27e+102.84e+117.03e+07Modern Greek (1453-)ellelel
39eng_Latn1.16e+112.86e+121.71e+134.39e+09Englishengenen
40epo_Latn2.04e+074.72e+082.98e+098.19e+05Esperantoepoeoeo
41est_Latn2.64e+084.74e+093.60e+108.45e+06Estonianestetet
42eus_Latn3.76e+077.77e+086.05e+091.97e+06Basqueeuseueu
43ewe_Latn1.43e+054.31e+062.13e+073.77e+03Eweeweeeee
44fao_Latn4.53e+069.34e+075.82e+082.40e+05Faroesefaofofo
45fij_Latn1.79e+057.26e+063.77e+078.91e+03Fijianfijfjfj
46fin_Latn9.77e+081.84e+101.56e+113.48e+07Finnishfinfifi
47fon_Latn1.48e+041.23e+065.34e+061.23e+03Fonfon
48fra_Latn1.06e+102.37e+111.46e+124.02e+08Frenchfrafrfr
49fur_Latn7.30e+052.08e+071.15e+083.67e+04Friulianfur
50fuv_Latn1.34e+055.14e+062.99e+077.76e+03Nigerian Fulfuldefuvfulff
51gaz_Latn9.74e+052.89e+072.19e+084.91e+04West Central Oromogazormom
52gla_Latn3.31e+068.07e+074.84e+081.37e+05Scottish Gaelicglagdgd
53gle_Latn1.10e+072.96e+081.75e+094.91e+05Irishglegaga
54glg_Latn6.12e+071.64e+091.01e+103.02e+06Galicianglgglgl
55grn_Latn1.71e+063.07e+072.19e+087.34e+04Guaranigrngngn
56guj_Gujr2.06e+075.77e+083.39e+091.13e+06Gujaratigujgugu
57hat_Latn4.64e+061.22e+086.39e+082.13e+05Haitianhaththt
58hau_Latn5.69e+061.53e+088.54e+083.16e+05Hausahauhaha
59heb_Hebr4.67e+089.97e+095.68e+101.71e+07Hebrewhebhehe
60hin_Deva2.67e+088.64e+094.40e+101.36e+07Hindihinhihi
61hne_Deva5.50e+042.20e+061.06e+072.81e+03Chhattisgarhihne
62hrv_Latn2.97e+087.31e+094.80e+101.23e+07Croatianhrvhbshrhr
63hun_Latn1.42e+093.05e+102.25e+115.19e+07Hungarianhunhuhu
64hye_Armn6.52e+071.40e+091.07e+103.60e+06Armenianhyehyhy
65ibo_Latn1.41e+063.83e+072.05e+085.63e+04Igboiboigig
66ilo_Latn1.12e+062.48e+071.57e+084.88e+04Ilokoilo
67ind_Latn2.39e+095.46e+103.84e+119.81e+07Indonesianindmsaidid
68isl_Latn6.96e+071.54e+099.59e+092.84e+06Icelandicislisis
69ita_Latn5.13e+091.27e+118.21e+112.22e+08Italianitaitit
70jav_Latn6.43e+061.38e+089.38e+081.96e+05Javanesejavjvjv
71jpn_Jpan2.33e+104.24e+109.01e+114.18e+08Japanesejpnjaja
72kab_Latn3.45e+059.22e+065.42e+071.51e+04Kabylekab
73kac_Latn1.59e+055.96e+062.84e+077.59e+03Kachinkac
74kam_Latn1.43e+046.74e+054.64e+061.18e+03Kamba (Kenya)kam
75kan_Knda2.49e+075.33e+084.30e+091.34e+06Kannadakanknkn
76kas_Arab2.71e+046.78e+053.47e+069.49e+02Kashmirikasksks
77kas_Deva1.36e+033.19e+041.85e+051.06e+02Kashmirikasksks
78kat_Geor6.37e+071.24e+091.02e+103.34e+06Georgiankatkaka
79kaz_Cyrl8.10e+071.41e+091.11e+102.64e+06Kazakhkazkkkk
80kbp_Latn4.68e+044.26e+062.09e+077.08e+03Kabiyèkbp
81kea_Latn4.39e+041.14e+066.14e+061.96e+03Kabuverdianukea
82khk_Cyrl5.35e+071.34e+099.33e+092.12e+06Halh Mongoliankhkmonmn
83khm_Khmr9.86e+061.14e+082.12e+097.01e+05Khmerkhmkmkm
84kik_Latn5.19e+041.43e+069.29e+064.00e+03Kikuyukikkiki
85kin_Latn1.92e+065.07e+073.67e+089.27e+04Kinyarwandakinrwrw
86kir_Cyrl1.00e+072.47e+081.92e+096.76e+05Kirghizkirkyky
87kmb_Latn1.18e+043.83e+052.07e+065.31e+02Kimbundukmb
88kmr_Latn7.15e+061.96e+081.12e+093.64e+05Northern Kurdishkmrkurku
89knc_Arab1.08e+042.62e+051.30e+062.45e+02Central Kanuriknckaukr
90knc_Latn1.05e+042.41e+061.20e+072.47e+03Central Kanuriknckaukr
91kon_Latn4.75e+041.94e+061.13e+072.54e+03Kongokonkgkg
92kor_Hang1.36e+091.97e+108.92e+103.89e+07Koreankorkoko
93lao_Laoo3.20e+055.18e+068.47e+072.95e+04Laolaololo
94lij_Latn1.58e+055.59e+063.15e+078.37e+03Ligurianlij
95lim_Latn7.14e+061.81e+081.12e+093.68e+05Limburganlimlili
96lin_Latn2.00e+055.56e+063.29e+077.59e+03Lingalalinlnln
97lit_Latn3.22e+086.68e+095.04e+101.33e+07Lithuanianlitltlt
98lmo_Latn2.12e+065.96e+073.45e+081.46e+05Lombardlmo
99ltg_Latn1.51e+053.79e+062.69e+079.21e+03Latgalianltglavlv
100ltz_Latn5.06e+061.07e+087.10e+082.47e+05Luxembourgishltzlblb
101lua_Latn3.87e+041.37e+069.00e+061.08e+03Luba-Lulualua
102lug_Latn4.08e+059.18e+066.80e+072.13e+04Gandaluglglg
103luo_Latn8.41e+043.73e+062.03e+074.15e+03Luo (Kenya and Tanzania)luo
104lus_Latn3.43e+061.25e+086.52e+081.60e+05Lushailus
105lvs_Latn1.74e+083.46e+092.52e+106.77e+06Standard Latvianlvslavlv
106mag_Deva1.93e+048.91e+054.28e+063.28e+02Magahimag
107mai_Deva6.46e+051.78e+079.67e+072.50e+04Maithilimai
108mal_Mlym4.80e+079.74e+089.49e+093.10e+06Malayalammalmlml
109mar_Deva3.63e+079.81e+086.62e+092.08e+06Marathimarmrmr
110min_Latn6.01e+051.10e+077.48e+072.50e+04Minangkabauminmsams
111mkd_Cyrl5.70e+071.48e+099.44e+093.57e+06Macedonianmkdmkmk
112mlt_Latn8.68e+061.96e+081.44e+093.67e+05Maltesemltmtmt
113mni_Beng6.58e+041.63e+061.18e+072.93e+03Manipurimni
114mos_Latn1.91e+048.08e+053.86e+069.31e+02Mossimos
115mri_Latn2.80e+068.68e+074.24e+081.08e+05Maorimrimimi
116mya_Mymr3.05e+074.53e+085.82e+091.37e+06Burmesemyamymy
117nld_Latn3.08e+097.14e+104.51e+111.39e+08Dutchnldnlnl
118nno_Latn3.46e+078.60e+085.40e+091.42e+06Norwegian Nynorsknnonornnnn
119nob_Latn6.76e+082.15e+101.33e+112.70e+07Norwegian Bokmålnobnornbnb
120npi_Deva3.71e+071.13e+097.26e+092.78e+06Nepali (individual language)npinepne
121nso_Latn1.43e+055.32e+062.75e+076.07e+03Pedinso
122nus_Latn8.51e+033.93e+051.88e+062.72e+02Nuernus
123nya_Latn1.34e+062.71e+072.03e+085.31e+04Nyanjanyanyny
124oci_Latn4.20e+061.03e+086.35e+081.90e+05Occitan (post 1500)ociococ
125ory_Orya3.60e+061.20e+087.82e+084.13e+05Odiaoryorior
126pag_Latn8.58e+045.66e+063.35e+076.90e+03Pangasinanpag
127pan_Guru1.17e+073.72e+081.90e+095.85e+05Panjabipanpapa
128pap_Latn1.39e+064.67e+072.54e+088.98e+04Papiamentopap
129pbt_Arab8.46e+062.79e+081.30e+094.66e+05Southern Pashtopbtpusps
130pes_Arab3.96e+098.86e+104.55e+119.05e+07Iranian Persianpesfasfa
131plt_Latn4.74e+061.17e+088.10e+082.08e+05Plateau Malagasypltmlgmg
132pol_Latn4.46e+098.95e+106.32e+111.75e+08Polishpolplpl
133por_Latn6.12e+091.46e+118.96e+112.38e+08Portugueseporptpt
134prs_Arab6.90e+071.84e+099.57e+092.84e+06Dariprsfasfa
135quy_Latn4.94e+051.73e+071.43e+083.69e+04Ayacucho Quechuaquyquequ
136ron_Latn1.70e+094.00e+102.51e+116.59e+07Romanianronroro
137run_Latn1.75e+064.44e+073.16e+081.37e+05Rundirunrnrn
138rus_Cyrl2.63e+105.41e+113.91e+128.85e+08Russianrusruru
139sag_Latn5.19e+043.61e+061.67e+073.16e+03Sangosagsgsg
140san_Deva3.28e+064.38e+073.59e+085.49e+04Sanskritsansasa
141sat_Olck4.58e+041.08e+066.27e+062.57e+03Santalisat
142scn_Latn1.65e+064.24e+072.52e+088.20e+04Sicilianscn
143shn_Mymr9.21e+041.65e+062.12e+076.00e+03Shanshn
144sin_Sinh3.37e+077.96e+084.98e+091.15e+06Sinhalasinsisi
145slk_Latn4.94e+081.06e+107.04e+102.18e+07Slovakslksksk
146slv_Latn2.39e+085.44e+093.53e+101.03e+07Slovenianslvslsl
147smo_Latn1.01e+063.71e+071.86e+084.59e+04Samoansmosmsm
148sna_Latn1.20e+062.39e+071.93e+086.11e+04Shonasnasnsn
149snd_Arab2.83e+068.95e+074.29e+081.00e+05Sindhisndsdsd
150som_Latn1.64e+073.89e+082.56e+099.66e+05Somalisomsoso
151sot_Latn1.08e+063.10e+071.72e+084.39e+04Southern Sothosotstst
152spa_Latn1.21e+103.22e+111.95e+125.03e+08Spanishspaeses
153srd_Latn9.17e+052.39e+071.49e+085.38e+04Sardiniansrdscsc
154srp_Cyrl9.38e+072.52e+091.62e+104.12e+06Serbiansrphbssrsr
155ssw_Latn6.21e+049.94e+058.82e+062.04e+03Swatisswssss
156sun_Latn3.24e+066.96e+074.75e+081.15e+05Sundanesesunsusu
157swe_Latn1.76e+094.01e+102.51e+116.68e+07Swedishswesvsv
158swh_Latn3.43e+077.18e+084.66e+091.37e+06Swahili (individual language)swhswasw
159szl_Latn6.37e+051.47e+071.04e+084.09e+04Silesianszl
160tam_Taml1.69e+082.98e+092.62e+106.11e+06Tamiltamtata
161taq_Latn1.39e+041.54e+068.84e+061.75e+03Tamasheqtaqtmh
162tat_Cyrl1.34e+072.97e+082.16e+096.31e+05Tatartattttt
163tel_Telu3.92e+078.35e+086.50e+092.06e+06Teluguteltete
164tgk_Cyrl2.48e+076.25e+084.59e+091.26e+06Tajiktgktgtg
165tgl_Latn5.29e+071.35e+098.13e+091.87e+06Tagalogtgltltl
166tha_Thai3.39e+083.51e+096.00e+101.77e+07Thaithathth
167tir_Ethi1.13e+063.67e+071.82e+086.47e+04Tigrinyatirtiti
168tpi_Latn2.82e+051.25e+076.45e+071.40e+04Tok Pisintpi
169tsn_Latn1.32e+055.27e+062.77e+076.05e+03Tswanatsntntn
170tso_Latn2.21e+058.67e+064.93e+071.10e+04Tsongatsotsts
171tuk_Latn3.36e+067.07e+075.70e+081.71e+05Turkmentuktktk
172tum_Latn9.90e+042.88e+062.11e+074.38e+03Tumbukatum
173tur_Latn2.58e+095.17e+103.90e+111.17e+08Turkishturtrtr
174twi_Latn1.26e+054.70e+062.42e+075.86e+03Twitwiakatwtw
175uig_Arab8.98e+062.24e+081.75e+094.42e+05Uighuruigugug
176ukr_Cyrl1.17e+092.52e+101.83e+114.74e+07Ukrainianukrukuk
177umb_Latn5.99e+042.43e+061.54e+072.47e+03Umbunduumb
178urd_Arab5.06e+072.13e+091.00e+103.19e+06Urduurdurur
179uzn_Latn1.48e+073.51e+082.85e+097.07e+05Northern Uzbekuznuzbuz
180vec_Latn1.58e+063.53e+072.18e+088.48e+04Venetianvec
181vie_Latn3.02e+098.32e+103.80e+111.01e+08Vietnamesevievivi
182war_Latn2.01e+055.89e+063.56e+071.39e+04Waray (Philippines)war
183wol_Latn1.62e+055.46e+062.75e+075.68e+03Wolofwolwowo
184xho_Latn1.82e+063.03e+072.59e+086.31e+04Xhosaxhoxhxh
185ydd_Hebr2.94e+067.75e+074.58e+081.28e+05Eastern Yiddishyddyidyi
186yor_Latn1.47e+064.28e+072.18e+086.61e+04Yorubayoryoyo
187yue_Hant1.24e+063.27e+067.43e+076.13e+04Yue Chineseyuezhozh
188zho_Hans4.24e+107.40e+102.35e+121.25e+09Chinesezhozhzh
189zho_Hant4.48e+099.51e+092.87e+111.57e+08Chinesezhozhzh
190zsm_Latn5.80e+081.15e+107.84e+101.84e+07Standard Malayzsmmsams
191zul_Latn2.71e+064.44e+073.81e+081.14e+05Zuluzulzuzu

Terms of Use and Takedown

Terms of Use

These data are released under these Terms of Use:

Notice and take down policy

Notice: Should you consider that our data contains material that is owned by you and should therefore not be reproduced here, please:

  • Clearly identify yourself, with detailed contact data such as an address, telephone number or email address at which you can be contacted.
  • Clearly identify the work claimed to be infringed.
  • Clearly identify the material that is claimed to be infringing and information reasonably sufficient to allow us to locate the material.
  • You can reach us at hplt-datasets@ufal.mff.cuni.cz

Take down: We will comply to legitimate requests by removing the affected sources from the next release of the corpora.

  • It is your responsibility that any use of the data complies with any applicable legal framework, such as, among others, the EU Copyright Directive 2019/790 and the General Data Protection Regulation 2018, as amended.

Cite us

@inproceedings{burchell-etal-2025-expanded,
    title = "An Expanded Massive Multilingual Dataset for High-Performance Language Technologies ({HPLT})",
    author = {Burchell, Laurie  and
      de Gibert, Ona  and
      Arefyev, Nikolay  and
      Aulamo, Mikko  and
      Ba{\~n}{\'o}n, Marta  and
      Chen, Pinzhen  and
      Fedorova, Mariia  and
      Guillou, Liane  and
      Haddow, Barry  and
      Haji{\v{c}}, Jan  and
      Helcl, Jind{\v{r}}ich  and
      Henriksson, Erik  and
      Klimaszewski, Mateusz  and
      Komulainen, Ville  and
      Kutuzov, Andrey  and
      Kyt{\"o}niemi, Joona  and
      Laippala, Veronika  and
      M{\ae}hlum, Petter  and
      Malik, Bhavitvya  and
      Mehryary, Farrokh  and
      Mikhailov, Vladislav  and
      Moghe, Nikita  and
      Myntti, Amanda  and
      O{'}Brien, Dayy{\'a}n  and
      Oepen, Stephan  and
      Pal, Proyag  and
      Piha, Jousia  and
      Pyysalo, Sampo  and
      Ram{\'i}rez-S{\'a}nchez, Gema  and
      Samuel, David  and
      Stepachev, Pavel  and
      Tiedemann, J{\"o}rg  and
      Vari{\v{s}}, Du{\v{s}}an  and
      Vojt{\v{e}}chov{\'a}, Tereza  and
      Zaragoza-Bernabeu, Jaume},
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.854/",
    doi = "10.18653/v1/2025.acl-long.854",
    pages = "17452--17485",
    ISBN = "979-8-89176-251-0",
    abstract = "Training state-of-the-art large language models requires vast amounts of clean and diverse textual data. However, building suitable multilingual datasets remains a challenge. In this work, we present HPLT v2, a collection of high-quality multilingual monolingual and parallel corpora, extending prior work of the HPLT project. The monolingual portion of the data contains 8T tokens covering 193 languages, while the parallel data contains 380M sentence pairs covering 51 languages. We document the entire data pipeline and release the code to reproduce it. We provide extensive analysis of the quality and characteristics of our data. Finally, we evaluate the performance of language models and machine translation systems trained on HPLT v2, demonstrating its value."
}
HPLT/HPLT2.0_cleaned · CoolFace