CoolFace
Datasetpublic

hotchpotch/multilingual_cc_news

hotchpotch/multilingual_cc_news Dataset Summary This dataset republishes multilingual CC-News data in a Hugging Face friendly layout with one subset per language. Source and transformation Original source datasets on the Hugging Face Hub: CloverSearch/cc-news-mutlilingual intfloat/multilingual_cc_news The intfloat version provides a loading script, but it can be difficult to use directly via the datasets library because it pulls raw JSONL files… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/multilingual_cc_news.

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes2.9kdownloads
Dataset Card

hotchpotch/multilingualccnews

Dataset Summary

This dataset republishes multilingual CC-News data in a Hugging Face friendly layout with one subset per language.

Source and transformation

Original source datasets on the Hugging Face Hub:

The intfloat version provides a loading script, but it can be difficult to use directly via the datasets library because it pulls raw JSONL files and relies on a custom builder. This repo republishes the same article-level content as pre-sharded Parquet subsets for easier loading. The text content is not semantically modified; the main change is packaging/layout.

Data Fields

  • title: string
  • maintext: string
  • url: string
  • date_publish: string

How to use this dataset

Each language is a dataset config. Load one language at a time:

python
from datasets import load_dataset

# Single language
train = load_dataset("hotchpotch/multilingual_cc_news", "af", split="train")

# Another language
train_ja = load_dataset("hotchpotch/multilingual_cc_news", "ja", split="train")

License

This dataset follows the license and usage terms of the original CC-News sources. The immediate upstream dataset cards are CloverSearch/cc-news-mutlilingual and intfloat/multilingual_cc_news. Because the content is derived from Common Crawl news data and source news articles, downstream users should also respect the applicable source-site and Common Crawl terms.

References

No dedicated paper is listed by the source dataset cards. For the CC-News dataset announcement, see:

Supported Languages (train split)

LanguageRows
af5,212
als652
am22,672
an23
arz9,806
as31,679
ast338
av41
az573,044
azb61
ba6,915
bar37
bcl34
be72,654
bg2,924,019
bh141
bn719,102
bo486
bpy65
br1,103
bs30,804
bxr165
ca851,043
cbk13
ce39
ceb2,449
ckb31,864
co6
cs2,695,727
cv54
cy37,692
da1,187,189
de2,242,000
diq64
dsb4
dty2
dv6
el6,772,358
eml81
en1,899,000
eo4,929
et1,098,270
eu76,444
fa3,443,176
fi1,536,679
fy39,731
ga1,652
gd10,988
gl113,208
gn140
gom173
gu75,382
gv98
he530,738
hi10,859,572
hif1
hr1,481,087
hsb10
ht3,790
hu2,485,688
hy223,261
ia109
id4,483,457
ie33
ilo222
io936
is196,123
ja4,306,405
jbo59
jv562
ka91,811
kk55,996
km2,630
kn290,071
ko5,572,465
krc134
ku7,566
kv116
kw210
ky100,863
la31,572
lb41,252
lez35
li9
lmo461
lo13,915
lt590,555
lv557,302
mai18
mg442
mhr1,600
min37
mk216,221
ml536,337
mn40,425
mr362,224
mrj6
ms31,690
mt990
mwl7
my71,512
myv10
mzn43
nah2
nap28
nds2,125
ne17,766
new65
nl5,616,536
nn171,271
no1,738,632
oc437
or50,586
os3
pa48,191
pam80
pfl1
pl3,508,134
pms275
pnb12,793
ps40,628
pt10,677,210
qu280
rm8,652
ro6,847,940
ru2,385,000
sa178
sah890
sc32
scn26
sco56
sd6,078
sh235,356
si19,193
sk1,241,434
sl889,313
so15,956
sq367,998
sr1,014,086
su579
sv3,285,173
sw58,793
ta1,937,554
te616,215
tg59,882
th174,151
tk5,600
tl61,195
tt14,760
tyv2
ug18
uk3,892,400
ur1,767,337
uz5,916
vec55
vep53
vi4,386,142
vls3
vo35
wa53
war602
wuu20
xal8
xmf1
yi115
yo5,264
yue16
zh6,133,244