CoolFace
Datasetpublic

data-is-better-together/fineweb-c

FineWeb-C: Educational content in many languages, labelled by the community Multilingual data is better together! Note: We are not actively working on this project anymore. You can continue to contribute annotations and we'll occasionally refresh the exported data. What is this? FineWeb-C is a collaborative, community-driven project that expands upon the FineWeb2 dataset. The goal is to create high-quality educational content annotations across… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/fineweb-c.

sourceHugging Faceupdated 8d agoView on Hugging Face
60likes4kdownloads
Dataset Card

FineWeb-C: Educational content in many languages, labelled by the community

<center> <img src="https://huggingface.co/spaces/data-is-better-together/fineweb-communications-pack/resolve/main/fineweb-c-card-header.png" alt="FineWeb 2: A sparkling update with 1000s of languages"> </center>

Multilingual data is better together!
Note: We are not actively working on this project anymore. You can continue to contribute annotations and we'll occasionally refresh the exported data.

What is this?

FineWeb-C is a collaborative, community-driven project that expands upon the FineWeb2 dataset. The goal is to create high-quality educational content annotations across hundreds of languages.

By enhancing web content with these annotations, we aim to improve the development of Large Language Models (LLMs) in all languages, making AI technology more accessible and effective globally.

The annotations in this dataset will help train AI systems to automatically identify high-quality educational content in more languages and in turn help build better Large Language Models for all languages.

What the community is doing:

  • For a given language, look at a page of web content from the FineWeb2 dataset in Argilla.
  • Rate how educational the content is.
  • Flag problematic content i.e. content that is malformed or in the wrong language.

Once a language reaches 1,000 annotations, the dataset will be included in this dataset! Alongside rating the educational quality of the content, different language communities are discussing other ways to improve the quality of data for their language in our Discord discussion channel.

What's been done so far?

So far 466 members of the Hugging Face community have submitted 58,188 annotations.

The following languages have reached the 10 annotation threshold to be included in the dataset.

Language CodeLanguage NameCompleted AnnotationsAnnotators
aeb_ArabTunisian Arabic5299
apc_ArabNorth Levantine Arabic2504
arb_ArabStandard Arabic100010
ars_ArabNajdi Arabic10001
ary_ArabMoroccan Arabic100015
arz_ArabEgyptian Arabic10009
asm_BengAssamese241
asm_LatnAssamese10005
ast_LatnAsturian801
bak_CyrlBashkir4424
bar_LatnBavarian10001
ben_BengBangla151
bho_DevaBhojpuri6431
bre_LatnBreton462
bul_CyrlBulgarian361
cat_LatnCatalan345
ces_LatnCzech4595
cmn_HaniMandarin Chinese10003
crh_LatnCrimean Tatar1881
dan_LatnDanish100018
deu_LatnGerman21219
ekk_LatnStandard Estonian1413
eus_LatnBasque24616
fao_LatnFaroese241
fas_ArabPersian10003
fil_LatnFilipino10002
fin_LatnFinnish10007
fra_LatnFrench100028
glg_LatnGalician3418
gmh_LatnMiddle High German10001
goh_LatnOld High German10005
gom_DevaGoan Konkani1601
gsw_LatnSwiss German10002
guj_GujrGujarati111
hin_DevaHindi62122
hsb_LatnUpper Sorbian271
hun_LatnHungarian1243
ind_LatnIndonesian293
ita_LatnItalian100026
jpn_JpanJapanese10005
kas_DevaKashmiri1862
kin_LatnKinyarwanda9123
kor_HangKorean63912
lat_LatnLatin1492
lez_CyrlLezghian551
lij_LatnLigurian10001
lit_LatnLithuanian3233
lug_LatnGanda1321
lvs_LatnStandard Latvian806
mal_MlymMalayalam101
mar_DevaMarathi562
nan_LatnMin Nan Chinese145
nds_LatnLow German3895
nld_LatnDutch36110
nob_LatnNorwegian Bokmål246
npi_DevaNepali (individual language)3875
npi_LatnNepali (individual language)122
pbt_ArabSouthern Pashto111
pcm_LatnNigerian Pidgin2403
pdc_LatnPennsylvania German6222
pfl_LatnPalatine German10001
pol_LatnPolish502
por_LatnPortuguese40418
quz_LatnCusco Quechua801
ron_LatnRomanian2684
rus_CyrlRussian10004
sco_LatnScots5883
sin_SinhSinhala2113
slk_LatnSlovak2215
som_LatnSomali111
spa_LatnSpanish100038
srp_CyrlSerbian934
srp_LatnSerbian111
swe_LatnSwedish10008
tam_TamlTamil10008
tat_LatnTatar2085
tel_TeluTelugu665
tha_ThaiThai4112
tir_EthiTigrinya7456
tok_LatnToki Pona111
tur_LatnTurkish3797
udm_CyrlUdmurt551
ukr_CyrlUkrainian10005
uzn_CyrlNorthern Uzbek202
uzn_LatnNorthern Uzbek442
vie_LatnVietnamese100011
vls_LatnWest Flemish10001
yor_LatnYoruba4437
yue_HaniCantonese10007
zsm_LatnStandard Malay10001

You can help contribute to the dataset [here](https://huggingface.co/spaces/data-is-better-together/fineweb-c).

Note: We are not actively supporting this effort anymore but you can continue to contribute annotations and we'll occasionally refresh the exported data.

Below is an overview of the number of annotations submitted for each language (updated daily).

<iframe src="https://huggingface.co/datasets/data-is-better-together/fineweb-c-progress/embed/sql-console/dhn8hw-" frameborder="0" width="100%" height="560px"></iframe>

Why are we doing this?

There are many languages in the world where no high quality LLMs exist. Having high quality data is a central part of building high quality LLMs. FineWeb2 is a crucial step in improving the availability of high quality data for many languages. We plan to go a step further.

FineWeb-Edu for every language?

FineWeb-Edu is a dataset built on the original FineWeb dataset. The dataset was constructed by developing an educational quality classifier using annotations generated by LLama3-70B-Instruct and using this classifier to retain only the most educational web pages.

FineWeb-Edu outperforms FineWeb on popular benchmark. Crucially, using this approach reduces the amount of data needed to train a high quality LLM reducing the barrier to building a high quality LLM for many languages.

We want to make it possible to build FineWeb-Edu datasets for all the worlds languages. To do this we need annotations in order to train an educational quality classifier.

This in turn will allow us to build the next generation of Large Language Models for many languages.

Why not use LLMs to annotate the data?

For high resources languages, using an LLM to generate educational quality annotations can be a good solution. However, for many languages LLMs are not able to generate high quality annotations — or we don't have enough data to validate whether the annotations are correct.

How can I help?

You can help by contributing to the dataset here and join the community discussions in Discord!

Why would I bother to contribute to this dataset?

Your contributions directly shape the future of AI in your language. Here's why this matters:

  1. 1.Break the AI language barrier: Most commercial AI companies focus on profitable languages, leaving many communities behind. Your work helps bring AI capabilities to more languages.
  1. 1.Keep it open: Unlike proprietary datasets locked away by companies, FineWeb2-C is an open dataset. This means anyone can use it to build AI systems that truly serve their community's needs. Through this open approach we also learn about which approaches work best for different languages.
  1. 1.Be part of something bigger: Just as Wikipedia showed how volunteers can build invaluable resources, the Hugging Face community has created numerous open models and datasets. You're joining a movement to democratize AI technology.

Every annotation counts. Whether you can contribute ten minutes or ten hours, your input helps build a more inclusive future for AI technology 🤗

Who contributed to this dataset so far?

These are the top 10 contributors to this release of the dataset. Make sure to give them a follow on the Hub to show your appreciation!

Hugging Face UsernameSubmissions
stefan-it4,614
tagay1n2,094
hannayukhymenko1,937
hasnachouikhi1,865
Aivis1,613
ivykopal1,365
gaydmi1,112
catastropiyush1,059
theblackcat1021,002
SamerAttrah1,000

Data work is the under appreciated foundation of AI and ML. This dataset is built by the community for the community. Below is a leaderboard that is updated daily and shows all the contributors to this annotation effort.

<iframe src="https://huggingface.co/datasets/data-is-better-together/fineweb-c-progress/embed/sql-console/DJ2n1Z0" frameborder="0" width="100%" height="560px"></iframe>

Language-specific Contributors

Below you can find a list of all the contributors to this release of the dataset for each language ❤️

<details> <summary>Detailed Contributor Statistics for each language</summary>

Assamese (asm_Beng)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
moyoor9724

</details>

Assamese (asm_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Asturian (ast_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
pablocuervo80

</details>

Bangla (ben_Beng)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
shetumohanto15

</details>

Bashkir (bak_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
vikkormallansohn442
MR973439
AigizK2
KarimUva2

</details>

Basque (eus_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Bavarian (bar_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
stefan-it1000

</details>

Bhojpuri (bho_Deva)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
theainerd643

</details>

Breton (bre_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
lbourdois43
Oktogazh3

</details>

Bulgarian (bul_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
krumeto36

</details>

Cantonese (yue_Hani)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Catalan (cat_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
mmarimon18
PereLluis1310
catbru3
JorgeAV2
ljaume1

</details>

Crimean Tatar (crh_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
gaydmi188

</details>

Cusco Quechua (quz_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
daqc80

</details>

Czech (ces_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
MrFrace253
ivykopal106
hroch80
hynky15
mlynatom5

</details>

Danish (dan_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Dutch (nld_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Egyptian Arabic (arz_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Faroese (fao_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
AnnikaSimonsen24

</details>

Filipino (fil_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
mhyles993
maryclara7

</details>

Finnish (fin_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

French (fra_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Galician (glg_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Ganda (lug_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Bronsn132

</details>

German (deu_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Goan Konkani (gom_Deva)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
archanalinguistamberkar160

</details>

Gujarati (guj_Gujr)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
diabolic604511

</details>

Hindi (hin_Deva)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Hindi (hin_Deva)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
catastropiyush926
pp73
Urmish1

</details>

Hungarian (hun_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Indonesian (ind_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
afaji11
prajnapras1910
ayameRushia8

</details>

Italian (ita_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Japanese (jpn_Jpan)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Kashmiri (kas_Deva)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Jagjeet2003125
aloobun61

</details>

Kinyarwanda (kin_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Kalvan20906
David1025
rutsam1

</details>

Korean (kor_Hang)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Latin (lat_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
nataliaElv142
Anna-Katharina7

</details>

Lezghian (lez_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
alialek55

</details>

Ligurian (lij_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
ConseggioLigure1000

</details>

Lithuanian (lit_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
GalvosasJ205
DeividasM109
karal1ius9

</details>

Low German (nds_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
johko213
Taylor658146
stefan-it30
bbunzeck1
dmarx1

</details>

Malayalam (mal_Mlym)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
kurianbenoy10

</details>

Mandarin Chinese (cmn_Hani)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
paperplanedeemo978
guokan-shang12
AdinaY10

</details>

Marathi (mar_Deva)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Disha50
adityapatkar6

</details>

Middle High German (gmh_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
stefan-it1000

</details>

Min Nan Chinese (nan_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Moroccan Arabic (ary_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Najdi Arabic (ars_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
SamerAttrah1000

</details>

Nepali (individual language) (npi_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
amitness10
kshitizkhanal72

</details>

Nepali (individual language) (npi_Deva)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Nigerian Pidgin (pcm_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Cherubikal153
basino300074
Linguistsam13

</details>

North Levantine Arabic (apc_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Northern Uzbek (uzn_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
murodbek32
ltim12

</details>

Northern Uzbek (uzn_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
ltim18
murodbek2

</details>

Norwegian Bokmål (nob_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Old High German (goh_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Palatine German (pfl_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
stefan-it1000

</details>

Pennsylvania German (pdc_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
stefan-it597
Anna-Katharina25

</details>

Persian (fas_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Maani985
mehrdadazizi14
kargaranamir1

</details>

Polish (pol_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
piotr-rybak45
mobarski5

</details>

Portuguese (por_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Romanian (ron_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Russian (rus_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
kitano-o593
kristaller486396
knyazer9
alialek5

</details>

Scots (sco_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
davanstrien564
burtenshaw16
owner8

</details>

Serbian (srp_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
JLouisBiz43
Stopwolf38
Wavelet11
Suzana1

</details>

Serbian (srp_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Wavelet11

</details>

Sinhala (sin_Sinh)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Suchinthana202
Ransaka5
harshanal4

</details>

Slovak (slk_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
ivykopal201
rbelanec199
mvyboh21
mrshu20
real-jiakai1

</details>

Somali (som_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Python223111

</details>

Southern Pashto (pbt_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
ihanif11

</details>

Spanish (spa_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Standard Arabic (arb_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
hasnachouikhi1000
alielfilali014

</details>

Standard Arabic (arb_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Standard Estonian (ekk_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
taidopurason64
TanelAlumae47
jpata30

</details>

Standard Latvian (lvs_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Aivis78
davispuh29
finnayeet22
slckl22
zemais7
minem992

</details>

Standard Malay (zsm_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
theblackcat1021000

</details>

Swedish (swe_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Swiss German (gsw_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
hannayukhymenko957
Anna-Katharina43

</details>

Tamil (tam_Taml)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Tatar (tat_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
tagay1n129
RostBat50
gaydmi16
nurAinur7
inov86

</details>

Telugu (tel_Telu)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Thai (tha_Thai)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Thaweewat359
wannaphong52

</details>

Tigrinya (tir_Ethi)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
meskil711
Gebrejwergs343
Hailay218
Aregawi168
wzbelo93
temesgenTom49

</details>

Toki Pona (tok_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Anna-Katharina11

</details>

Tunisian Arabic (aeb_Arab)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Turkish (tur_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Udmurt (udm_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
codemurt55

</details>

Ukrainian (ukr_Cyrl)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

Upper Sorbian (hsb_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
Korla27

</details>

Vietnamese (vie_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

West Flemish (vls_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

UsernameSubmissions
mariedewulf1000

</details>

Yoruba (yor_Latn)

<details> <summary>User Statistics Table (Minimum 1 submissions)</summary>

</details>

</details>

Using this dataset

The dataset has a default config that contains all the language and configs per language.

To download the dataset using the Hugging Face datasets library, you can use the following code:

python
from datasets import load_dataset

dataset = load_dataset("data-is-better-together/fineweb-c-edu")

To download a specific language, you can use the following code:

python
dataset = load_dataset("data-is-better-together/fineweb-c-edu", name="cmn_Hani")

You can also download the dataset using Pandas

python
import pandas as pd

# Login using e.g. `huggingface-cli login` to access this dataset
df = pd.read_parquet("hf://datasets/data-is-better-together/fineweb-c-edu/arb_Arab/train-00000-of-00001.parquet")

or polars

python

import polars as pl

# Login using e.g. `huggingface-cli login` to access this dataset
df = pl.read_parquet('hf://datasets/davanstrien/fineweb-c-exported-data-test/arb_Arab/train-00000-of-00001.parquet')

Annotations with source metadata

The with_metadata config contains the same 45,160 annotation rows as default, with 58 columns including source URLs, language scores, duplicate-cluster sizes, complete selected-source records and versioned join provenance. All eight original annotation columns are preserved. Existing default and language configs keep their current data.

python
from datasets import load_dataset

dataset = load_dataset(
    "data-is-better-together/fineweb-c", name="with_metadata", split="train"
)

Source metadata is recovered for 43,086 rows (95.4%); 2,073 unmatched rows and one ambiguous row are retained with null selected-source fields. Most links refer to pinned historical sample datasets. Full-FineWeb2 membership was separately checked only for Kinyarwanda and Tatar. Labels continue to describe the original text, including when a later source version differs.

See metadata documentation for join rules, every added column, coverage by language, pinned sources and downloadable candidate audit tables.

Data Fields

The dataset contains the following columns:

Column NameTypeDescription
idstringA unique identifier for each annotation record
textstringThe text of the web page
educationalvaluelabelslist[string]A list of labels indicating the educational value of the web page rated by the community
annotator_idsstringA string ID for the annotator
problematiccontentlabel_presentbooleanA flag indicating the presence of at leaste one 'problematic' label being assigned to the text
problematiccontentlabel_agreementfloatThe agreement of the annotator with the problematic content label
language_namesstrThe name of the language page
language_codestrThe code of the language

The main things to note (we'll update this as we get more data)

  • Some languages already have multiple annotations per page. So far we haven't done any processing on these rows so people are free to calculate the agreement of the annotators in whatever way they want.
  • For languages with many active annotators, we may increase the overlap of annotations over time to further improve the quality of the dataset.
  • Some languages contain many problematic content labels. These often occur when the language detection was not correct. There is a problematic_content_label_present boolean column that indicates if the page contains at least one problematic content label. If you want to remove these rows you can do so by filtering on this column. Alternatively, you can use the problematic_content_label_agreement column to filter on the agreement of the annotators i.e. only remove rows where the annotators agree on the problematic content label. For many of the most active language efforts we're working with the community to improve the quality of the data so we hope the number of problematic content labels will decrease over time.

Licensing Information

The dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 license. The use of this dataset is also subject to CommonCrawl's Terms of Use.

Citation

Citation information needs to be added

Last Updated

2025-07-08