CoolFace
Datasetpublic

NagaYu/tessera-mixed-pages

Tessera mixed-page corpus Synthetic multilingual web pages with exact span boundaries and multi-label language annotations, for training and evaluating span-level language identification. Why this is synthetic, stated up front No public corpus annotates span boundaries and languages inside real multilingual web pages at scale. Annotating one by hand across 479 languages, most of them low-resource, is not feasible and would itself be error-prone in exactly the… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/tessera-mixed-pages.

sourceHugging Facecc-by-sa-4.0updated 21d agoView on Hugging Face
0likes46downloads
filetest_adversarial.parquet3.0 MBdownload
filetest_romanized.parquet2.0 MBdownload
filetest.parquet5.2 MBdownload
filetrain.parquet51.5 MBdownload
filevalidation.parquet3.9 MBdownload

NagaYu/tessera-mixed-pages · main · files are served by the source, never re-hosted here