NagaYu/tessera-mixed-pages
Tessera mixed-page corpus Synthetic multilingual web pages with exact span boundaries and multi-label language annotations, for training and evaluating span-level language identification. Why this is synthetic, stated up front No public corpus annotates span boundaries and languages inside real multilingual web pages at scale. Annotating one by hand across 479 languages, most of them low-resource, is not feasible and would itself be error-prone in exactly the… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/tessera-mixed-pages.
046
