NagaYu/tessera-mixed-pages
Tessera mixed-page corpus Synthetic multilingual web pages with exact span boundaries and multi-label language annotations, for training and evaluating span-level language identification. Why this is synthetic, stated up front No public corpus annotates span boundaries and languages inside real multilingual web pages at scale. Annotating one by hand across 479 languages, most of them low-resource, is not feasible and would itself be error-prone in exactly the… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/tessera-mixed-pages.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face