datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lid-tyv
spellman LID corpus — tyv
182,189 rows / 65.9M chars of tyv-language text with
provenance, built for training and evaluating the
spellman language detector. Rows are
single-line texts: news articles and library works (paragraph/sentence-group
sized) plus Telegram/VK community posts and comments (wild register — emoji, hashtags,
code-switching artefacts, folk spellings kept on purpose).
Row length: median 173 chars (p10 48, p90 743).
Configs
config
contents… See the full description on the dataset page: https://huggingface.co/datasets/vpermilp/lid-tyv.icml2026-Tyv61ZKb9s-repro-traces
Agent traces
Agent sessions published from a Trackio Logbook.
tyv-rus-200k
tyv-rus-200k data card
This data was collected via www.tyvan.ru platform by linguists, scientists, journalists, volunteers, etc.
Actually here 296k rows. Almost 300k
Dataset Details
Dataset Description
Curated by: Ali Kuzhuget (tech and data), Ondar Choygan (data) contributors
Language(s) (NLP): Tyvan (Tuvan), Russian
License:: CC BY 4.0.
Below is the brief information about the languages
Language
Language code on the website
ISO 639-3
Glottolog… See the full description on the dataset page: https://huggingface.co/datasets/Agisight/tyv-rus-200k.tyvan-russian-parallel-50kA 50K sample from the Russian-Tyvan parallel corpus collected at https://tyvan.ru.
