ansarzeinulla/Nogai-Unified-Corpus-v1
Nogai Unified Corpus (NUC) v1 Monolingual text in Nogai (Kipchak Turkic, Cyrillic script) for continued pre-training of language models. It was built for NogaiLLM. Size Rows (sentences / short paragraphs) 163,531: 155,354 train, 8,177 validation (95 / 5) Words 2.35 M Characters 18.1 M UTF-8 text 33.4 MB (files: 35.6 MB) Tokens (Qwen2.5 tokenizer) 9.8 M Format JSONL, one {"text": ...} per row Sources Newspapers: digitised… See the full description on the dataset page: https://huggingface.co/datasets/ansarzeinulla/Nogai-Unified-Corpus-v1.
This repository belongs to ansarzeinulla on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
