datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nogai-Russian-SFT-Biblical-v1
Nogai-Russian SFT Biblical Corpus v1 (superseded — use v2)
Use Nogai-Russian-SFT-Biblical-v2 instead.
This version is kept unchanged because the published SFT adapter was trained on it. Its splits are not suitable for evaluation (see below).
Russian↔Nogai translation instructions in ChatML format, from human Bible translations by the Institute for Bible Translation (IBT). Used for Phase 2 SFT of NogaiLLM.
What is in it (measured September 2026)
Rows… See the full description on the dataset page: https://huggingface.co/datasets/ansarzeinulla/Nogai-Russian-SFT-Biblical-v1.Nogai-Russian-SFT-Biblical-v2
Nogai-Russian SFT Biblical Corpus v2
Russian↔Nogai translation instructions (ChatML) from human Bible translations by the Institute for Bible Translation (IBT). This is a clean rebuild of v1 with splits that can be used for evaluation. Built with build_sft_clean.py (seed 42).
Splits
Split
Pairs
Rows (2 directions per pair)
train
507
1,014
validation
60
120
test
58
116
How it was built from v1
4,310 v1 rows reduce to 650 unique… See the full description on the dataset page: https://huggingface.co/datasets/ansarzeinulla/Nogai-Russian-SFT-Biblical-v2.
