CoolFace
Datasetpublicgated

Tohirju/sl-saiga

sl-saiga Bilingual (Tajik + Russian) corpus of the laws of the Republic of Tajikistan, 1990–2025, prepared for retrieval-augmented generation. Text is cleaned with legacy Tajik-font/glyph restoration; every record carries temporal metadata. chunks.jsonl — one JSON object per passage: chunk_id, text, law, year, lang (tg|ru), article, status (in_force|repealed), doc_type (law|amnesty|constitution|conventions), source, versions. 72,920 passages (38,393 tg + 34,527 ru); 51,857… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/sl-saiga.

sourceHugging Faceotherupdated 16d agoView on Hugging Face
0likes15downloads

No commit history came back for main. The revision may not exist, or the source declined the request.