Tohirju/sl-saiga
sl-saiga Bilingual (Tajik + Russian) corpus of the laws of the Republic of Tajikistan, 1990–2025, prepared for retrieval-augmented generation. Text is cleaned with legacy Tajik-font/glyph restoration; every record carries temporal metadata. chunks.jsonl — one JSON object per passage: chunk_id, text, law, year, lang (tg|ru), article, status (in_force|repealed), doc_type (law|amnesty|constitution|conventions), source, versions. 72,920 passages (38,393 tg + 34,527 ru); 51,857… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/sl-saiga.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face