CoolFace
Datasetpublicgated

Tohirju/sl-saiga

sl-saiga Bilingual (Tajik + Russian) corpus of the laws of the Republic of Tajikistan, 1990–2025, prepared for retrieval-augmented generation. Text is cleaned with legacy Tajik-font/glyph restoration; every record carries temporal metadata. chunks.jsonl — one JSON object per passage: chunk_id, text, law, year, lang (tg|ru), article, status (in_force|repealed), doc_type (law|amnesty|constitution|conventions), source, versions. 72,920 passages (38,393 tg + 34,527 ru); 51,857… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/sl-saiga.

sourceHugging Faceotherupdated 16d agoView on Hugging Face
0likes15downloads

Tohirju/sl-saiga · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.