slsa
Datasets
All datasets matching “slsa”sl-sample-120sl-saiga
sl-saiga
Bilingual (Tajik + Russian) corpus of the laws of the Republic of Tajikistan, 1990–2025,
prepared for retrieval-augmented generation. Text is cleaned with legacy Tajik-font/glyph
restoration; every record carries temporal metadata.
chunks.jsonl — one JSON object per passage:
chunk_id, text, law, year, lang (tg|ru), article, status (in_force|repealed), doc_type (law|amnesty|constitution|conventions), source, versions.
72,920 passages (38,393 tg + 34,527 ru); 51,857… See the full description on the dataset page: https://huggingface.co/datasets/Tohirju/sl-saiga.pinkpen
pinkpen
This dataset was generated using a phospho starter pack.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
School
