tajik
Datasets
All datasets matching “tajik”tajik-curated-benchmark
Tajik Curated General Benchmark
Dataset Summary
The Tajik Curated General Benchmark is a multi-domain multiple-choice evaluation dataset for the Tajik language. It comprises 2,266 questions distributed across six domains: History, Biology, Literature, Tajik Language (grammar, morphology, and syntax), Law, and Geography. Questions were manually curated from diverse Tajik-language sources to assess cross-domain breadth rather than depth in any single area.
Its defining… See the full description on the dataset page: https://huggingface.co/datasets/zehnlab/tajik-curated-benchmark.ipfs_tajikistan_laws_ir
Tajikistan legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_tajikistan_laws (revision a65dcc660a4365d4b07820f83c6e2bad1f629b20) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Tajikistan prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_tajikistan_laws_ir.Tajiktajik-asr-corpus-v3
tajik-asr-corpus-v3
1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled)
plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind
Peacockery/omni-ctc-300m-tajik
(16.9% WER on FLEURS test, 37.6% on held-out conversational speech).
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/.
Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list),
and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.tajik-mmlu
Tajik MMLU
Dataset Summary
Tajik MMLU is a Tajik-language translation of the original English MMLU benchmark, covering the same 57 subject categories across STEM, humanities, and social sciences, and preserving the original question structure and answer options.
Its role is complementary to English MMLU: while the latter measures retention of general knowledge, Tajik MMLU measures whether a model can access and apply the same knowledge when prompted in Tajik. Gaps between… See the full description on the dataset page: https://huggingface.co/datasets/zehnlab/tajik-mmlu.Weather-Tajikistan-1940-2026
🌤️ Weather-Tajikistan-1940-2026
Comprehensive Hourly Weather Dataset for 11 Regions of Tajikistan (1940–2026)
This dataset provides 8,337,912 hourly weather records from 11 cities/regions of Tajikistan, spanning from 1940-01-01 to 2026-06-20 . It is designed for climate research, time-series analysis, and machine learning applications.
✨ Key Features
📊 11 regions across Tajikistan
🕐 Hourly resolution from 1940 to 2026 (86 years)
🌡️ 20+ meteorological… See the full description on the dataset page: https://huggingface.co/datasets/arabovs-ai-lab/Weather-Tajikistan-1940-2026.
