datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TajikWeather-Tajikistan-1940-2026
🌤️ Weather-Tajikistan-1940-2026
Comprehensive Hourly Weather Dataset for 11 Regions of Tajikistan (1940–2026)
This dataset provides 8,337,912 hourly weather records from 11 cities/regions of Tajikistan, spanning from 1940-01-01 to 2026-06-20 . It is designed for climate research, time-series analysis, and machine learning applications.
✨ Key Features
📊 11 regions across Tajikistan
🕐 Hourly resolution from 1940 to 2026 (86 years)
🌡️ 20+ meteorological… See the full description on the dataset page: https://huggingface.co/datasets/arabovs-ai-lab/Weather-Tajikistan-1940-2026.tajik-lora-qlora-benchmark
Tajik LoRA/QLoRA Benchmark
📊 Description
This benchmark contains the complete results of fine-tuning 15+ language models (from 124M to 7B parameters) on a subset of the Tajik language (1000 sentences from the TajikNLPWorld/tajik-web-corpus).The study compares full fine-tuning versus LoRA/QLoRA, evaluating model quality (perplexity), GPU memory usage, and training time.
Key Findings
GPT‑2 medium (full fine-tuning) achieves the lowest perplexity (3.48), but… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-lora-qlora-benchmark.tajik-web-corpus
Dataset Card for Tajik Web Corpus
Dataset Details
Dataset Description
The Tajik Web Corpus is a large-scale collection of 319,298 documents in the Tajik language, totaling approximately 1.11 billion characters and 168.5 million words. The data has been cleaned, normalized, and deduplicated, and is provided in JSONL format with the following fields: title, text, category, source, date, and URL. It covers various domains including news, Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-web-corpus.tajik-wiki-corpus
Dataset Card for Tajik Wikipedia Corpus
Dataset Details
Dataset Description
The Tajik Wikipedia Corpus is a collection of 79,985 Wikipedia articles in the Tajik language, totaling approximately 101.6 million characters and 15.3 million words. The data has been extracted from the Tajik Wikipedia dump and processed to ensure clean, well‑structured text suitable for NLP applications. The corpus includes articles with titles, categories, and metadata.… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-wiki-corpus.TajPersParallelCorpus
Dataset Card for Tajik–Persian Parallel Corpus
Dataset Details
Dataset Description
The Tajik–Persian Parallel Corpus is a large-scale parallel corpus containing 328,253 aligned Tajik–Persian sentence pairs collected from multiple sources, including news, poetry, prose, lexical resources, and named-entity lists. It is intended for machine translation, cross-lingual retrieval, linguistic analysis, tokenizer evaluation, and other NLP tasks.
Curated… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/TajPersParallelCorpus.TajPersParallelLexicalCorpus
Dataset Card for TajPersParallelLexicalCorpus
Dataset Details
Dataset Description
The TajPersParallelLexicalCorpus is a parallel lexical resource for the Tajik–Persian language pair. It contains 43,819 records with word‑level translations, part‑of‑speech annotations, and usage examples from classical and modern literature. The corpus is designed for machine translation, cross‑lingual NLP, and linguistic research involving Tajik (Cyrillic script) and… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/TajPersParallelLexicalCorpus.khf_news
Dataset Card for KHF News Dataset
Dataset Details
Dataset Description
This dataset contains news articles published on the official website of the Committee for Emergency Situations and Civil Defense under the Government of the Republic of Tajikistan (www.khf.tj). The articles cover a wide range of topics related to emergency situations, disaster response, civil defense, and government activities in Tajikistan. The dataset is intended for NLP… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/khf_news.
