CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alifbank /Tajiktext-generation0 likes153 downloads2y agoHugging Face02arabovs-ai-lab /Weather-Tajikistan-1940-2026gated 🌤️ Weather-Tajikistan-1940-2026 Comprehensive Hourly Weather Dataset for 11 Regions of Tajikistan (1940–2026) This dataset provides 8,337,912 hourly weather records from 11 cities/regions of Tajikistan, spanning from 1940-01-01 to 2026-06-20 . It is designed for climate research, time-series analysis, and machine learning applications. ✨ Key Features 📊 11 regions across Tajikistan 🕐 Hourly resolution from 1940 to 2026 (86 years) 🌡️ 20+ meteorological… See the full description on the dataset page: https://huggingface.co/datasets/arabovs-ai-lab/Weather-Tajikistan-1940-2026.table-question-answering1M<n<10M0 likes117 downloads3mo agoHugging Face03TajikNLPWorld /tajik-lora-qlora-benchmark Tajik LoRA/QLoRA Benchmark 📊 Description This benchmark contains the complete results of fine-tuning 15+ language models (from 124M to 7B parameters) on a subset of the Tajik language (1000 sentences from the TajikNLPWorld/tajik-web-corpus).The study compares full fine-tuning versus LoRA/QLoRA, evaluating model quality (perplexity), GPU memory usage, and training time. Key Findings GPT‑2 medium (full fine-tuning) achieves the lowest perplexity (3.48), but… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-lora-qlora-benchmark.imagetext-generationn<1K0 likes70 downloads6mo agoHugging Face04TajikNLPWorld /tajik-web-corpusgated Dataset Card for Tajik Web Corpus Dataset Details Dataset Description The Tajik Web Corpus is a large-scale collection of 319,298 documents in the Tajik language, totaling approximately 1.11 billion characters and 168.5 million words. The data has been cleaned, normalized, and deduplicated, and is provided in JSONL format with the following fields: title, text, category, source, date, and URL. It covers various domains including news, Wikipedia… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-web-corpus.texttext-classification100K<n<1M0 likes52 downloads29d agoHugging Face05TajikNLPWorld /tajik-wiki-corpusgated Dataset Card for Tajik Wikipedia Corpus Dataset Details Dataset Description The Tajik Wikipedia Corpus is a collection of 79,985 Wikipedia articles in the Tajik language, totaling approximately 101.6 million characters and 15.3 million words. The data has been extracted from the Tajik Wikipedia dump and processed to ensure clean, well‑structured text suitable for NLP applications. The corpus includes articles with titles, categories, and metadata.… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-wiki-corpus.texttext-classification10K<n<100K1 likes35 downloads29d agoHugging Face06TajikNLPWorld /TajPersParallelCorpusgated Dataset Card for Tajik–Persian Parallel Corpus Dataset Details Dataset Description The Tajik–Persian Parallel Corpus is a large-scale parallel corpus containing 328,253 aligned Tajik–Persian sentence pairs collected from multiple sources, including news, poetry, prose, lexical resources, and named-entity lists. It is intended for machine translation, cross-lingual retrieval, linguistic analysis, tokenizer evaluation, and other NLP tasks. Curated… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/TajPersParallelCorpus.texttranslation100K<n<1M0 likes27 downloads28d agoHugging Face07TajikNLPWorld /TajPersParallelLexicalCorpusgated Dataset Card for TajPersParallelLexicalCorpus Dataset Details Dataset Description The TajPersParallelLexicalCorpus is a parallel lexical resource for the Tajik–Persian language pair. It contains 43,819 records with word‑level translations, part‑of‑speech annotations, and usage examples from classical and modern literature. The corpus is designed for machine translation, cross‑lingual NLP, and linguistic research involving Tajik (Cyrillic script) and… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/TajPersParallelLexicalCorpus.texttranslation10K<n<100K2 likes24 downloads29d agoHugging Face08TajikNLPWorld /khf_newsgated Dataset Card for KHF News Dataset Dataset Details Dataset Description This dataset contains news articles published on the official website of the Committee for Emergency Situations and Civil Defense under the Government of the Republic of Tajikistan (www.khf.tj). The articles cover a wide range of topics related to emergency situations, disaster response, civil defense, and government activities in Tajikistan. The dataset is intended for NLP… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/khf_news.texttext-classification1K<n<10K0 likes2 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.