CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01UAzimov /uzbek-instruct-llmuzbek-instruct-llm is a corpus of more than 15,000 records. It's made for instruct fine-tuning large language models for Uzbek language. It's mostly translated from other instruct datasets with some extra data added texttext-generation10K<n<100K9 likes160 downloads2y agoHugging Face02uzinfocom-edu-ai /uzbek-asr-curated-701h Uzbek ASR Curated Dataset (701 hours) A curated multi-source Uzbek speech dataset for automatic speech recognition (ASR) training and evaluation. Dataset Description Language Uzbek (Latin script with okina ʻ) Total utterances 337,920 Total duration ~701 hours Audio format 16 kHz mono WAV (PCM_16) Manifest format NeMo JSONL Splits train (94%) / val (3%) / test (3%) Splits Split Utterances Hours Train 317,655… See the full description on the dataset page: https://huggingface.co/datasets/uzinfocom-edu-ai/uzbek-asr-curated-701h.audioautomatic-speech-recognition100K<n<1M1 likes112 downloads3mo agoHugging Face03javohirmat /uzbek-legal-corpus Uzbek Legal Corpus (Oʻzbek huquqiy korpusi) 25 cleaned, audited, machine-readable texts of the Republic of Uzbekistan's Constitution, 20 codes, and 4 major laws — in Uzbek (Latin script), sourced from the official National Database of Legislation (lex.uz). Snapshot: July 2026. 7,368 articles. Built for NLP / LLM / RAG work on Uzbek legal text. Released by Tomaris AI. Contents Path What it is data/raw/*.txt Cleaned plain text, one file per code/law… See the full description on the dataset page: https://huggingface.co/datasets/javohirmat/uzbek-legal-corpus.texttext-retrieval1K<n<10K1 likes111 downloads2mo agoHugging Face04risqaliyevds /uzbek_ner Uzbek NER Dataset About the Dataset This dataset is created for Named Entity Recognition (NER) in Uzbek texts. The dataset includes named entities from various categories such as persons, places, organizations, dates, and more. Data Structure The data is provided in JSON format with the following structure: { "LOC": ["Location names"], "ORG": ["Organization names"], "PERSON": ["Person names"], "DATE": ["Date expressions"], "MONEY":… See the full description on the dataset page: https://huggingface.co/datasets/risqaliyevds/uzbek_ner.texttoken-classification10K<n<100K4 likes63 downloads2y agoHugging Face05rustam1221 /uzbek-asr-train-manifests Uzbek ASR Training Manifests The exact training, validation and test splits behind rustam1221/uzbek-asr-gigaam: 974 hours of Uzbek speech drawn from seven public corpora, filtered, text-normalized, and split by speaker. No audio is copied. Each row is a pointer — a parquet file plus a row index in the upstream dataset — and the training dataloader decodes the audio when the batch is built. That keeps the whole corpus definition at 200 MB instead of roughly a terabyte of… See the full description on the dataset page: https://huggingface.co/datasets/rustam1221/uzbek-asr-train-manifests.textautomatic-speech-recognition1K<n<10K0 likes58 downloads21d agoHugging Face06Rajan2026 /soas-english-uzbek-rag-evaluation SOAS English-Uzbek Retrieval Pilot Dataset Summary This folder documents a bilingual English-Uzbek retrieval evaluation benchmark for culturally grounded RAG systems. The 400-row public pilot release is retrieval-only: it contains questions and source-document targets, but it intentionally excludes answer, context, excerpt, and source-text fields. This is a pilot benchmark with documented quality flags, template-generated examples, and domain mismatches. The rows… See the full description on the dataset page: https://huggingface.co/datasets/Rajan2026/soas-english-uzbek-rag-evaluation.texttext-retrievaln<1K0 likes49 downloads19d agoHugging Face07sukhrobnurali /uzbek_spam_dataset Uzbek Spam Detection Dataset A synthetic dataset for training spam detection models for Uzbek language, specifically targeting Telegram-style messages. Dataset Description This dataset contains 2000 Uzbek text messages labeled as either "spam" or "normal". Dataset Summary Total samples: 2000 Train split: 1800 Test split: 200 Labels: spam, normal Language: Uzbek (Latin and Cyrillic scripts) Spam Categories Covered Aggressive advertising and… See the full description on the dataset page: https://huggingface.co/datasets/sukhrobnurali/uzbek_spam_dataset.texttext-classification1K<n<10K1 likes44 downloads9mo agoHugging Face08sukhrobnurali /uzbek_constitution Constitution of the Republic of Uzbekistan (Trilingual Dataset) Dataset Summary This dataset contains the Constitution of the Republic of Uzbekistan (New Edition, adopted via referendum on April 30, 2023) aligned in three languages: 🇺🇿 Uzbek (Latin script) 🇷🇺 Russian 🇬🇧 English The data was carefully scraped and processed from the official National Database of Legislation of the Republic of Uzbekistan. It serves as a high-quality parallel corpus for legal NLP… See the full description on the dataset page: https://huggingface.co/datasets/sukhrobnurali/uzbek_constitution.tabulartranslationn<1K4 likes35 downloads10mo agoHugging Face09ML-Jonibek /English-Uzbek-Translation-1 🌐 English–Uzbek Translation Dataset A parallel corpus for English ↔ Uzbek machine translation, curated to support research and development of NLP models for the Uzbek language — one of the most underrepresented Turkic languages in open-source datasets. 📖 Dataset Description This dataset contains aligned sentence pairs in English and Uzbek, designed for training, fine-tuning, and evaluating neural machine translation (NMT) models. The dataset aims to bridge the resource… See the full description on the dataset page: https://huggingface.co/datasets/ML-Jonibek/English-Uzbek-Translation-1.texttranslation10K<n<100K1 likes26 downloads5mo agoHugging Face10shokhjakhon /uzbektext100K<n<1M0 likes26 downloads9d agoHugging Face11Mehriddin1997 /lex-uzbek-laws Lex.uz — “Xavfsizlik va huquq-tartibot muhofazasi” (Partial) Dataset (Uzbek) Dataset Summary This dataset contains Uzbek legal texts collected from the lex.uz portal, specifically from the “Xavfsizlik va huquq-tartibot muhofazasi” (Security and law enforcement protection) section.The collection is partial: the full section has not been fully scraped yet, and the dataset may be expanded in future updates. The dataset is intended for: domain adaptation / fine-tuning of… See the full description on the dataset page: https://huggingface.co/datasets/Mehriddin1997/lex-uzbek-laws.texttext-generationn<1K0 likes20 downloads7mo agoHugging Face12nickoo004 /uzbekdatatexttext-generation1K<n<10K2 likes16 downloads2y agoHugging Face13Jonibek21 /English-Uzbek-Translationtext10K<n<100K2 likes15 downloads5mo agoHugging Face14Tohirju /uzbek-asr-full-datagated Saidzoda Lab — Gated Research Dataset Part of Saidzoda Lab's Central-Asian language research (Tajik, Uzbek, Kazakh, Kyrgyz). Dataset contents, provenance, and statistics are not publicly disclosed. Access is granted manually on request. text100K<n<1M0 likes8 downloads1mo agoHugging Face15tulkinyusuf /customer-support-uzbek-120btext1K<n<10K0 likes3 downloads8mo agoHugging Face16pythoninverter /English-Uzbek-Translationtext10K<n<100K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.