CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LocalDoc /azerbaijani-pretrain-corpus Azerbaijani Pretraining Corpus (merged & deduplicated) A cleaned Azerbaijani text corpus assembled for language-model pretraining, merging two curated sources and removing exact duplicates. Contents Documents: 6,931,898 Tokens: ~5.36B (measured with the o200k_base tokenizer; an Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base segments agglutinative Azerbaijani inefficiently) Avg tokens/document: ~773 Fields text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.texttext-generation1M<n<10M0 likes385 downloads4mo agoHugging Face02LocalDoc /community_oscar_azerbaijani Community-OSCAR Azerbaijani This is Azerbaijani version Community OSCAR dataset https://huggingface.co/datasets/oscar-corpus/community-oscar. Dataset Statistics (Aggregate) Metric Value Language Azerbaijani (az) Average per release 3.36 GiB, 603,832 documents Words per release ~408.8M words Characters per release ~3.12B characters Total size (all releases) 137.62 GiB Total lines 24.76M Total words 16.76B words Total characters 128.07B characters… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani.tabulartext-generation10M<n<100M0 likes91 downloads11mo agoHugging Face03LocalDoc /community_oscar_azerbaijani_scored Azerbaijani Web Corpus with Quality Scores This dataset is the full Azerbaijani web corpus LocalDoc/community_oscar_azerbaijani with a continuous quality score attached to every document. It is intended as the filtering layer for building a clean Azerbaijani pretraining corpus: each document carries a score that lets you keep, clean, or drop it according to your own thresholds. What was done Every document in the source corpus was scored by the model… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani_scored.texttext-classification10M<n<100M0 likes63 downloads4mo agoHugging Face04khazarai /AARA_Azerbaijani_LLM_Benchmark AARA: Azerbaijani Advanced Reasoning Assessment This dataset is the Azerbaijani-translated version of the emre/TARA_Turkish_LLM_Benchmark. textquestion-answeringn<1K1 likes43 downloads6mo agoHugging Face05aznlp /azerbaijani-blogs Azerbaijani Blogs dataset Dataset Details Dataset Description This dataset provides blogs written in azerbaijani language with categories and tags for each. Language(s) (NLP): Azerbaijani License: Apache license 2.0 Data Source All the data was found in public resources of kayzen.az blogging website without any restriction. texttext-classification1K<n<10K3 likes38 downloads2y agoHugging Face06kamaalg /azerbaijani-instructions Azerbaijani Instruction Dataset (v0) Azerbaijani (instruction, response) pairs for supervised fine-tuning (SFT) of Azerbaijani language models — part of an open Azerbaijani LLM stack. Instruction data is genuinely scarce for Azerbaijani; this is both training data for our models and a reusable standalone artifact for anyone building Azerbaijani instruction-following models. Contents seeds_az.jsonl — 45 hand-authored, high-quality seed pairs spanning 16 task… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-instructions.texttext-generationn<1K1 likes37 downloads3mo agoHugging Face07LocalDoc /numina_math_azerbaijaniThis is part of the translated version of the original dataset: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT texttext-generation100K<n<1M1 likes24 downloads9mo agoHugging Face08Yusiko /azerbaijani-wiki-instruct-alpacaAn Azerbaijani instruction-following dataset in Alpaca format (instruction, input, output).Useful for supervised fine-tuning (SFT) to improve instruction following and long-form, explanatory answers in Azerbaijani. Quick facts Rows: 167,590 Split: train only License: MIT Main file: azerbaijani_wiki_instruct.jsonl (~432 MB) Auto-converted Parquet: ~225 MB Data schema Each record contains: instruction (string): the task/prompt in Azerbaijani input (string): optional… See the full description on the dataset page: https://huggingface.co/datasets/Yusiko/azerbaijani-wiki-instruct-alpaca.texttext-generation100K<n<1M0 likes23 downloads9mo agoHugging Face09Yusiko /little_dataset_Azerbaijani🇦🇿 Azerbaijani Instruction Dataset 📚 About this dataset This dataset contains 400 Azerbaijani-language instruction–response pairs, created for fine-tuning conversational and educational AI models. It follows the Alpaca-style format with "input" and "output" fields and focuses on clarity, accuracy, and linguistic richness. 🧩 Structure input → The user’s instruction or question (in Azerbaijani) output → The correct and natural Azerbaijani response Each topic includes 100 high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Yusiko/little_dataset_Azerbaijani.texttext-generationn<1K0 likes21 downloads11mo agoHugging Face10kamaalg /azerbaijani-corpus-v0 Azerbaijani Pretraining Corpus (v0) A cleaned, deduplicated, PII-redacted ~1.0 billion token Latin-script Azerbaijani corpus for language-model pretraining, built with a reproducible datatrove pipeline from open multilingual web + encyclopedic sources. Full provenance, methodology, and limitations are in the Datasheet (Gebru-style). Summary Language Azerbaijani (az/azj), Latin script only Documents 1,711,442 Tokens ~1.0B (az_unigram_32k; train… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-corpus-v0.texttext-generation1K<n<10K1 likes19 downloads3mo agoHugging Face11eljanmahammadli /alpaca-azerbaijani-gpt-4o-mini Dataset Details This is a translated version of the Alpaca dataset into the Azerbaijani language, using GPT-4o-mini. textquestion-answering10K<n<100K1 likes17 downloads2y agoHugging Face12LocalDoc /Finance-Instruct-AzerbaijaniThis is part of a translated version of the original dataset: https://huggingface.co/datasets/Josephgflowers/Finance-Instruct-500k tabulartext-generation10K<n<100K0 likes14 downloads11mo agoHugging Face13LocalDoc /medical-o1-reasoning-SFT-azerbaijaniThis is a translated version of the dataset https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT textquestion-answering10K<n<100K1 likes12 downloads1y agoHugging Face14nazrinburz /internet_archive_azerbaijanitexttext-generationn<1K0 likes8 downloads1y agoHugging Face15LocalDoc /alpaca_cleaned_azerbaijanigatedThis is translated into Azerbaijani Alpaca-Cleaned dataset. texttext-generation10K<n<100K1 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.