datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
azerbaijani-pretrain-corpus
Azerbaijani Pretraining Corpus (merged & deduplicated)
A cleaned Azerbaijani text corpus assembled for language-model pretraining,
merging two curated sources and removing exact duplicates.
Contents
Documents: 6,931,898
Tokens: ~5.36B (measured with the o200k_base tokenizer; an
Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base
segments agglutinative Azerbaijani inefficiently)
Avg tokens/document: ~773
Fields
text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.community_oscar_azerbaijani
Community-OSCAR Azerbaijani
This is Azerbaijani version Community OSCAR dataset https://huggingface.co/datasets/oscar-corpus/community-oscar.
Dataset Statistics (Aggregate)
Metric
Value
Language
Azerbaijani (az)
Average per release
3.36 GiB, 603,832 documents
Words per release
~408.8M words
Characters per release
~3.12B characters
Total size (all releases)
137.62 GiB
Total lines
24.76M
Total words
16.76B words
Total characters
128.07B characters… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani.community_oscar_azerbaijani_scored
Azerbaijani Web Corpus with Quality Scores
This dataset is the full Azerbaijani web corpus
LocalDoc/community_oscar_azerbaijani
with a continuous quality score attached to every document. It is intended
as the filtering layer for building a clean Azerbaijani pretraining corpus:
each document carries a score that lets you keep, clean, or drop it according
to your own thresholds.
What was done
Every document in the source corpus was scored by the model… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani_scored.AARA_Azerbaijani_LLM_Benchmark
AARA: Azerbaijani Advanced Reasoning Assessment
This dataset is the Azerbaijani-translated version of the emre/TARA_Turkish_LLM_Benchmark.
azerbaijani-blogs
Azerbaijani Blogs dataset
Dataset Details
Dataset Description
This dataset provides blogs written in azerbaijani language with categories and tags for each.
Language(s) (NLP): Azerbaijani
License: Apache license 2.0
Data Source
All the data was found in public resources of kayzen.az blogging website without any restriction.
azerbaijani-instructions
Azerbaijani Instruction Dataset (v0)
Azerbaijani (instruction, response) pairs for supervised fine-tuning (SFT) of Azerbaijani language
models — part of an open Azerbaijani LLM stack. Instruction data is genuinely scarce for Azerbaijani;
this is both training data for our models and a reusable standalone artifact for anyone building
Azerbaijani instruction-following models.
Contents
seeds_az.jsonl — 45 hand-authored, high-quality seed pairs spanning 16 task… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-instructions.numina_math_azerbaijaniThis is part of the translated version of the original dataset: https://huggingface.co/datasets/AI-MO/NuminaMath-CoT
azerbaijani-wiki-instruct-alpacaAn Azerbaijani instruction-following dataset in Alpaca format (instruction, input, output).Useful for supervised fine-tuning (SFT) to improve instruction following and long-form, explanatory answers in Azerbaijani.
Quick facts
Rows: 167,590
Split: train only
License: MIT
Main file: azerbaijani_wiki_instruct.jsonl (~432 MB)
Auto-converted Parquet: ~225 MB
Data schema
Each record contains:
instruction (string): the task/prompt in Azerbaijani
input (string): optional… See the full description on the dataset page: https://huggingface.co/datasets/Yusiko/azerbaijani-wiki-instruct-alpaca.little_dataset_Azerbaijani🇦🇿 Azerbaijani Instruction Dataset
📚 About this dataset
This dataset contains 400 Azerbaijani-language instruction–response pairs, created for fine-tuning conversational and educational AI models.
It follows the Alpaca-style format with "input" and "output" fields and focuses on clarity, accuracy, and linguistic richness.
🧩 Structure
input → The user’s instruction or question (in Azerbaijani)
output → The correct and natural Azerbaijani response
Each topic includes 100 high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Yusiko/little_dataset_Azerbaijani.azerbaijani-corpus-v0
Azerbaijani Pretraining Corpus (v0)
A cleaned, deduplicated, PII-redacted ~1.0 billion token Latin-script Azerbaijani corpus for
language-model pretraining, built with a reproducible datatrove
pipeline from open multilingual web + encyclopedic sources. Full provenance, methodology, and limitations
are in the Datasheet (Gebru-style).
Summary
Language
Azerbaijani (az/azj), Latin script only
Documents
1,711,442
Tokens
~1.0B (az_unigram_32k; train… See the full description on the dataset page: https://huggingface.co/datasets/kamaalg/azerbaijani-corpus-v0.alpaca-azerbaijani-gpt-4o-mini
Dataset Details
This is a translated version of the Alpaca dataset into the Azerbaijani language, using GPT-4o-mini.
Finance-Instruct-AzerbaijaniThis is part of a translated version of the original dataset: https://huggingface.co/datasets/Josephgflowers/Finance-Instruct-500k
medical-o1-reasoning-SFT-azerbaijaniThis is a translated version of the dataset https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT
internet_archive_azerbaijanialpaca_cleaned_azerbaijaniThis is translated into Azerbaijani Alpaca-Cleaned dataset.
