datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
azerbaijani_retriever_corpus
A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training
Dataset Description
This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents.
The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.community_oscar_azerbaijani
Community-OSCAR Azerbaijani
This is Azerbaijani version Community OSCAR dataset https://huggingface.co/datasets/oscar-corpus/community-oscar.
Dataset Statistics (Aggregate)
Metric
Value
Language
Azerbaijani (az)
Average per release
3.36 GiB, 603,832 documents
Words per release
~408.8M words
Characters per release
~3.12B characters
Total size (all releases)
137.62 GiB
Total lines
24.76M
Total words
16.76B words
Total characters
128.07B characters… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani.azerbaijani_review_sentiment_classificationAzerbaijani Sentiment Classification Dataset with ~160K reviews.
Dataset contains 3 columns: Content, Score, Upvotes
glue-mrpc-azerbaijaniThis dataset represents a translated version of the GLUE/MRPC dataset, generated using the Google Translate API.
azerbaijani_books_retriever_corpus-reranked
Azerbaijani Books Retrieval Dataset (Reranked)
A large-scale retrieval dataset built from LocalDoc/books_dataset — a collection of 2,804 Azerbaijani-language books with 7.8M sentences spanning politics, history, literature, science, and more. Designed for training and evaluating information retrieval, semantic search, and RAG pipelines in Azerbaijani.
Dataset Configs
The dataset consists of three configs that can be joined via passage_id and query_id:… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_books_retriever_corpus-reranked.azerbaijani-text-quality-labeled
Azerbaijani Text Quality — Labeled Dataset
249,949 Azerbaijani web documents annotated with a quality score 0-3.
Used to train a document-level quality classifier for filtering a web
corpus before language-model pretraining.
Source and labeling
Texts: sampled from LocalDoc/community_oscar_azerbaijani,
an OSCAR-derived Common Crawl corpus. The texts are NOT original to this dataset.
Labels: generated by the LLM Mistral-Small-24B-Instruct-2501, not by humans.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-text-quality-labeled.hellaswag-azerbaijani
Dataset Details
This dataset is a translated version of Rowan/hellaswag into Azerbaijani. Only the ctx and endings columns have been translated, resulting in the new columns ctx_az and endings_az. The translation was done using Gemini Flash 2.0. A few samples (around 200) were removed due to errors arising from JSON parsing of the LLM response.
@inproceedings{zellers2019hellaswag,
title={HellaSwag: Can a Machine Really Finish Your Sentence?},
author={Zellers, Rowan and… See the full description on the dataset page: https://huggingface.co/datasets/eljanmahammadli/hellaswag-azerbaijani.Azerbaijani-sickr-stsmedical-thinking-azerbaijaniAzerbaijani-sts12-stsFinance-Instruct-AzerbaijaniThis is part of a translated version of the original dataset: https://huggingface.co/datasets/Josephgflowers/Finance-Instruct-500k
Azerbaijani-STSBenchmarkAzerbaijani-biosses-stsAzerbaijani-sts13-stsAzerbaijani-sts16-stsgpt-4o-mini_Azerbaijani_Hist_MCazerbaijani_retriever_corpus-reranked
Azerbaijan Legislation Retrieval Corpus — Reranked
Reranked version of LocalDoc/azerbaijani_retriever_corpus.
Hard negatives were re-scored with BAAI/bge-reranker-v2-m3 cross-encoder. False negatives (score > 95% of positive score) were filtered out. Remaining negatives are sorted by score descending (hardest first).
Configs
Config
Rows
Description
corpus
65,188
Passage chunks: chunk_id, passage
queries
188,941
Queries: query_id, chunk_id, query… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus-reranked.Azerbaijani-sts15-stsgpt-4o-mini_Azerbaijani_Lit_MCLlama-3.2-1B-Instruct-Q3_K_L_Azerbaijani_Hist_MCgpt-4o-mini_Azerbaijani_Lang_MCLlama-3.2-1B-Instruct-Q3_K_L_Azerbaijani_Lang_MCLlama-3.2-1B-Instruct-Q3_K_L_Azerbaijani_Lit_MC
