CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BHOSAI /QA_3sualaz_on_Azerbaijani Question-Answering Dataset for Azerbaijani Language based on Intellectual Games (3sual.az) Baku Higher Oil School Research and Development Center on AI introduces a dataset to fine-tune the NLP models to manage it as a question answering. This dataset contains 4697 questions with answers and explanations. In some cases answer does not exist therefore that slot is empty. Dataset have been collected from 3sual.az and copyright belongs to corresponding website (3sual.az) and its owner Bahruz… See the full description on the dataset page: https://huggingface.co/datasets/BHOSAI/QA_3sualaz_on_Azerbaijani.question-answering1K<n<10K1 likes3k downloads2y agoHugging Face02LocalDoc /azerbaijani_asr Azerbaijani ASR Dataset Dataset Description This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks. Dataset Summary Language: Azerbaijani (az) Task: Automatic Speech Recognition Total Duration: ~328 hours Total Samples: ~345,643 audio-text pairs Audio Format: WAV, 16kHz sampling rate License: CC-BY-4.0 Dataset Structure Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.audioautomatic-speech-recognition100K<n<1M4 likes755 downloads2mo agoHugging Face03shunyalabs /azerbaijani-speech-datasetaudio100K<n<1M1 likes428 downloads1y agoHugging Face04LocalDoc /azerbaijani-pretrain-corpus Azerbaijani Pretraining Corpus (merged & deduplicated) A cleaned Azerbaijani text corpus assembled for language-model pretraining, merging two curated sources and removing exact duplicates. Contents Documents: 6,931,898 Tokens: ~5.36B (measured with the o200k_base tokenizer; an Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base segments agglutinative Azerbaijani inefficiently) Avg tokens/document: ~773 Fields text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.texttext-generation1M<n<10M0 likes385 downloads4mo agoHugging Face05LocalDoc /azerbaijani_retriever_corpus A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training Dataset Description This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents. The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.tabularsentence-similarity100K<n<1M0 likes383 downloads1y agoHugging Face06LocalDoc /court_cases_azerbaijani Court Cases Of The Republic Of Azerbaijan This dataset consists of court cases from the Republic of Azerbaijan. Overview It was formed based on 1,200,000 court cases. The data has been preliminarily normalized and split into sentences. The dataset consists of 37 million sentences and approximately 500-600 million tokens. Dataset Structure Each row represents a single sentence extracted from a court case document. Column Type Description case_id… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/court_cases_azerbaijani.text10M<n<100M4 likes148 downloads6mo agoHugging Face07LocalDoc /azerbaijani-ocr-linesgated Azerbaijani OCR Lines Line-level training data for Azerbaijani text recognition, in Latin and Cyrillic script, extracted from scanned books. Fields field description image cropped text line, grayscale, height 48 px text transcription script az_latin or az_cyrillic book anonymised source-book id How it was built Pages come from scanned PDFs that already carried an OCR text layer. Line boxes were taken from that layer, rendered… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-lines.imageimage-to-text100K<n<1M1 likes122 downloads16d agoHugging Face08ughurabbasov /azerbaijani-tts-datasetaudio10K<n<100K1 likes121 downloads7mo agoHugging Face09LocalDoc /azerbaijani-ocr-benchmark Azerbaijani OCR Benchmark Line-level OCR benchmark for Azerbaijani in both Latin and Cyrillic script, built from scanned books. Fields field description image cropped text line, grayscale, height 48 px text verbatim transcription script az_latin or az_cyrillic book anonymised source-book id How labels were produced Every line carries a label agreed on independently by three sources: the OCR text layer already present in the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-ocr-benchmark.imageimage-to-text1K<n<10K1 likes107 downloads20d agoHugging Face10LocalDoc /azerbaijani-htr-synthetic Azerbaijani Synthetic Handwritten OCR Dataset A large-scale synthetic dataset for training handwritten text recognition (HTR) models on Azerbaijani Latin script. Generated using a procedural pipeline that combines real-world handwriting fonts with realistic scan-style augmentations. This dataset addresses the lack of publicly available Azerbaijani handwriting OCR data — a low-resource language for which no IAM-equivalent corpus exists. Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-htr-synthetic.imageimage-to-text1M<n<10M1 likes105 downloads5mo agoHugging Face11LocalDoc /fleurs-azerbaijani-asr FLEURS Azerbaijani ASR Benchmark Azerbaijani (az_az) subset of FLEURS, reformatted for ASR benchmarking and fine-tuning. Source Based on FLEURS dataset by Google (Conneau et al., 2022). Licensed under CC-BY-4.0. Structure Split Samples Duration train 2656 9.28h dev 400 1.35h test 921 3.23h Fields audio — 16kHz mono WAV sentence — transcription (original casing and punctuation) sentence_normalized — normalized (lowercase, no… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/fleurs-azerbaijani-asr.audioautomatic-speech-recognition1K<n<10K0 likes92 downloads6mo agoHugging Face12LocalDoc /community_oscar_azerbaijani Community-OSCAR Azerbaijani This is Azerbaijani version Community OSCAR dataset https://huggingface.co/datasets/oscar-corpus/community-oscar. Dataset Statistics (Aggregate) Metric Value Language Azerbaijani (az) Average per release 3.36 GiB, 603,832 documents Words per release ~408.8M words Characters per release ~3.12B characters Total size (all releases) 137.62 GiB Total lines 24.76M Total words 16.76B words Total characters 128.07B characters… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani.tabulartext-generation10M<n<100M0 likes91 downloads11mo agoHugging Face13tahmaz /azerbaijani-asr-zenfira Dataset Card for "azerbaijani-asr-zenfira" More Information needed audio10K<n<100K2 likes77 downloads7mo agoHugging Face14omar07ibrahim /alpaca-cleaned_AZERBAIJANItext10K<n<100K14 likes74 downloads3y agoHugging Face15bahramhasanov /azerbaijani-carpet-fullsizeimagen<1K0 likes68 downloads2y agoHugging Face16LocalDoc /azerbaijani-english-parallel-corpus Azerbaijani English Parallel Corpus This dataset contains 4,141,966 pairs of high-quality sentences translated from Azerbaijani to English. The data was collected from various resources such as websites, news, books, wikipedia, legislation, scientific articles and etc. License CC-BY-4.0 Contact For more information, questions, or issues, please contact LocalDoc at [v.resad.89@gmail.com]. texttranslation1M<n<10M1 likes66 downloads6mo agoHugging Face17LocalDoc /community_oscar_azerbaijani_scored Azerbaijani Web Corpus with Quality Scores This dataset is the full Azerbaijani web corpus LocalDoc/community_oscar_azerbaijani with a continuous quality score attached to every document. It is intended as the filtering layer for building a clean Azerbaijani pretraining corpus: each document carries a score that lets you keep, clean, or drop it according to your own thresholds. What was done Every document in the source corpus was scored by the model… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/community_oscar_azerbaijani_scored.texttext-classification10M<n<100M0 likes63 downloads4mo agoHugging Face18tahmaz /azerbaijani_asr_th_clean4audio100K<n<1M0 likes59 downloads9mo agoHugging Face19ARMammadli /azerbaijani-cuisine Azerbaijani Cuisine Dataset A curated image dataset of traditional Azerbaijani dishes for computer vision and image classification tasks. Dataset Description This dataset contains images of five traditional Azerbaijani dish categories. It is organized into standard training, validation, and test splits to facilitate machine learning model development and evaluation. Features 5 Food Categories: Dolma, Kebabs, Pakhlava, Plov, and Soups 324 Total Images: Properly… See the full description on the dataset page: https://huggingface.co/datasets/ARMammadli/azerbaijani-cuisine.imagen<1K1 likes56 downloads1y agoHugging Face20hajili /azerbaijani_review_sentiment_classificationAzerbaijani Sentiment Classification Dataset with ~160K reviews. Dataset contains 3 columns: Content, Score, Upvotes tabulartext-classification100K<n<1M6 likes53 downloads3y agoHugging Face21Kartal-Ol /English-Azerbaijani-Arabic-Script-Parallel-Corpustext1M<n<10M1 likes51 downloads2y agoHugging Face22tahmaz /azerbaijani-asr-news1audio10K<n<100K0 likes47 downloads1y agoHugging Face23eljanmahammadli /glue-mrpc-azerbaijaniThis dataset represents a translated version of the GLUE/MRPC dataset, generated using the Google Translate API. tabulartext-classification1K<n<10K0 likes46 downloads3y agoHugging Face24aznlp /azerbaijani-blogs Azerbaijani Blogs dataset Dataset Details Dataset Description This dataset provides blogs written in azerbaijani language with categories and tags for each. Language(s) (NLP): Azerbaijani License: Apache license 2.0 Data Source All the data was found in public resources of kayzen.az blogging website without any restriction. texttext-classification1K<n<10K3 likes44 downloads2y agoHugging Face25khazarai /AARA_Azerbaijani_LLM_Benchmark AARA: Azerbaijani Advanced Reasoning Assessment This dataset is the Azerbaijani-translated version of the emre/TARA_Turkish_LLM_Benchmark. textquestion-answeringn<1K1 likes43 downloads6mo agoHugging Face26LocalDoc /azerbaijani_books_retriever_corpus-reranked Azerbaijani Books Retrieval Dataset (Reranked) A large-scale retrieval dataset built from LocalDoc/books_dataset — a collection of 2,804 Azerbaijani-language books with 7.8M sentences spanning politics, history, literature, science, and more. Designed for training and evaluating information retrieval, semantic search, and RAG pipelines in Azerbaijani. Dataset Configs The dataset consists of three configs that can be joined via passage_id and query_id:… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_books_retriever_corpus-reranked.tabularsentence-similarity1M<n<10M0 likes43 downloads7mo agoHugging Face27DGurgurov /azerbaijani_sa Sentiment Analysis Data for the Azerbaijani Language Dataset Description: This dataset contains a sentiment analysis dataset from LocalDoc (2024). Data Structure: The data was used for the project on improving word embeddings with graph knowledge for Low Resource Languages. Citation: @source{azerbaijanisent, title={Sentiment Analysis Datset for Azerbaijani}, author={LocalDoc}, link={https://huggingface.co/LocalDoc}, year={2024} } texttext-classification10K<n<100K0 likes42 downloads2y agoHugging Face28tahmaz /azerbaijani_asr_th_clean6audio10K<n<100K0 likes41 downloads9mo agoHugging Face29asmarhajizada /azerbaijani-audiobooksaudio1K<n<10K0 likes41 downloads7mo agoHugging Face30regional122 /azerbaijani_spelling_dictionary_2021This dataset contains words from the Spelling Dictionary of the Azerbaijani Language, 7th edition, which was published in 2021. text10K<n<100K0 likes39 downloads21d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.