CoolFace
20 results

azerbaijan

BHOSAI /QA_3sualaz_on_Azerbaijani Question-Answering Dataset for Azerbaijani Language based on Intellectual Games (3sual.az) Baku Higher Oil School Research and Development Center on AI introduces a dataset to fine-tune the NLP models to manage it as a question answering. This dataset contains 4697 questions with answers and explanations. In some cases answer does not exist therefore that slot is empty. Dataset have been collected from 3sual.az and copyright belongs to corresponding website (3sual.az) and its owner Bahruz… See the full description on the dataset page: https://huggingface.co/datasets/BHOSAI/QA_3sualaz_on_Azerbaijani.question-answering1K<n<10K1 likes3k downloads2y agoHugging Faceendomorphosis /ipfs_azerbaijan_laws Laws of Azerbaijan Research snapshot of official legislation collected from e-qanun.az downloadDetailPdf. Not legal advice. Official gazettes / government portals prevail over this corpus. Snapshot Field Value Snapshot date 2026-09-22 Coverage catalog-backed incomplete Source e-qanun.az downloadDetailPdf Collector scrapers/collect_az.py Laws / instruments 7101 Articles 6241 Language az Jurisdiction Azerbaijan License az-eqanun… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_azerbaijan_laws.texttext-retrieval10K<n<100K0 likes2.7k downloads2d agoHugging FaceLocalDoc /azerbaijani_asr Azerbaijani ASR Dataset Dataset Description This dataset contains Azerbaijani speech data for Automatic Speech Recognition (ASR) tasks. Dataset Summary Language: Azerbaijani (az) Task: Automatic Speech Recognition Total Duration: ~328 hours Total Samples: ~345,643 audio-text pairs Audio Format: WAV, 16kHz sampling rate License: CC-BY-4.0 Dataset Structure Each audio segment is specially numbered so that you can merge them if you… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_asr.audioautomatic-speech-recognition100K<n<1M4 likes671 downloads2mo agoHugging Faceismatsamadov /azerbaijan-court-data Azerbaijan Court System Dataset The most comprehensive open dataset of Azerbaijan's judicial system — 1.64 million structured records and 1.54 million court decision PDFs (~160 GB) covering court decisions, active cases, scheduled hearings, court registries, judges, lawyers, and mediator organizations. Built for AI engineers, legal tech startups, and researchers who need real-world legal data at scale. Quick Start Load with Hugging Face datasets from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ismatsamadov/azerbaijan-court-data.imagetext-classification1M<n<10M2 likes413 downloads6mo agoHugging Facejusticedao /ipfs_azerbaijan_laws_ir Azerbaijan legislation IR (CID-keyed sparse GraphRAG) Research retrieval release of endomorphosis/ipfs_azerbaijan_laws (revision 01e4eb269e4de3aa301daaa5ca725c7de218b3b6) packaged as country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir). Not legal advice. This is a research snapshot. The official gazette / authentic source of Azerbaijan prevails over this corpus. Retrieved documents and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_azerbaijan_laws_ir.tabulartext-retrieval100K<n<1M0 likes406 downloads9h agoHugging FaceLocalDoc /azerbaijani-pretrain-corpus Azerbaijani Pretraining Corpus (merged & deduplicated) A cleaned Azerbaijani text corpus assembled for language-model pretraining, merging two curated sources and removing exact duplicates. Contents Documents: 6,931,898 Tokens: ~5.36B (measured with the o200k_base tokenizer; an Azerbaijani-specific tokenizer will yield fewer tokens, as o200k_base segments agglutinative Azerbaijani inefficiently) Avg tokens/document: ~773 Fields text — the… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani-pretrain-corpus.texttext-generation1M<n<10M0 likes380 downloads4mo agoHugging Face