CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mteb /ArmenianParaphrasePC ArmenianParaphrasePC An MTEB dataset Massive Text Embedding Benchmark asparius/Armenian-Paraphrase-PC Task category t2t Domains News, Written Reference https://github.com/ivannikov-lab/arpa-paraphrase-corpus How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["ArmenianParaphrasePC"]) evaluator = mteb.MTEB(task) model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ArmenianParaphrasePC.texttext-classificationn<1K0 likes553 downloads1y agoHugging Face02Chillarmo /common_voice_20_armenian Common Voice 20 - Armenian This dataset is the Armenian portion of Mozilla's Common Voice 20.0 release, a massively multilingual collection of transcribed speech intended for speech technology research and development. Dataset Details Language: Armenian (hy) Source: Mozilla Common Voice Version: 20.0 License: CC0-1.0 audioautomatic-speech-recognition10K<n<100K1 likes352 downloads1y agoHugging Face03Karavet /ARPA-Armenian-Paraphrase-Corpus Dataset Description We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language. Dataset Summary The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.text1K<n<10K3 likes282 downloads4y agoHugging Face04Karavet /pioNER-Armenian-Named-Entity pioNER - named entity annotated datasets pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language. Alongside the datasets, we release 50-, 100-, 200-, and 300-dimensional GloVe word embeddings trained on a collection of Armenian texts from Wikipedia, news, blogs, and encyclopedia. Silver-standard dataset The generated corpus is automatically extracted and annotated using Armenian Wikipedia. We used a modification of… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/pioNER-Armenian-Named-Entity.2 likes155 downloads4y agoHugging Face05tetrak /armenian-ocr-crops Tetrak Armenian OCR crops Training data for tetrak_hy, the Armenian text recogniser we are building as an EasyOCR custom model in tetrak-hy-trainer for Tetrak, an OCR pipeline for community archives. The dataset has three configurations: corpus — 1,190 proofread pages of the Armenian Soviet Encyclopedia, as plain text with full Wikisource provenance. crops — the v0 synthetic pre-training set: 181,800 rendered word crops with transcriptions. crops-v1 — the v1 synthetic training… See the full description on the dataset page: https://huggingface.co/datasets/tetrak/armenian-ocr-crops.imageimage-to-text100K<n<1M1 likes123 downloads26d agoHugging Face06albertgrigoryan /eastern_armenian_tigran_nune2 likes87 downloads2y agoHugging Face07nomikos-project /armenian-manuscript-htr Armenian Manuscript HTR (BnF Arménien 172 and UCLA Armenian MS 72) This release contains handwritten text recognition ground truth for two Armenian canon law manuscripts: 1,704 transcribed lines with line polygons and baselines across 34 pages. Page images ship for both manuscripts: the BnF manuscript under Gallica terms and the UCLA manuscript by decision of the NOMOS project. The language and the manuscripts Armenian is an Indo-European language, attested in… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/armenian-manuscript-htr.imageimage-to-text1K<n<10K0 likes55 downloads5d agoHugging Face08saillab /alpaca-armenian-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-armenian-cleaned.text10K<n<100K1 likes40 downloads2y agoHugging Face09shunyalabs /armenian-speech-datasetaudio1K<n<10K0 likes31 downloads1y agoHugging Face10arjunkl /Armeniandna0 likes30 downloads15d agoHugging Face11davit312 /Armenian-speech-hygreek-myths :: total length 404.67 mins -> 6 hrs 44 mins monte-cristo :: total length 185.39 mins -> 3 hrs 5 mins 0 likes29 downloads2y agoHugging Face12ArmGPT /gsm8k-armenian GSM8K Տվյալների շտեմարան (Հայերեն տարբերակ՝ թվային պատասխաններով) ԿԱՐԵՎՈՐ ԾԱՆՈՒՑՈՒՄ. Սույն տվյալների հավաքածուն հանդիսանում է OpenAI-ի հեղինակած GSM8K (Grade School Math 8K) բնօրինակ շտեմարանի հայերեն թարգմանությունը։ Բոլոր հեղինակային իրավունքները և բովանդակության սեփականությունը պատկանում են OpenAI-ին: Այս տարբերակը կազմվել է հայալեզու մոդելների արագ և արդյունավետ ստուգաչափման (benchmarking) նպատակով։ Տվյալները ներկայացված են հստակ կառուցվածքով, որտեղ յուրաքանչյուր հարցի դիմաց… See the full description on the dataset page: https://huggingface.co/datasets/ArmGPT/gsm8k-armenian.textquestion-answering1K<n<10K0 likes29 downloads8mo agoHugging Face13catherinearnett /classical_armenian_pd Classical Armenian Public Domain Literature This dataset consists of 102 Classical Armenian texts in the public domain, which were collected from the Eastern Armenian National Corpus. A list of the works is provided below. Full list of works List of Works Աբովյան Խաչատուր՝ Առաջին սերը (First Love by Khachatur Abovian) Աբովյան Խաչատուր՝ Պարապ վախտի խաղալիք (Idle Time Toy by Khachatur Abovian) Աբովյան Խաչատուր՝ Թուրքի աղջիկը (The Turkish Girl by Khachatur Abovian)… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/classical_armenian_pd.texttext-generationn<1K0 likes27 downloads6mo agoHugging Face14Narek889 /armenian_gold_dataset Armenian Gold Dataset This repository contains the Armenian Gold Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification. Dataset Description The Armenian Gold Dataset provides a clean, well-structured, and verified collection of Armenian text. It is designed to… See the full description on the dataset page: https://huggingface.co/datasets/Narek889/armenian_gold_dataset.texttranslation100K<n<1M1 likes26 downloads5mo agoHugging Face15sarahooker /armenian-news-samples armenian_news_samples A collection of text samples in Armenian, covering news-related topics such as politics, economics, legal issues, and sports. The dataset includes sentences from news reports, official statements, and religious references. It reflects current events and societal developments in Armenia and the surrounding region. This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. Quality of Remastered Dataset… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/armenian-news-samples.imagen<1K0 likes23 downloads7mo agoHugging Face16LenaVolkova /ArmenianAddresses Dataset Card for Armenian Address Extraction Dataset This dataset consists of 2,271 records containing raw Armenian utility maintenance/planned outage announcement texts paired with structured address components extracted from them. The raw texts primarily originate from public announcements (such as those by the Electric Networks of Armenia) notifying the public of scheduled maintenance, and the dataset breaks down these notices into granular, queryable geographic attributes.… See the full description on the dataset page: https://huggingface.co/datasets/LenaVolkova/ArmenianAddresses.texttoken-classification1K<n<10K0 likes23 downloads3mo agoHugging Face17asparius /Armenian-Paraphrase-PC Armenian Paraphrase Detection Corpus This data is orinally from https://github.com/ivannikov-lab/arpa-paraphrase-corpus BibTeX Citation If you use this dataset, please cite following paper: @misc{malajyan2020arpa, title={ARPA: Armenian Paraphrase Detection Corpus and Models}, author={Arthur Malajyan and Karen Avetisyan and Tsolak Ghukasyan}, year={2020}, eprint={2009.12615}, archivePrefix={arXiv}, primaryClass={cs.CL} }… See the full description on the dataset page: https://huggingface.co/datasets/asparius/Armenian-Paraphrase-PC.textn<1K1 likes22 downloads2y agoHugging Face18Speech-data /armenian-speech-dataset 🎧 Armenian Speech Dataset 📘 Overview The Armenian Speech Dataset is a high-quality speech audio dataset designed for building, training, and evaluating modern AI voice technologies. It provides structured audio data optimized for deep learning workflows in speech processing. The dataset includes 76 hours of audio data distributed across 558 files, delivered in MP3 and WAV formats, with a total size of 189 MB. This carefully curated audio dataset ensures balanced and… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/armenian-speech-dataset.audioautomatic-speech-recognitionn<1K0 likes18 downloads6mo agoHugging Face19k-mktr /western_armenian_hy Western Armenian NT Description The Western Armenian New Testament is a translation of the New Testament into Western Armenian, the dialect of the Armenian language spoken by the Armenian diaspora (originating from Ottoman Armenia). Following the Armenian Genocide of 1915, Western Armenian became primarily a diaspora language. This translation is a vital cultural and linguistic artifact for the Armenian community worldwide, preserving the biblical text in a… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/western_armenian_hy.tabular1K<n<10K0 likes18 downloads2mo agoHugging Face20saillab /alpaca_armenian_tacoThis repository contains the dataset used for the TaCo paper. The dataset follows the style outlined in the TaCo paper, as follows: { "instruction": "instruction in xx", "input": "input in xx", "output": "Instruction in English: instruction in en , Response in English: response in en , Response in xx: response in xx " } Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_armenian_taco.text10K<n<100K1 likes16 downloads2y agoHugging Face21Metric-AI /armenian-history-test-2025-1textn<1K0 likes14 downloads2y agoHugging Face22Metric-AI /armenian-history-test-2025-3textn<1K0 likes14 downloads2y agoHugging Face23andovirab /armenian_heritage_small_dataset Armenian Heritage Dataset This repository contains the Armenian Heritage Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification. Dataset Description The Armenian Heritage Dataset provides a clean, well-structured, and verified collection of Armenian text. It is… See the full description on the dataset page: https://huggingface.co/datasets/andovirab/armenian_heritage_small_dataset.textn<1K0 likes13 downloads5mo agoHugging Face24EdUarD0110 /daily_dialog_armenian DailyDialog Armenian (Seq2Seq Format) This dataset is a translated version of the DailyDialog dataset, where all dialog lines have been translated into Armenian. It is structured in Seq2Seq format, which makes it suitable for training conversational AI models, machine translation models, or general-purpose sequence-to-sequence learning systems. 📚 Dataset Description The original DailyDialog dataset consists of multi-turn dialogues on daily life topics. Each dialogue… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/daily_dialog_armenian.text10K<n<100K0 likes11 downloads1y agoHugging Face25Metric-AI /armenian-history-test-2025-2textn<1K0 likes10 downloads2y agoHugging Face26k-mktr /armenian_bible_hy Armenian Bible (Eastern) Description The Armenian translation of the Bible has a history dating back to the 5th century, when Mesrop Mashtots created the Armenian alphabet specifically to translate the Scriptures. This public domain edition represents the Eastern Armenian version, translated from the original Hebrew and Greek texts. Armenia was the first nation to adopt Christianity as a state religion (301 AD), and the Armenian Bible is a cornerstone of Armenian… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/armenian_bible_hy.tabular1K<n<10K0 likes10 downloads2mo agoHugging Face27Metric-AI /armenian-language-test-2025-2textn<1K0 likes9 downloads2y agoHugging Face28Codename324 /cv17-armenian-processed1K<n<10K0 likes8 downloads1y agoHugging Face29edisimon /armenian-clean-textgated Armenian Clean Corpus (pretraining + SFT bundle) Combined, deduplicated, cleaned Armenian text assembled for pretraining and supervised fine-tuning of small language models. Built via the pipeline at https://github.com/EdikSimonian/armenian-gpt: python 1_download.py # fetch sources python 2_prepare.py # clean + dedup + merge python 1_download.py --upload # push this bundle Contents corpus/clean_text.txt.zst zstd-compressed merged corpus… See the full description on the dataset page: https://huggingface.co/datasets/edisimon/armenian-clean-text.10B<n<100B1 likes8 downloads4mo agoHugging Face30Metric-AI /armenian-history-test-2025-4textn<1K0 likes7 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.