datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ArmenianParaphrasePC
ArmenianParaphrasePC
An MTEB dataset
Massive Text Embedding Benchmark
asparius/Armenian-Paraphrase-PC
Task category
t2t
Domains
News, Written
Reference
https://github.com/ivannikov-lab/arpa-paraphrase-corpus
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["ArmenianParaphrasePC"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ArmenianParaphrasePC.common_voice_20_armenian
Common Voice 20 - Armenian
This dataset is the Armenian portion of Mozilla's Common Voice 20.0 release,
a massively multilingual collection of transcribed speech intended for speech technology research and development.
Dataset Details
Language: Armenian (hy)
Source: Mozilla Common Voice
Version: 20.0
License: CC0-1.0
ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.pioNER-Armenian-Named-Entity
pioNER - named entity annotated datasets
pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
Alongside the datasets, we release 50-, 100-, 200-, and 300-dimensional GloVe word embeddings trained on a collection of Armenian texts from Wikipedia, news, blogs, and encyclopedia.
Silver-standard dataset
The generated corpus is automatically extracted and annotated using Armenian Wikipedia. We used a modification of… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/pioNER-Armenian-Named-Entity.armenian-ocr-crops
Tetrak Armenian OCR crops
Training data for tetrak_hy, the Armenian text recogniser we are
building as an EasyOCR custom model in
tetrak-hy-trainer
for Tetrak, an OCR pipeline for community
archives.
The dataset has three configurations:
corpus — 1,190 proofread pages of the Armenian Soviet
Encyclopedia, as plain text with full Wikisource provenance.
crops — the v0 synthetic pre-training set: 181,800 rendered
word crops with transcriptions.
crops-v1 — the v1 synthetic training… See the full description on the dataset page: https://huggingface.co/datasets/tetrak/armenian-ocr-crops.eastern_armenian_tigran_nunearmenian-manuscript-htr
Armenian Manuscript HTR (BnF Arménien 172 and UCLA Armenian MS 72)
This release contains handwritten text recognition ground truth for two Armenian canon law manuscripts: 1,704 transcribed lines with line polygons and baselines across 34 pages. Page images ship for both manuscripts: the BnF manuscript under Gallica terms and the UCLA manuscript by decision of the NOMOS project.
The language and the manuscripts
Armenian is an Indo-European language, attested in… See the full description on the dataset page: https://huggingface.co/datasets/nomikos-project/armenian-manuscript-htr.alpaca-armenian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-armenian-cleaned.armenian-speech-datasetArmeniandnaArmenian-speech-hygreek-myths ::
total length 404.67 mins -> 6 hrs 44 mins
monte-cristo ::
total length 185.39 mins -> 3 hrs 5 mins
gsm8k-armenian
GSM8K Տվյալների շտեմարան (Հայերեն տարբերակ՝ թվային պատասխաններով)
ԿԱՐԵՎՈՐ ԾԱՆՈՒՑՈՒՄ. Սույն տվյալների հավաքածուն հանդիսանում է OpenAI-ի հեղինակած GSM8K (Grade School Math 8K) բնօրինակ շտեմարանի հայերեն թարգմանությունը։ Բոլոր հեղինակային իրավունքները և բովանդակության սեփականությունը պատկանում են OpenAI-ին:
Այս տարբերակը կազմվել է հայալեզու մոդելների արագ և արդյունավետ ստուգաչափման (benchmarking) նպատակով։ Տվյալները ներկայացված են հստակ կառուցվածքով, որտեղ յուրաքանչյուր հարցի դիմաց… See the full description on the dataset page: https://huggingface.co/datasets/ArmGPT/gsm8k-armenian.classical_armenian_pd
Classical Armenian Public Domain Literature
This dataset consists of 102 Classical Armenian texts in the public domain, which were collected from the Eastern Armenian National Corpus.
A list of the works is provided below.
Full list of works
List of Works
Աբովյան Խաչատուր՝ Առաջին սերը (First Love by Khachatur Abovian)
Աբովյան Խաչատուր՝ Պարապ վախտի խաղալիք (Idle Time Toy by Khachatur Abovian)
Աբովյան Խաչատուր՝ Թուրքի աղջիկը (The Turkish Girl by Khachatur Abovian)… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/classical_armenian_pd.armenian_gold_dataset
Armenian Gold Dataset
This repository contains the Armenian Gold Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification.
Dataset Description
The Armenian Gold Dataset provides a clean, well-structured, and verified collection of Armenian text. It is designed to… See the full description on the dataset page: https://huggingface.co/datasets/Narek889/armenian_gold_dataset.armenian-news-samples
armenian_news_samples
A collection of text samples in Armenian, covering news-related topics such as politics, economics, legal issues, and sports. The dataset includes sentences from news reports, official statements, and religious references. It reflects current events and societal developments in Armenia and the surrounding region.
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
Quality of Remastered Dataset… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/armenian-news-samples.ArmenianAddresses
Dataset Card for Armenian Address Extraction Dataset
This dataset consists of 2,271 records containing raw Armenian utility maintenance/planned outage announcement texts paired with structured address components extracted from them. The raw texts primarily originate from public announcements (such as those by the Electric Networks of Armenia) notifying the public of scheduled maintenance, and the dataset breaks down these notices into granular, queryable geographic attributes.… See the full description on the dataset page: https://huggingface.co/datasets/LenaVolkova/ArmenianAddresses.Armenian-Paraphrase-PC
Armenian Paraphrase Detection Corpus
This data is orinally from https://github.com/ivannikov-lab/arpa-paraphrase-corpus
BibTeX Citation
If you use this dataset, please cite following paper:
@misc{malajyan2020arpa,
title={ARPA: Armenian Paraphrase Detection Corpus and Models},
author={Arthur Malajyan and Karen Avetisyan and Tsolak Ghukasyan},
year={2020},
eprint={2009.12615},
archivePrefix={arXiv},
primaryClass={cs.CL}
}… See the full description on the dataset page: https://huggingface.co/datasets/asparius/Armenian-Paraphrase-PC.armenian-speech-dataset
🎧 Armenian Speech Dataset
📘 Overview
The Armenian Speech Dataset is a high-quality speech audio dataset designed for building, training, and evaluating modern AI voice technologies. It provides structured audio data optimized for deep learning workflows in speech processing. The dataset includes 76 hours of audio data distributed across 558 files, delivered in MP3 and WAV formats, with a total size of 189 MB.
This carefully curated audio dataset ensures balanced and… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/armenian-speech-dataset.western_armenian_hy
Western Armenian NT
Description
The Western Armenian New Testament is a translation of the New Testament into Western Armenian, the dialect of the Armenian language spoken by the Armenian diaspora (originating from Ottoman Armenia). Following the Armenian Genocide of 1915, Western Armenian became primarily a diaspora language. This translation is a vital cultural and linguistic artifact for the Armenian community worldwide, preserving the biblical text in a… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/western_armenian_hy.alpaca_armenian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_armenian_taco.armenian-history-test-2025-1armenian-history-test-2025-3armenian_heritage_small_dataset
Armenian Heritage Dataset
This repository contains the Armenian Heritage Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification.
Dataset Description
The Armenian Heritage Dataset provides a clean, well-structured, and verified collection of Armenian text. It is… See the full description on the dataset page: https://huggingface.co/datasets/andovirab/armenian_heritage_small_dataset.daily_dialog_armenian
DailyDialog Armenian (Seq2Seq Format)
This dataset is a translated version of the DailyDialog dataset, where all dialog lines have been translated into Armenian. It is structured in Seq2Seq format, which makes it suitable for training conversational AI models, machine translation models, or general-purpose sequence-to-sequence learning systems.
📚 Dataset Description
The original DailyDialog dataset consists of multi-turn dialogues on daily life topics. Each dialogue… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/daily_dialog_armenian.armenian-history-test-2025-2armenian_bible_hy
Armenian Bible (Eastern)
Description
The Armenian translation of the Bible has a history dating back to the 5th century, when Mesrop Mashtots created the Armenian alphabet specifically to translate the Scriptures. This public domain edition represents the Eastern Armenian version, translated from the original Hebrew and Greek texts. Armenia was the first nation to adopt Christianity as a state religion (301 AD), and the Armenian Bible is a cornerstone of Armenian… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/armenian_bible_hy.armenian-language-test-2025-2cv17-armenian-processedarmenian-clean-text
Armenian Clean Corpus (pretraining + SFT bundle)
Combined, deduplicated, cleaned Armenian text assembled for pretraining
and supervised fine-tuning of small language models. Built via the
pipeline at https://github.com/EdikSimonian/armenian-gpt:
python 1_download.py # fetch sources
python 2_prepare.py # clean + dedup + merge
python 1_download.py --upload # push this bundle
Contents
corpus/clean_text.txt.zst zstd-compressed merged corpus… See the full description on the dataset page: https://huggingface.co/datasets/edisimon/armenian-clean-text.armenian-history-test-2025-4
