Armenian
Datasets
All datasets matching “Armenian”ArmenianParaphrasePC
ArmenianParaphrasePC
An MTEB dataset
Massive Text Embedding Benchmark
asparius/Armenian-Paraphrase-PC
Task category
t2t
Domains
News, Written
Reference
https://github.com/ivannikov-lab/arpa-paraphrase-corpus
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["ArmenianParaphrasePC"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ArmenianParaphrasePC.common_voice_20_armenian
Common Voice 20 - Armenian
This dataset is the Armenian portion of Mozilla's Common Voice 20.0 release,
a massively multilingual collection of transcribed speech intended for speech technology research and development.
Dataset Details
Language: Armenian (hy)
Source: Mozilla Common Voice
Version: 20.0
License: CC0-1.0
ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.pioNER-Armenian-Named-Entity
pioNER - named entity annotated datasets
pioNER corpus provides gold-standard and automatically generated named-entity datasets for the Armenian language.
Alongside the datasets, we release 50-, 100-, 200-, and 300-dimensional GloVe word embeddings trained on a collection of Armenian texts from Wikipedia, news, blogs, and encyclopedia.
Silver-standard dataset
The generated corpus is automatically extracted and annotated using Armenian Wikipedia. We used a modification of… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/pioNER-Armenian-Named-Entity.armenian-ocr-crops
Tetrak Armenian OCR crops
Training data for tetrak_hy, the Armenian text recogniser we are
building as an EasyOCR custom model in
tetrak-hy-trainer
for Tetrak, an OCR pipeline for community
archives.
The dataset has three configurations:
corpus — 1,190 proofread pages of the Armenian Soviet
Encyclopedia, as plain text with full Wikisource provenance.
crops — the v0 synthetic pre-training set: 181,800 rendered
word crops with transcriptions.
crops-v1 — the v1 synthetic training… See the full description on the dataset page: https://huggingface.co/datasets/tetrak/armenian-ocr-crops.eastern_armenian_tigran_nune
