datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ARPA-Armenian-Paraphrase-Corpus
Dataset Description
We provide sentential paraphrase detection train, test datasets as well as BERT-based models for the Armenian language.
Dataset Summary
The sentences in the dataset are taken from Hetq and Panarmenian news articles. To generate paraphrase for the sentences, we used back translation from Armenian to English. We repeated the step twice, after which the generated paraphrases were manually reviewed. Invalid sentences were filtered out, while the rest were… See the full description on the dataset page: https://huggingface.co/datasets/Karavet/ARPA-Armenian-Paraphrase-Corpus.gsm8k-armenian
GSM8K Տվյալների շտեմարան (Հայերեն տարբերակ՝ թվային պատասխաններով)
ԿԱՐԵՎՈՐ ԾԱՆՈՒՑՈՒՄ. Սույն տվյալների հավաքածուն հանդիսանում է OpenAI-ի հեղինակած GSM8K (Grade School Math 8K) բնօրինակ շտեմարանի հայերեն թարգմանությունը։ Բոլոր հեղինակային իրավունքները և բովանդակության սեփականությունը պատկանում են OpenAI-ին:
Այս տարբերակը կազմվել է հայալեզու մոդելների արագ և արդյունավետ ստուգաչափման (benchmarking) նպատակով։ Տվյալները ներկայացված են հստակ կառուցվածքով, որտեղ յուրաքանչյուր հարցի դիմաց… See the full description on the dataset page: https://huggingface.co/datasets/ArmGPT/gsm8k-armenian.armenian-speech-dataset
🎧 Armenian Speech Dataset
📘 Overview
The Armenian Speech Dataset is a high-quality speech audio dataset designed for building, training, and evaluating modern AI voice technologies. It provides structured audio data optimized for deep learning workflows in speech processing. The dataset includes 76 hours of audio data distributed across 558 files, delivered in MP3 and WAV formats, with a total size of 189 MB.
This carefully curated audio dataset ensures balanced and… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/armenian-speech-dataset.daily_dialog_armenian
DailyDialog Armenian (Seq2Seq Format)
This dataset is a translated version of the DailyDialog dataset, where all dialog lines have been translated into Armenian. It is structured in Seq2Seq format, which makes it suitable for training conversational AI models, machine translation models, or general-purpose sequence-to-sequence learning systems.
📚 Dataset Description
The original DailyDialog dataset consists of multi-turn dialogues on daily life topics. Each dialogue… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/daily_dialog_armenian.Armenian_sentences_100000armenian_poems_dataset
Armenian 320 Poems Dataset 🇦🇲📜
This dataset contains 320 Armenian poems, collected and formatted for use in natural language processing, literary analysis, or educational purposes.
📑 Dataset Structure
The dataset is stored in a CSV format with 2 columns:
վերնագիր (Title)
poem (Poem Text)
Երգ եղբայրության
…
Իմ Հայաստան
…
Column Descriptions:
վերնագիր: The title of the poem in Armenian.
poem: The full text of the poem. Line breaks within… See the full description on the dataset page: https://huggingface.co/datasets/EdUarD0110/armenian_poems_dataset.
