CoolFace
Datasetpublic

andovirab/armenian_heritage_small_dataset

Armenian Heritage Dataset This repository contains the Armenian Heritage Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification. Dataset Description The Armenian Heritage Dataset provides a clean, well-structured, and verified collection of Armenian text.… See the full description on the dataset page: https://huggingface.co/datasets/andovirab/armenian_heritage_small_dataset.

sourceHugging Facecc-by-sa-4.0updated 5mo agoView on Hugging Face
0likes13downloads
Dataset Card

Armenian Heritage Dataset

This repository contains the Armenian Heritage Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification.


Dataset Description

The Armenian Heritage Dataset provides a clean, well-structured, and verified collection of Armenian text. It is designed to advance research and development for the Armenian language in AI.

  • —Curator: [Andranik Virabyan]
  • —Language: Armenian (hy), English (en) ---

Uses

Direct Use

  • —Fine-tuning large language models (LLMs) on high-quality Armenian text.
  • —Masked language modeling (MLM) for Armenian BERT-style models.
  • —Benchmarking spelling correction, text classification, and linguistic tasks.

Out-of-Scope Use

  • —Translating without proper contextual verification.
  • —Any usage that violates the terms of the CC BY-SA 4.0 license.

Dataset Structure

The dataset is provided in standard formats (e.g., CSV, JSON, or Parquet) and contains the following main features:

  • —`text`: The primary text content in the Armenian language.
  • —`id` (if applicable): A unique identifier for each entry.
  • —`source` (if applicable): The origin or category of the specific text entry.

Licensing & Attribution

This dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) License.

You are free to:

  • —Share — Copy and redistribute the material in any medium or format.
  • —Adapt — Remix, transform, and build upon the material for any purpose, even commercially.

Under the following terms:

  • —Attribution — You must give appropriate credit, provide a link to the dataset, and indicate if changes were made.
  • —ShareAlike — If you remix, transform, or build upon the material, you must distribute your contributions under the same license.