andovirab/armenian_heritage_small_dataset
Armenian Heritage Dataset This repository contains the Armenian Heritage Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification. Dataset Description The Armenian Heritage Dataset provides a clean, well-structured, and verified collection of Armenian text.… See the full description on the dataset page: https://huggingface.co/datasets/andovirab/armenian_heritage_small_dataset.
Armenian Heritage Dataset
This repository contains the Armenian Heritage Dataset, a high-quality, curated dataset designed for Armenian Natural Language Processing (NLP) tasks. It serves as a benchmark and training resource for various downstream applications, including text generation, masked language modeling, and token classification.
Dataset Description
The Armenian Heritage Dataset provides a clean, well-structured, and verified collection of Armenian text. It is designed to advance research and development for the Armenian language in AI.
- Curator: [Andranik Virabyan]
- Language: Armenian (
hy), English (en) ---
Uses
Direct Use
- Fine-tuning large language models (LLMs) on high-quality Armenian text.
- Masked language modeling (MLM) for Armenian BERT-style models.
- Benchmarking spelling correction, text classification, and linguistic tasks.
Out-of-Scope Use
- Translating without proper contextual verification.
- Any usage that violates the terms of the CC BY-SA 4.0 license.
Dataset Structure
The dataset is provided in standard formats (e.g., CSV, JSON, or Parquet) and contains the following main features:
- `text`: The primary text content in the Armenian language.
- `id` (if applicable): A unique identifier for each entry.
- `source` (if applicable): The origin or category of the specific text entry.
Licensing & Attribution
This dataset is licensed under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) License.
You are free to:
- Share — Copy and redistribute the material in any medium or format.
- Adapt — Remix, transform, and build upon the material for any purpose, even commercially.
Under the following terms:
- Attribution — You must give appropriate credit, provide a link to the dataset, and indicate if changes were made.
- ShareAlike — If you remix, transform, or build upon the material, you must distribute your contributions under the same license.
