edisimon/armenian-clean-text
Armenian Clean Corpus (pretraining + SFT bundle) Combined, deduplicated, cleaned Armenian text assembled for pretraining and supervised fine-tuning of small language models. Built via the pipeline at https://github.com/EdikSimonian/armenian-gpt: python 1_download.py # fetch sources python 2_prepare.py # clean + dedup + merge python 1_download.py --upload # push this bundle Contents corpus/clean_text.txt.zst zstd-compressed merged corpus… See the full description on the dataset page: https://huggingface.co/datasets/edisimon/armenian-clean-text.
18
No commit history came back for main. The revision may not exist, or the source declined the request.
