CoolFace
Datasetpublic

LVSTCK/macedonian-corpus-raw

Macedonian Corpus - Raw ๐ŸŒŸ Key Highlights Size: 37.6 GB, Word Count: 3.53 billion Includes data from 10+ sources, including academic texts, public archives, and online resources. Minimal preprocessing applied. Examples include academic papers, books, scraped web content, and more. ๐Ÿ“‹ Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as farโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes45downloads
settings

This repository belongs to LVSTCK on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemacedonian-corpus-raw
visibilitypublic
licencecc-by-4.0
gatedno
ownerLVSTCK
Account settings
LVSTCK/macedonian-corpus-raw ยท CoolFace