CoolFace
Datasetpublic

LVSTCK/macedonian-corpus-raw

Macedonian Corpus - Raw 🌟 Key Highlights Size: 37.6 GB, Word Count: 3.53 billion Includes data from 10+ sources, including academic texts, public archives, and online resources. Minimal preprocessing applied. Examples include academic papers, books, scraped web content, and more. 📋 Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as far… See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes45downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
LVSTCK/macedonian-corpus-raw · CoolFace