LVSTCK/macedonian-corpus-raw
Macedonian Corpus - Raw ๐ Key Highlights Size: 37.6 GB, Word Count: 3.53 billion Includes data from 10+ sources, including academic texts, public archives, and online resources. Minimal preprocessing applied. Examples include academic papers, books, scraped web content, and more. ๐ Overview Macedonian is widely recognized as a low-resource language in the field of NLP. Publicly available resources in Macedonian are extremely limited, and as farโฆ See the full description on the dataset page: https://huggingface.co/datasets/LVSTCK/macedonian-corpus-raw.
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload macedonian_corpus_raw.jsonl.gz with huggingface_hub
Delete macedonian_corpus_raw.jsonl.gz
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Rename corpus-mk.jsonl.gz to macedonian_corpus_raw.jsonl.gz
Update README.md
Update README.md
Upload corpus-mk.jsonl.gz with huggingface_hub
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Upload corpus-mk.jsonl.gz with huggingface_hub
Upload corpus-mk.jsonl.gz with huggingface_hub
Update README.md
Update README.md
Update README.md
initial commit
