RichNachos/georgian-corpus
Georgian Corpus Dataset The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository. This dataset was developed as part of a bachelor’s project by: Georgi Kldiashvili Luka Paichadze Saba Shoshiashvili Dataset Summary Content: High-quality text in Georgian, suitable for NLP tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/RichNachos/georgian-corpus.
Update README.md
Upload Report.pdf
Update README.md
Create README.md
Upload folder using huggingface_hub
Upload folder using huggingface_hub
Upload folder using huggingface_hub
initial commit
