RichNachos/georgian-corpus
Georgian Corpus Dataset The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository. This dataset was developed as part of a bachelor’s project by: Georgi Kldiashvili Luka Paichadze Saba Shoshiashvili Dataset Summary Content: High-quality text in Georgian, suitable for NLP tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/RichNachos/georgian-corpus.
This repository belongs to RichNachos on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
