RichNachos/georgian-corpus
Georgian Corpus Dataset The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository. This dataset was developed as part of a bachelor’s project by: Georgi Kldiashvili Luka Paichadze Saba Shoshiashvili Dataset Summary Content: High-quality text in Georgian, suitable for NLP tasks like text… See the full description on the dataset page: https://huggingface.co/datasets/RichNachos/georgian-corpus.
Georgian Corpus Dataset
The Georgian Corpus Dataset is an open-source dataset designed to advance the NLP community in Georgia. It contains 5 million rows of filtered, cleaned, and deduplicated text data extracted from the Common Crawl repository.
This dataset was developed as part of a bachelor’s project by:
Dataset Summary
- Content: High-quality text in Georgian, suitable for NLP tasks like text classification, translation, and language modeling.
- Source: Extracted from Common Crawl, then processed to ensure data quality through filtering, cleaning, and deduplication.
- Size: Slightly more than 5 million rows.
Motivation
We created this dataset to address the limited availability of Georgian-language datasets and to contribute to the global and local open-source NLP community.
How to Use
The dataset can be used for a variety of tasks, including:
- Language Modeling: Pretraining or fine-tuning Georgian language models.
- Machine Translation: Building models for translation between Georgian and other languages.
- Text Classification: Training classifiers for sentiment analysis or topic categorization.
For usage examples and additional details, please refer to the project report.
License
This dataset is distributed under the GPL-3.0 License.
Citation
If you use this dataset in your research or projects, please cite us.
