CoolFace
Datasetpublic

Marmara-NLP/CSE4078S25_Grp6_Text_Classification

In this project, we aim to develop a text classification system to classify Turkish texts into specific categories. We are currently in the first phase of our project and in this context, we are researching and collecting Turkish text classification datasets from internet sources (HuggingFace, Kaggle, etc.). Dataset Statistics The combined dataset consists of a total of 1,227,879 instructions. The average length for each component is as follows: Instruction Length (in characters): 83.87 Input… See the full description on the dataset page: https://huggingface.co/datasets/Marmara-NLP/CSE4078S25_Grp6_Text_Classification.

sourceHugging Faceupdated 1y agoView on Hugging Face
0likes6downloads

Marmara-NLP/CSE4078S25_Grp6_Text_Classification · main · files are served by the source, never re-hosted here