CoolFace
Datasetpublic

SEACrowd/indolem_sentiment

IndoLEM (Indonesian Language Evaluation Montage) is a comprehensive Indonesian benchmark that comprises of seven tasks for the Indonesian language. This benchmark is categorized into three pillars of NLP tasks: morpho-syntax, semantics, and discourse. This dataset is based on binary classification (positive and negative), with distribution: * Train: 3638 sentences * Development: 399 sentences * Test: 1011 sentences The data is sourced from 1) Twitter [(Koto and Rahmaningtyas, 2017)](https://www.researchgate.net/publication/321757985_InSet_Lexicon_Evaluation_of_a_Word_List_for_Indonesian_Sentiment_Analysis_in_Microblogs) and 2) [hotel reviews](https://github.com/annisanurulazhar/absa-playground/). The experiment is based on 5-fold cross validation.

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes82downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
SEACrowd/indolem_sentiment · CoolFace