CoolFace
14 results

indolem

indolem /IndoMMLU IndoMMLU Fajri Koto, Nurul Aisyah, Haonan Li, Timothy Baldwin 📄 Paper • 🏆 Leaderboard • 🤗 Dataset Introduction We introduce IndoMMLU, the first multi-task language understanding benchmark for Indonesian culture and languages, which consists of questions from primary school to university entrance exams in Indonesia. By employing professional teachers, we obtain 14,906 questions across 63 tasks and education… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoMMLU.question-answering10K<n<100K20 likes185 downloads3y agoHugging Faceindolem /IndoCareer Introduction IndoCareer is a dataset comprising 8,834 multiple-choice questions designed to evaluate performance in vocational and professional certification exams across various fields. With a focus on Indonesia, IndoCareer provides rich local contexts, spanning six key sectors: (1) healthcare, (2) insurance and finance, (3) creative and design, (4) tourism and hospitality, (5) education and training, and (6) law. Data Each question in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoCareer.tabularquestion-answering10K<n<100K5 likes120 downloads2y agoHugging FaceSEACrowd /indolem_ntpNTP (Next Tweet prediction) is one of the comprehensive Indonesian benchmarks that given a list of tweets and an option, we predict if the option is the next tweet or not. This task is similar to the next sentence prediction (NSP) task used to train BERT (Devlin et al., 2019). In NTP, each instance consists of a Twitter thread (containing 2 to 4 tweets) that we call the premise, and four possible options for the next tweet, one of which is the actual response from the original thread. Train: 5681 threads Development: 811 threads Test: 1890 threads0 likes95 downloads2y agoHugging FaceSEACrowd /indolem_ner_ugmNER UGM is a Named Entity Recognition dataset that comprises 2,343 sentences from news articles, and was constructed at the University of Gajah Mada based on five named entity classes: person, organization, location, time, and quantity.0 likes82 downloads2y agoHugging FaceSEACrowd /indolem_sentimentIndoLEM (Indonesian Language Evaluation Montage) is a comprehensive Indonesian benchmark that comprises of seven tasks for the Indonesian language. This benchmark is categorized into three pillars of NLP tasks: morpho-syntax, semantics, and discourse. This dataset is based on binary classification (positive and negative), with distribution: * Train: 3638 sentences * Development: 399 sentences * Test: 1011 sentences The data is sourced from 1) Twitter [(Koto and Rahmaningtyas, 2017)](https://www.researchgate.net/publication/321757985_InSet_Lexicon_Evaluation_of_a_Word_List_for_Indonesian_Sentiment_Analysis_in_Microblogs) and 2) [hotel reviews](https://github.com/annisanurulazhar/absa-playground/). The experiment is based on 5-fold cross validation.0 likes81 downloads2y agoHugging FaceSEACrowd /indolem_neruiNER UI is a Named Entity Recognition dataset that contains 2,125 sentences obtained via an annotation assignment in an NLP course at the University of Indonesia in 2016. The corpus has three named entity classes: location, organisation, and person with training/dev/test distribution: 1,530/170/42 and based on 5-fold cross validation.1 likes69 downloads2y agoHugging Face