CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01indolem /IndoMMLU IndoMMLU Fajri Koto, Nurul Aisyah, Haonan Li, Timothy Baldwin 📄 Paper • 🏆 Leaderboard • 🤗 Dataset Introduction We introduce IndoMMLU, the first multi-task language understanding benchmark for Indonesian culture and languages, which consists of questions from primary school to university entrance exams in Indonesia. By employing professional teachers, we obtain 14,906 questions across 63 tasks and education… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoMMLU.question-answering10K<n<100K20 likes185 downloads3y agoHugging Face02indolem /IndoCareer Introduction IndoCareer is a dataset comprising 8,834 multiple-choice questions designed to evaluate performance in vocational and professional certification exams across various fields. With a focus on Indonesia, IndoCareer provides rich local contexts, spanning six key sectors: (1) healthcare, (2) insurance and finance, (3) creative and design, (4) tourism and hospitality, (5) education and training, and (6) law. Data Each question in the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/indolem/IndoCareer.tabularquestion-answering10K<n<100K5 likes120 downloads2y agoHugging Face03SEACrowd /indolem_ntpNTP (Next Tweet prediction) is one of the comprehensive Indonesian benchmarks that given a list of tweets and an option, we predict if the option is the next tweet or not. This task is similar to the next sentence prediction (NSP) task used to train BERT (Devlin et al., 2019). In NTP, each instance consists of a Twitter thread (containing 2 to 4 tweets) that we call the premise, and four possible options for the next tweet, one of which is the actual response from the original thread. Train: 5681 threads Development: 811 threads Test: 1890 threads0 likes95 downloads2y agoHugging Face04SEACrowd /indolem_ner_ugmNER UGM is a Named Entity Recognition dataset that comprises 2,343 sentences from news articles, and was constructed at the University of Gajah Mada based on five named entity classes: person, organization, location, time, and quantity.0 likes82 downloads2y agoHugging Face05SEACrowd /indolem_sentimentIndoLEM (Indonesian Language Evaluation Montage) is a comprehensive Indonesian benchmark that comprises of seven tasks for the Indonesian language. This benchmark is categorized into three pillars of NLP tasks: morpho-syntax, semantics, and discourse. This dataset is based on binary classification (positive and negative), with distribution: * Train: 3638 sentences * Development: 399 sentences * Test: 1011 sentences The data is sourced from 1) Twitter [(Koto and Rahmaningtyas, 2017)](https://www.researchgate.net/publication/321757985_InSet_Lexicon_Evaluation_of_a_Word_List_for_Indonesian_Sentiment_Analysis_in_Microblogs) and 2) [hotel reviews](https://github.com/annisanurulazhar/absa-playground/). The experiment is based on 5-fold cross validation.0 likes81 downloads2y agoHugging Face06SEACrowd /indolem_neruiNER UI is a Named Entity Recognition dataset that contains 2,125 sentences obtained via an annotation assignment in an NLP course at the University of Indonesia in 2016. The corpus has three named entity classes: location, organisation, and person with training/dev/test distribution: 1,530/170/42 and based on 5-fold cross validation.1 likes69 downloads2y agoHugging Face07indolem /indo_story_cloze IndoCloze About We hired seven Indonesian university students to each write 500 short stories over a period of one month. This paper wins Best Paper Award at CSRR (ACL 2022). Paper Fajri Koto, Timothy Baldwin, and Jey Han Lau. Cloze Evaluation for Deeper Understanding of Commonsense Stories in Indonesian. In In Proceedings of Commonsense Representation and Reasoning Workshop 2022 (CSRR at ACL 2022), Dublin, Ireland. Dataset A story in our dataset… See the full description on the dataset page: https://huggingface.co/datasets/indolem/indo_story_cloze.3 likes68 downloads3y agoHugging Face08SEACrowd /indolem_tweet_orderingIndoLEM (Indonesian Language Evaluation Montage) is a comprehensive Indonesian benchmark that comprises of seven tasks for the Indonesian language. This benchmark is categorized into three pillars of NLP tasks: morpho-syntax, semantics, and discourse. This task is based on the sentence ordering task of Barzilay and Lapata (2008) to assess text relatedness. We construct the data by shuffling Twitter threads (containing 3 to 5 tweets), and assessing the predicted ordering in terms of rank correlation (p) with the original. The experiment is based on 5-fold cross validation. Train: 4327 threads Development: 760 threads Test: 1521 threads0 likes47 downloads2y agoHugging Face09SEACrowd /indolem_ud_id_pud1 of 8 sub-datasets of IndoLEM, a comprehensive dataset encompassing 7 NLP tasks (Koto et al., 2020). This dataset is part of [Parallel Universal Dependencies (PUD)](http://universaldependencies.org/conll17/) project. This is based on the first corrected version by Alfina et al. (2019), contains 1,000 sentences.0 likes41 downloads2y agoHugging Face10SEACrowd /indolem_ud_id_gsdThe Indonesian-GSD treebank consists of 5598 sentences and 122k words split into train/dev/test of 97k/12k/11k words. The treebank was originally converted from the content head version of the universal dependency treebank v2.0 (legacy) in 2015.In order to comply with the latest Indonesian annotation guidelines, the treebank has undergone a major revision between UD releases v2.8 and v2.9 (2021).0 likes40 downloads2y agoHugging Face11indolem /IndoCulturetext1K<n<10K11 likes40 downloads2y agoHugging Face12treamyracle /indolem-ugm-nertext1K<n<10K0 likes36 downloads2mo agoHugging Face13treamyracle /indolem-ui-nertext1K<n<10K0 likes20 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.