CoolFace
20 results

qana

qanastek /MASSIVEMASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations for the Natural Language Understanding tasks of intent prediction and slot annotation. Utterances span 60 intents and include 55 slot types. MASSIVE was created by localizing the SLURP dataset, composed of general Intelligent Voice Assistant single-shot interactions.text-classification100K<n<1M28 likes7.7k downloads4y agoHugging Faceqanastek /ELRC-Medical-V2 ELRC-Medical-V2 : European parallel corpus for healthcare machine translation Dataset Summary ELRC-Medical-V2 is a parallel corpus for neural machine translation funded by the European Commission and coordinated by the German Research Center for Artificial Intelligence. Supported Tasks and Leaderboards translation: The dataset can be used to train a model for translation. Languages In our case, the corpora consists of a pair of source and target… See the full description on the dataset page: https://huggingface.co/datasets/qanastek/ELRC-Medical-V2.texttranslation100K<n<1M17 likes551 downloads4y agoHugging FaceQanadil /ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets" Note About Sentiment_label_confidence "Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.tabulartext-classification10K<n<100K1 likes408 downloads2y agoHugging Faceqanastek /EMEA-V3 EMEA-V3 : European parallel translation corpus from the European Medicines Agency Dataset Summary EMEA-V3 is a parallel corpus for neural machine translation collected and aligned by Tiedemann, Jorg during the OPUS project. Supported Tasks and Leaderboards translation: The dataset can be used to train a model for translation. Languages In our case, the corpora consists of a pair of source and target sentences for all 22 different languages from the… See the full description on the dataset page: https://huggingface.co/datasets/qanastek/EMEA-V3.texttranslation10M<n<100M9 likes393 downloads4y agoHugging Faceqanastek /ECDC ECDC : An overview of the European Union's highly multilingual parallel corpora Dataset Summary In October 2012, the European Union (EU) agency 'European Centre for Disease Prevention and Control' (ECDC) released a translation memory (TM), i.e. a collection of sentences and their professionally produced translations, in twenty-five languages. The data gets distributed via the web pages of the EC's Joint Research Centre (JRC). Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/qanastek/ECDC.texttranslation10K<n<100K2 likes295 downloads4y agoHugging FaceQanadil /ASTD_Arabic_Sentiment_Tweets_Dataset Citation (https://aclanthology.org/D15-1299/) texttext-classification1K<n<10K0 likes293 downloads2y agoHugging Face