qana
Datasets
All datasets matching “qana”MASSIVEMASSIVE is a parallel dataset of > 1M utterances across 51 languages with annotations
for the Natural Language Understanding tasks of intent prediction and slot annotation.
Utterances span 60 intents and include 55 slot types. MASSIVE was created by localizing
the SLURP dataset, composed of general Intelligent Voice Assistant single-shot interactions.ELRC-Medical-V2
ELRC-Medical-V2 : European parallel corpus for healthcare machine translation
Dataset Summary
ELRC-Medical-V2 is a parallel corpus for neural machine translation funded by the European Commission and coordinated by the German Research Center for Artificial Intelligence.
Supported Tasks and Leaderboards
translation: The dataset can be used to train a model for translation.
Languages
In our case, the corpora consists of a pair of source and target… See the full description on the dataset page: https://huggingface.co/datasets/qanastek/ELRC-Medical-V2.ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets
ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets
Dataset Card for "ArSAS: An Arabic Speech-Act and Sentiment Corpus of Tweets"
Note About Sentiment_label_confidence
"Crowdflower provides a confidence score with each annotated tweet that represents the confidence in the quality of the label. For a three annotators per tweet setup, the confidence score would range between 0.3 and 1 according to two factors: 1) annotator quality level; and 2) agreement… See the full description on the dataset page: https://huggingface.co/datasets/Qanadil/ArSAS_An_Arabic_Speech-Act_and_Sentiment_Corpus_of_Tweets.EMEA-V3
EMEA-V3 : European parallel translation corpus from the European Medicines Agency
Dataset Summary
EMEA-V3 is a parallel corpus for neural machine translation collected and aligned by Tiedemann, Jorg during the OPUS project.
Supported Tasks and Leaderboards
translation: The dataset can be used to train a model for translation.
Languages
In our case, the corpora consists of a pair of source and target sentences for all 22 different languages from the… See the full description on the dataset page: https://huggingface.co/datasets/qanastek/EMEA-V3.ECDC
ECDC : An overview of the European Union's highly multilingual parallel corpora
Dataset Summary
In October 2012, the European Union (EU) agency 'European Centre for Disease Prevention and Control' (ECDC) released a translation memory (TM), i.e. a collection of sentences and their professionally produced translations, in twenty-five languages. The data gets distributed via the web pages of the EC's Joint Research Centre (JRC).
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/qanastek/ECDC.ASTD_Arabic_Sentiment_Tweets_Dataset
Citation
(https://aclanthology.org/D15-1299/)
