CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ZurichNLP /mlit-guanaco Description Guanaco dataset subsets used for experiments in the paper Turning English-centric LLMs Into Polyglots: How Much Multilinguality Is Needed? We extend the original Guanaco dataset with language tags, with languages identified using OpenLID. The following subsets were used to train our experimental models: config name languages ml1 en ml2, mtml2 en, es ml3, mtml3 en, es, ru ml4, mtml4 en, es, ru, de ml5, mtml5 en, es, ru, de, zh ml6, mtml6 en, es, ru… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mlit-guanaco.tabular10K<n<100K2 likes455 downloads3y agoHugging Face02ZurichNLP /rsd-ists-2016Training and test data for the task of Recognizing Semantic Differences (RSD). See the paper for details on how the dataset was created, and see our code at https://github.com/ZurichNLP/recognizing-semantic-differences for an example of how to use the data for evaluation. The data are derived from the SemEval-2016 Task 2 for Interpretable Semantic Textual Similarity organized by Agirre et al. (2016). The original URLs of the data are: Train:… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/rsd-ists-2016.texttoken-classification10K<n<100K0 likes190 downloads1y agoHugging Face03ZurichNLP /mlit-alpaca-eval Description Translated versions of the AlpacaEval prompt dataset for evaluating the performance of chat LLMs. Translations were generated using gpt-3.5-turbo-0613 using the following prompt template (adapted from Lai et al, 2023): You are a helpful assistant. Translate the following text into {{target_language}}. Keep the structure of the original text and preserve things like code and names. Please ensure that your response contains only the translated text. The translation must… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/mlit-alpaca-eval.text10K<n<100K1 likes132 downloads3y agoHugging Face04ZurichNLP /swissner SwissNER A multilingual test set for named entity recognition (NER) on Swiss news articles. Description SwissNER is a dataset for named entity recognition based on manually annotated news articles in Swiss Standard German, French, Italian, and Romansh Grischun. We have manually annotated a selection of articles that have been published in February 2023 in the categories "Switzerland" or "Regional" on the following online news portals: Swiss Standard German: srf.ch… See the full description on the dataset page: https://huggingface.co/datasets/ZurichNLP/swissner.texttoken-classificationn<1K2 likes49 downloads3y agoHugging Face05adityarra07 /zurich_data Dataset Card for "zurich_data" More Information needed audio1K<n<10K0 likes12 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.