faroese
Datasets
All datasets matching “faroese”FaroeseSTS
FaroeseSTS
An MTEB dataset
Massive Text Embedding Benchmark
Semantic Text Similarity (STS) corpus for Faroese.
Task category
t2t
Domains
News, Web, Written
Reference
https://aclanthology.org/2023.nodalida-1.74.pdf
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["FaroeseSTS"])
evaluator = mteb.MTEB(task)
model = mteb.get_model(YOUR_MODEL)
evaluator.run(model)… See the full description on the dataset page: https://huggingface.co/datasets/mteb/FaroeseSTS.faroese-stsThis is a Semantic Text Similarity (STS) corpus for Faroese, Fo-STS, it was created by translating the English STS dataset.
If you find this dataset useful, please cite
@inproceedings{snaebjarnarson-etal-2023-transfer,
title = "{T}ransfer to a Low-Resource Language via Close Relatives: The Case Study on Faroese",
author = "Snæbjarnarson, Vésteinn and
Simonsen, Annika and
Glavaš, Goran and
Vulić, Ivan",
booktitle = "Proceedings of the 24th Nordic Conference on… See the full description on the dataset page: https://huggingface.co/datasets/vesteinn/faroese-sts.faroese-dynaword
🧨 Faroese Dynaword
Version
0.0.7 (Changelog)
Language
Faroese (fo, fao)
License
Openly Licensed, See the respective dataset
Models
Currently there are no models trained on this dataset
Contact
If you have question about this project please create an issue here
Dataset Description
Number of samples: 405.81K
Number of tokens (Llama 3): 45.40M
Average document length in tokens (min, max): 111.87 (2, 109.50K)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dynaword.scandinavian_faroesefaroese-dyna-instruct
🧨 Faroese dyna-instruct
Version
0.1.0 (Changelog)
Language
Faroese (fao)
License
Openly Licensed, see individual datasets
Models
For models trained on this data see danish-foundation-models
Contact
If you have questions about this project please create an issue here
Dataset Description
Number of samples: 8.61K
Number of tokens (Llama 3): 2.64M
Average conversation length in tokens (min, max): 306.67 (98, 1.24K)
Average number of… See the full description on the dataset page: https://huggingface.co/datasets/danish-foundation-models/faroese-dyna-instruct.faroese-blimp-single-error
