CoolFace
Datasetpublic

mteb/BornholmBitextMining

BornholmBitextMining An MTEB dataset Massive Text Embedding Benchmark Danish Bornholmsk Parallel Corpus. Bornholmsk is a Danish dialect spoken on the island of Bornholm, Denmark. Historically it is a part of east Danish which was also spoken in Scania and Halland, Sweden. Task category t2t Domains Web, Social, Fiction, Written Reference https://aclanthology.org/W19-6138/ Source datasets: strombergnlp/bornholmsk_parallel How to evaluate on this task… See the full description on the dataset page: https://huggingface.co/datasets/mteb/BornholmBitextMining.

sourceHugging Facecc-by-4.0updated 7mo agoView on Hugging Face
0likes11kdownloads
5 commits on main
4e0a8607mo ago

Add eval config

Samoed
5b020481y ago

Add dataset card

Samoed
bc1455b1y ago

Add dataset card

Samoed
1d06a9f1y ago

Upload dataset

Samoed
3b361f41y ago

initial commit

Samoed