datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
faroese-stsThis is a Semantic Text Similarity (STS) corpus for Faroese, Fo-STS, it was created by translating the English STS dataset.
If you find this dataset useful, please cite
@inproceedings{snaebjarnarson-etal-2023-transfer,
title = "{T}ransfer to a Low-Resource Language via Close Relatives: The Case Study on Faroese",
author = "Snæbjarnarson, Vésteinn and
Simonsen, Annika and
Glavaš, Goran and
Vulić, Ivan",
booktitle = "Proceedings of the 24th Nordic Conference on… See the full description on the dataset page: https://huggingface.co/datasets/vesteinn/faroese-sts.stsb
Glue STS-B
This dataset is a port of the official sts-b dataset on the Hub.
This is not a classification task, so the label_text column is only included for consistency
Note that the sentence1 and sentence2 columns have been renamed to text1 and text2 respectively.
Also, the test split is not labeled; the label column values are always -1.
ko-en-code-mixing-sts
Korean–English Code-Mixing STS Dataset
This dataset contains 1,500 Korean–English code-mixed pairs derived from KLUE-STS. We keep the original sentence_a and apply insertion-only code-mixing to sentence_b using an LLM (Gemini 2.5 Flash), recording where and how code-mixing occurred.
Interactive Dashboard
🌐 Explore the dataset interactively: https://watchstep.github.io/ko-en-cm/
The dashboard provides:
Interactive data exploration and filtering
Sample visualization with… See the full description on the dataset page: https://huggingface.co/datasets/watchstep/ko-en-code-mixing-sts.hf-sts-doc-example
HF STS doc example (verbatim control)
Exact JSONL from https://huggingface.co/docs/hub/session-traces-format
sts-probe
