CoolFace
Datasetpublic

puttatidam/stsbiosses-crosslingual-mya-sts

stsbiosses-crosslingual-mya-sts Deduplicated copy of kornwtp/stsbiosses-crosslingual-mya-sts, part of the SEA-BED data-quality work. Source dataset: kornwtp/stsbiosses-crosslingual-mya-sts Deduplicated on: 2026-09-04 Task type: sts Splits: train What changed Kept in this dataset's ORIGINAL schema (sentence1/sentence2/score). Identical sentence pairs are collapsed to one row -- a repeat is counted twice in the rank correlation and so carries double weight for no… See the full description on the dataset page: https://huggingface.co/datasets/puttatidam/stsbiosses-crosslingual-mya-sts.

sourceHugging Faceupdated 12d agoView on Hugging Face
0likes61downloads
Dataset Card

stsbiosses-crosslingual-mya-sts

Deduplicated copy of `kornwtp/stsbiosses-crosslingual-mya-sts`, part of the SEA-BED data-quality work.

  • —Source dataset: kornwtp/stsbiosses-crosslingual-mya-sts
  • —Deduplicated on: 2026-09-04
  • —Task type: sts
  • —Splits: train

What changed

Kept in this dataset's ORIGINAL schema (sentence1/sentence2/score). Identical sentence pairs are collapsed to one row -- a repeat is counted twice in the rank correlation and so carries double weight for no reason -- taking the mean of the group's gold scores, which agree with each other by construction. A repeated pair whose gold scores DISAGREE by more than 25% of the declared scale is collapsed to one row with the competing scores recorded in a new conflict_score column (NULL elsewhere); score keeps the kept row's own value rather than inventing one, so a downstream human/LLM pass can resolve it and write back. Pairs where sentence1 == sentence2, and scores outside the declared range, are reported in the summary CSV but left in place: neither is duplication.

Row counts and the exact policy used are recorded in sts_dedup_summary.csv in the analysis repo. This dataset is not identical to the published SEA-BED data -- scores computed on it are not directly comparable to published SEA-BED numbers.