CoolFace
Datasetpublic

mjbommar/opengloss-v1.2-hard-negative-pairs

OpenGloss Hard Negative Pairs v1.2 This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains. Dataset Summary Total records: 73,244 Unique lexemes: 11,522 Relation Distribution Relation Type Count same_domain_wrong_entity 27,216 near_fact_confusion 22,095 style_variant 11… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-hard-negative-pairs.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
0likes12downloads
Dataset Card

OpenGloss Hard Negative Pairs v1.2

This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains.

Dataset Summary

  • —Total records: 73,244
  • —Unique lexemes: 11,522

Relation Distribution

Relation TypeCount
samedomainwrong_entity27,216
nearfactconfusion22,095
style_variant11,519
true_match11,517
sibling_concept897

Domain Distribution

DomainCount
geography23,619
history18,628
art7,431
civics4,879
biology4,758
religion4,678
general_academic2,080
law2,070
education2,050
linguistics2,016
literature605
anthropology195
chemistry145
technology70
astronomy6
psychiatry4
biologyandmedicine2
government2
medicine2
physics2
technologyandneuroscience2

Difficulty Distribution

DifficultyCount
medium38,735
hard22,992
easy11,517

Label Distribution

LabelCount
0.2027,216
0.3522,992
0.8011,519
1.0011,517

Recommended Use

  • —calibration supervision for embedding models
  • —reducing over-scoring of sibling concepts and same-domain wrong entities
  • —improving low-label score behavior in the 0.20 to 0.35 range

Record Schema

Each record includes:

  • —id
  • —text_a
  • —text_b
  • —label
  • —relation_type
  • —domain
  • —difficulty
  • —anchor_lexeme
  • —candidate_lexeme
  • —lexeme_id_a
  • —lexeme_id_b
  • —optional source
  • —optional notes

License

This dataset is released under Creative Commons Attribution 4.0 International (CC-BY 4.0).

Related Datasets


Generated from the OpenGloss v1.2 lexicon.