mjbommar/opengloss-v1.2-hard-negative-pairs
OpenGloss Hard Negative Pairs v1.2 This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains. Dataset Summary Total records: 73,244 Unique lexemes: 11,522 Relation Distribution Relation Type Count same_domain_wrong_entity 27,216 near_fact_confusion 22,095 style_variant 11… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-hard-negative-pairs.
012
