mjbommar/opengloss-v1.2-hard-negative-pairs
OpenGloss Hard Negative Pairs v1.2 This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains. Dataset Summary Total records: 73,244 Unique lexemes: 11,522 Relation Distribution Relation Type Count same_domain_wrong_entity 27,216 near_fact_confusion 22,095 style_variant 11… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.2-hard-negative-pairs.
OpenGloss Hard Negative Pairs v1.2
This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains.
Dataset Summary
- Total records: 73,244
- Unique lexemes: 11,522
Relation Distribution
Domain Distribution
Difficulty Distribution
Label Distribution
Recommended Use
- calibration supervision for embedding models
- reducing over-scoring of sibling concepts and same-domain wrong entities
- improving low-label score behavior in the 0.20 to 0.35 range
Record Schema
Each record includes:
idtext_atext_blabelrelation_typedomaindifficultyanchor_lexemecandidate_lexemelexeme_id_alexeme_id_b- optional
source - optional
notes
License
This dataset is released under Creative Commons Attribution 4.0 International (CC-BY 4.0).
Related Datasets
- OpenGloss v1.2 Dictionary
- OpenGloss v1.2 Definitions
- OpenGloss v1.2 Query Examples
- OpenGloss v1.2 Contrastive Examples
Generated from the OpenGloss v1.2 lexicon.
