CoolFace
Datasetpublic

mjbommar/opengloss-v1.3-hard-negative-pairs

See also OpenGloss v2.1 (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list. OpenGloss Hard Negative Pairs v1.3 This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v1.3-hard-negative-pairs.

sourceHugging Facecc-by-4.0updated 20d agoView on Hugging Face
0likes58downloads
Dataset Card
See also [OpenGloss v2.1](https://huggingface.co/datasets/mjbommar/opengloss-v2.1-senses) (2026-09-07): a deeper release of 109,633 of these headwords — sense-level ids, four reading levels, sense-tagged examples with spans, a judged relation graph, and retrieval supervision — published as a 16-dataset family. v1.3 remains the broader headword list.

OpenGloss Hard Negative Pairs v1.3

This dataset contains calibration-oriented positive and low-label similarity pairs for embedding training. It is designed to reduce over-scoring of related-but-wrong matches and improve score separation in weak domains.

Dataset Summary

  • —Total records: 1,131,241
  • —Unique lexemes: 205,967

Relation Distribution

Relation TypeCount
samedomainwrong_entity566,913
style_variant205,963
true_match205,956
nearfactconfusion137,232
sibling_concept15,177

Domain Distribution

DomainCount
general737,516
history145,655
geography144,860
linguistics35,517
art10,134
religion9,736
science7,550
language7,265
education6,258
technology4,830
civics3,885
anthropology2,790
life-sciences2,350
society2,065
mathematics1,590
law1,440
economics1,124
arts1,025
biology782
literature690
philosophy610
food445
sports420
medicine375
chemistry310
zoology290
pharmacology170
politics160
botany110
anatomy105
physics65
music60
mineralogy55
entomology50
biochemistry45
taxonomy45
astronomy40
materials_science40
psychology35
genetics30
geology30
molecular_biology30
architecture25
ophthalmology25
ornithology25
computing20
construction20
ichthyology20
life_sciences20
mycology20
neuroscience20
textiles20
clothing12
agriculture6
archaeology6
biography6
bullfighting6
construction_engineering6
ecology6
economicsandaccounting6
endocrinology6
fantasy6
foodanddrink6
historyandliterature6
historyandpolitics6
informal_speech6
law/civics6
legalandfinancial6
meteorology6
military6
mythology6
onomastics6
politicsandlaw6
publishing6
rhetoric6
transport6
transportation6
typography6
psychiatry4
accountinganddocumentation2
acoustics2
administrative_communication2
aeronautics2
animal_husbandry2
architecture,_housing2
artandliterature2
artandmedia2
art_history2
aviation2
aviation_security2
biologyandmedicine2
biology_genetics2
business2
business,economics,management2
cardiology2
cartography2
ceramics2
chronology2
classical_studies2
color2
comics2
communication2
crafts2
craftsandrestoration2
cryptography2
culinary2
dance2
domestic_service2
earth_science2
ecologyandzoology2
economics,operationsmanagement2
economicsandlaw2
electrical_engineering2
electronics2
embryology2
environmental_science2
equine2
fashion2
filmandmedia2
finance2
finance_technology2
foodandlanguage2
food_preparation2
food_science2
forestry2
funerary2
furnitureandrestoration2
government2
graph_theory2
heraldry2
historical_administration2
historicalandhonorific_usage2
historicalandinstitutional2
historical_architecture2
historicalsocialwelfare2
historical_transport2
historical_weapons2
historyandclassical_mythology2
historyandgeography2
historyandphilosophy2
human-computerinteraction,interface_design2
kinship2
languageandgeneral2
languageandprinting2
languageandvisual_perception2
law,finance,administration2
law,_politics2
law/politics2
lawandcivics2
lawandpolitics2
law_enforcement2
legalandgeneral2
legalandgeneral_use2
legalandpolitical2
limnology2
literary2
literary_studies2
literatureandfilm2
manufacturing2
marketing,_business2
mathematicsandcomputer_science2
mathematicsandphysics2
mathematicsandstatistics2
medical_imaging2
metallurgy2
military_drill2
musicandspeech2
music_history2
music_theory2
mythologyandliterature2
nautical2
nautical_engineering2
neuroanatomy2
neuropsychology2
neuroscienceandophthalmology2
nonequilibrium_thermodynamics2
nutrition,behavioralscience2
occult2
occult_studies2
occupation2
oenology2
onoma2
organic_chemistry2
otolaryngology2
ottoman_studies2
packaging2
paleontology2
particle_physics2
petroleum_refining2
phonetics2
photometry2
physical_geography2
physicsandmaterials_science2
physics_chemistry2
physiology2
poetry2
political_geography2
political_vocabulary2
postal_services2
prosody2
religionandhistorical_language2
scienceandengineering2
sleep_disorders2
sleep_medicine2
sports_betting2
statistics2
symbolism2
technologyandmedia2
technologyandmilitary2
technologyandneuroscience2
textile2
theatre_theory2
transportation_engineering2
travel2
urban_planning2
videogamehistory2
video_games2
woodworking2

Difficulty Distribution

DifficultyCount
medium772,876
easy205,956
hard152,409

Label Distribution

LabelCount
0.20566,913
0.35152,409
0.80205,963
1.00205,956

Recommended Use

  • —calibration supervision for embedding models
  • —reducing over-scoring of sibling concepts and same-domain wrong entities
  • —improving low-label score behavior in the 0.20 to 0.35 range

Record Schema

Each record includes:

  • —id
  • —text_a
  • —text_b
  • —label
  • —relation_type
  • —domain
  • —difficulty
  • —anchor_lexeme
  • —candidate_lexeme
  • —lexeme_id_a
  • —lexeme_id_b
  • —optional source
  • —optional notes

License

This dataset is released under Creative Commons Attribution 4.0 International (CC-BY 4.0).

Related Datasets


Generated from the OpenGloss v1.3 lexicon.