CoolFace
Datasetpublic

NorskHelsenett/eti-embedding-training-data-2048-triplets

ETI Embedding Training Data — Triplets with Hard Negatives This dataset contains 330,120 (anchor, positive, negative) triplets for training and fine-tuning Norwegian-language embedding models, particularly for health-related retrieval and RAG applications. How this dataset was created Source data The triplets were mined from the source dataset NorskHelsenett/eti-embedding-training-data-2048, which contains 78,888 anchor-positive pairs of Norwegian… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-triplets.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes57downloads
7 commits on main
9ef05cd5mo ago

Add NorskHelsenett attribution

thivy
9c4cb4c5mo ago

Update references to NorskHelsenett org

thivy
f8a33a17mo ago

Add dataset card with methodology and statistics

thivy
c79dac57mo ago

Add dataset card with methodology and statistics

thivy
4b0e6287mo ago

Add dataset card with methodology and statistics

thivy
b74aa507mo ago

Upload dataset

thivy
55cbe057mo ago

initial commit

thivy