NorskHelsenett/eti-embedding-training-data-2048-triplets
ETI Embedding Training Data — Triplets with Hard Negatives This dataset contains 330,120 (anchor, positive, negative) triplets for training and fine-tuning Norwegian-language embedding models, particularly for health-related retrieval and RAG applications. How this dataset was created Source data The triplets were mined from the source dataset NorskHelsenett/eti-embedding-training-data-2048, which contains 78,888 anchor-positive pairs of Norwegian… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-triplets.
Add NorskHelsenett attribution
Update references to NorskHelsenett org
Add dataset card with methodology and statistics
Add dataset card with methodology and statistics
Add dataset card with methodology and statistics
Upload dataset
initial commit
