thivy/eti-embedding-training-data-2048-triplets-v4
ETI Embedding Training Data — v4 (5-style, LLM-judged triplets) Norwegian (nb) retrieval triplets built from the NorskHelsenett/LOS_Document_classification_ETI corpus of public-service / welfare / health documents. Each row is a triplet (anchor, positive, negative) plus three metadata columns (style, category, doc_url) you can use to filter, weight, or build curriculum stages. 76,408 triplets · 2,530 source documents · 38,629 distinct anchors. Schema field… See the full description on the dataset page: https://huggingface.co/datasets/thivy/eti-embedding-training-data-2048-triplets-v4.
This repository belongs to thivy on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
