norsk
Datasets
All datasets matching “norsk”eti-embedding-training-data-2048-v3
Version note (v3): third generation of the ETI dense training data. Reuses the questions of NorskHelsenett/eti-embedding-training-data-v2 (which trained eti-embeddinggemma-v2), but re-maps every question to its article URL, re-cuts positives at a 2048-token window, re-mines hard negatives with granite + reranker validation (pos_score/neg_score), adds the keyword half, and filters junk anchors. Trained NorskHelsenett/eti-embeddinggemma-v3.
eti-embedding-training-data-2048… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-v3.eti-embedding-training-data-2048-triplets
ETI Embedding Training Data — Triplets with Hard Negatives
This dataset contains 330,120 (anchor, positive, negative) triplets for training and fine-tuning Norwegian-language embedding models, particularly for health-related retrieval and RAG applications.
How this dataset was created
Source data
The triplets were mined from the source dataset NorskHelsenett/eti-embedding-training-data-2048, which contains 78,888 anchor-positive pairs of Norwegian health… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-2048-triplets.eti-embedding-training-data-v2
ETI Embedding Training Data v2
576,708 Norwegian-language (anchor, positive, negative) triplets for contrastive fine-tuning of embedding, sparse-retrieval, and reranker models in the Norwegian health and welfare domain.
This dataset is a weighted, deduplicated merge of four upstream triplet datasets, designed to give a single high-quality training signal for two-stage RAG pipelines that retrieve over Norwegian public health information.
The model trained on it —… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/eti-embedding-training-data-v2.norsk-olje-gass-LLM
Dataset Card for ynuwara/norsk-olje-gass-LLM
Dataset Description
This dataset contains images converted from PDFs using the PDFs to Page Images Converter Space.
Number of images: 236
Number of PDFs processed: 1
Sample size per PDF: 100
Created on: 2024-12-14 15:00:48
Dataset Creation
Source Data
The images in this dataset were generated from user-uploaded PDF files.
Processing Steps
PDF files were uploaded to the PDFs to Page Images… See the full description on the dataset page: https://huggingface.co/datasets/ynuwara/norsk-olje-gass-LLM.norsk-olje-gass-QnA-ColPaliLOS_Document_classification_ETI
LOS Document Classification (ETI)
Norwegian public-sector documents labelled with their top-level LOS category
(Felles vokabular / common vocabulary, level 1). Built to train a small,
self-hostable text classifier by distilling LLM-generated labels.
Splits
split
rows
labels
use
train
~1982
silver — LLM-generated (Claude Haiku 4.5, few-shot)
training
test
551
human gold — manually annotated
honest evaluation
The train labels are weak supervision… See the full description on the dataset page: https://huggingface.co/datasets/NorskHelsenett/LOS_Document_classification_ETI.
