alexxxmf/survaiv-embeddinggemma-300m
SurvAIv EmbeddingGemma 300M — epoch 1
SurvAIv's fine-tuned embedding checkpoint, derived from Google's EmbeddingGemma 300M. This is a complete checkpoint, not a LoRA adapter. The root contains the original trained Sentence Transformers modules, configuration and tokenizer. gguf/embedding-q8_0.gguf is the Q8_0 deployment export currently used by the local SurvAIv RAG, with 768-dimensional vectors.
The recorded training data has 4,000 train, 592 validation and 408 test queries. The selected checkpoint is epoch 1 of the recorded epoch sweep. Training used maximum sequence length 512 and seed 42. Preserve the search formatting:
- Query prefix:
task: search result | query: - Document prefix:
title: none | text:
The original base-model card is retained in UPSTREAM_MODEL_CARD.md for attribution and background; its base-model capabilities and results are not new measurements of this fine-tune. The saved run configuration does not pin an immutable base revision. The previously missing mined negatives have been recovered from the original v1 output archive reused by the epoch sweep. Exact cross-environment retraining reproducibility is not claimed.
Training inputs, splits, mined negatives and provenance: SurvAIv embedding training dataset (private; authorized access required).
This model is a retrieval component, not a clinically validated first-aid system. No standalone accuracy or safety claim is made here. Gemma's terms apply; see LICENSE.html for the saved upstream terms and NOTICE.md for attribution. This public release retains the original epoch-1 weights; it is not one of the later checkpoints trained on the -fix dataset.
Licence and distribution
Use, reproduction, modification and distribution of this model are governed by the Gemma Terms of Use, including the Section 3.2 restrictions and incorporated Prohibited Use Policy. These restrictions apply to recipients and subsequent distributions; this repository grants no exception to them. Include the terms and notice when redistributing the model. This is not an unrestricted Apache/MIT release.
Training-data access is separate: the linked dataset remains private pending source-text licensing and attribution preparation. The model licence is not a licence to redistribute the source documents or training passages. No independent legal clearance of the training process is represented.
