Tarka-AIR/Tarka-Embedding-10M-V1-Preview
322
Training details and a stable version of the model will be released soon. In the meantime, feel free to give this model a try.
Model details
Tarka-Embedding-10M-V1 has the following features:
- Model Type: Text Embedding
- Supported Languages: English.
- Number of Paramaters: 10M
- Context Length: Optimal performance is observed with inputs under 1K tokens
- Embedding Dimension: 1024
Training Details
- Initialization: Based on Qwen3/Qwen3-Embedding-0.6B
- Architecture Modifications: The tokenizer is replaced with modernbert tokenizer . We use SVD decomposition with rank of 64 for the compression of both Transformer layers and the embedding layer.
- Teacher Model: Qwen3/Qwen3-Embedding-0.6B
