CoolFace
Modelpublic

ethicalabs/Echo-DSRN-v0.1.3-Embed-Intent

sourceHugging Faceapache-2.0updated 1mo agoView on Hugging Face
0likes164downloads
Model Card

Echo-DSRN-v0.1.3-Embed-Intent

![GitHub](https://github.com/ethicalabs-ai/Echo-DSRN/) ![License](https://opensource.org/licenses/Apache-2.0) ![Python](https://www.python.org/downloads/) ![Model Collection](https://huggingface.co/collections/ethicalabs/echo-dsrn) ![Hybrid Collection](https://huggingface.co/collections/ethicalabs/echo-dsrn-hybrid) ![Working Paper](https://github.com/ethicalabs-ai/Echo-DSRN/blob/main/PAPER.md)

98M-parameter multilingual intent classification embedding model based on the Echo-DSRN architecture (Dual-State Recurrent Neural Network) ◦ Recurrent Hybrid.

Fine-tuned from `ethicalabs/Echo-DSRN-v0.1.3-Embed-Exp` on Amazon MASSIVE across all 51 languages using MultipleNegativesRankingLoss (MNRL).

Model specs

PropertyValue
ArchitectureEcho-DSRN (Recurrent Neural Network)
Parameters98,266,629 (~98M)
Layers8 DSRN blocks
Hidden dim512
Attention heads4
Vocab size32,017 tokens
Precisionfp32
Base modelethicalabs/Echo-DSRN-v0.1.3-Embed-Exp
GPUAMD Radeon AI Pro R9700 (ROCm 7.2)

MTEB Results

MassiveIntentClassification (60 intents, 51 languages)

MetricScore
Accuracy (mean)72.42%
Accuracy (min)62.98%
Accuracy (max)78.33%
F1 (mean)66.27%
F1 (min)56.52%
F1 (max)71.94%

MassiveScenarioClassification (17 scenarios, 51 languages)

MetricScore
Accuracy (mean)79.00%
Accuracy (min)71.62%
Accuracy (max)84.30%
F1 (mean)78.28%
F1 (min)70.05%
F1 (max)84.08%

Evaluated via MTEB v2.12.30 logistic regression protocol on frozen embeddings. Per-language scores available in the model-index metadata.

Training

  • Base model: ethicalabs/Echo-DSRN-v0.1.3-Embed-Exp (STS-pretrained, 0.753 avg Spearman on MTEB STS)
  • Dataset: Amazon MASSIVE, all 51 locales (~1M training utterances)
  • Loss: MultipleNegativesRankingLoss with intent-grouped positive pairs
  • Pooling: mean_c_all (2048-dim recurrent slow state)
  • Convergence: Early stopping at epoch 1.2; linear accuracy gain (+2 pts/1k steps), no grokking plateau
  • Random baseline: ~1.7% (60-class 1-NN)

Example Usage

python
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("ethicalabs/Echo-DSRN-v0.1.3-Embed-Intent", trust_remote_code=True, device="cpu")

sentences = [
  "Can I order a pizza?",
  "I am so hungry. what about pizza?",
  "I like spaghetti."
]
embeddings = model.encode(sentences)

similarities = model.similarity(embeddings, embeddings)
print(similarities.shape)

Console output

torch.Size([3, 3])
>>> print(similarities)
tensor([[1.0000, 0.9333, 0.7381],
        [0.9333, 1.0000, 0.8411],
        [0.7381, 0.8411, 1.0000]])

Intent Vector Space Behavior

Here is what the model is actually doing under the hood for each pair:

  1. 1.`Sim(0, 1) = 0.9333` — Matching Actionable Intent
  2. 2.Sentence 0: "Can I order a pizza?"
  3. 3.Sentence 1: "I am so hungry. what about pizza?"
  4. 4.Analysis: Despite using completely different phrasing and syntax (one is a direct question, the other is a multi-sentence conversational prompt), the model maps them to nearly the same spot in vector space. The recurrent slow state identifies the underlying action (order_food) and topic (pizza), yielding a massive 0.9333 correlation.
  1. 1.`Sim(0, 2) = 0.7381` — Action vs. Statement Separation
  2. 2.Sentence 0: "Can I order a pizza?"
  3. 3.Sentence 2: "I like spaghetti."
  4. 4.Analysis: Notice the significant drop down to 0.7381. Even though both sentences live in the general domain of Italian food, the model correctly separates an actionable transactional request ("Can I order...") from a static statement of personal preference ("I like..."). This is where the fine-tuning on the MASSIVE dataset with MultipleNegativesRankingLoss shines: it prevents the model from relying purely on lexical topic overlap.
  1. 1.`Sim(1, 2) = 0.8411` — Conversational Context
  2. 2.Sentence 1: "I am so hungry. what about pizza?"
  3. 3.Sentence 2: "I like spaghetti."
  4. 4.Analysis: This pair scores higher (0.8411) than (0, 2). Because Sentence 1 expresses a state/desire ("I am so hungry"), its semantic profile sits naturally between an explicit ordering command and a preference statement.

Citation

bibtex
@software{echo_dsrn_embed_intent,
  author = {Massimo Roberto Scamarcia},
  title = {Echo-DSRN-v0.1.3-Embed-Intent: Multilingual Intent Classification Embeddings},
  year = {2026},
  url = {https://huggingface.co/Echo-DSRN-v0.1.3-Embed-Intent}
}

Note on tokenizer padding

The benchmark results on this card were measured with left padding (padding_side: left), and this model version reproduces them under that convention. A right-padded training version is planned: right padding keeps padded-batch embeddings consistent with single-request embeddings (leading pad tokens do not pollute the recurrent state), so future checkpoints will be batch-composition independent.