spider-ai/arachnia-nano-datasets
Arachnia Nano 270m Arachnia Nano aims to be a lightweight retrieval-based context-aware proactive assistant for: Retrieval-augmented factual autocompletion Email auto completion General-purpose web text completion Suggestions Architecture Arachnia Nano is based on these two models: Gemma 3 270m IT EmbeddingGemma We aim to bridge these two models into a single architecture to sync what generative model wants to see and what the embedding layer shows it.… See the full description on the dataset page: https://huggingface.co/datasets/spider-ai/arachnia-nano-datasets.
Arachnia Nano 270m
Arachnia Nano aims to be a lightweight retrieval-based context-aware proactive assistant for:
- Retrieval-augmented factual autocompletion
- Email auto completion
- General-purpose web text completion
- Suggestions
Architecture
Arachnia Nano is based on these two models:
- Gemma 3 270m IT
- EmbeddingGemma
We aim to bridge these two models into a single architecture to sync what generative model wants to see and what the embedding layer shows it.
Template {#template}
The template is based on the Gemma 3 chat template. There are four turns, each starting with <start_of_turn> and ending with <end_of_turn>
systembrowser_stateusermodel
browser_state has three parameters:
ACTIVE_TABREFERENCE_TAB
context_retrieval contains excerpts of important data filtered with EmbeddingGemma.
Suggestion example:
<start_of_turn> system
You are Arachnia, a proactive browser assistant. You analyze open tabs to provide autocomplete and context-aware suggestions.
<end_of_turn>
<start_of_turn> browser_state
ACTIVE_TAB: "Auvo Nikasios - Reading comprehension" (docs.google.com)
OTHER_TABS: "Harvard University - Wikipedia" (en.wikipedia.org), “ChatGPT” (chatgpt.com), “Harvard - Google Search” (google.com)
<end_of_turn>
<start_of_turn> context_retrieval
"Harvard University - Wikipedia" (en.wikipedia.org)
Harvard University is a private research university in Cambridge, Massachusetts. It was established in 1636.
<end_of_turn>
<start_of_turn> user
[Cursor @ Line 12]: "Harvard i"
<end_of_turn>
<start_of_turn> model
<thought> User is starting to define Harvard based on the Wikipedia tab. </thought>
<autocomplete> is a highly residential research university known for the world's largest academic library system. </autocomplete>
<suggestion> Would you like me *to summarize the Harvard University article on “en.wikipedia.org”*? </suggestion>
<end_of_turn>Summary example:
<start_of_turn> system
You are Arachnia, a proactive browser assistant. The user has accepted your suggestion to summarize the Harvard University article on “en.wikipedia.org”.
<end_of_turn>
<start_of_turn> context_retrieval
"Harvard University - Wikipedia" (en.wikipedia.org)
Harvard University is a private Ivy League research university in Cambridge, Massachusetts, United States. Founded in 1636, and named Harvard College in 1639 in honor of its first benefactor, Puritan clergyman John Harvard, it is the oldest institution of higher learning in the United States. Its influence, wealth, and rankings have made it one of the most prestigious universities in the world…
<end_of_turn>
<start_of_turn> model
Harvard University, located in Cambridge, Massachusetts, is a private Ivy League research university and the oldest institution of higher learning in the United States!
Key Facts
* Founding: Established in 1636 and named after its first benefactor, John Harvard.
* Wealth: It holds the world's largest academic endowment, valued at $55.7 billion in 2025.
* Academics: The university is composed of 13 schools, including Harvard College (undergraduate), Harvard Medical School, and Harvard Law School.
* Resources: It houses the Harvard Library, the largest academic library system in the world with over 20 million volumes.Datasets
The stage 1 dataset contains synthetic data generated by Gemma 4 26B A4B to teach the model the basic template.
Process
Training Arachnia Nano will follow these stages:
Stage 1: Training generative model
Before anything else, we must train the generative model based on Gemma 3 270m to understand the template.
To truly prevent collapse during a one-shot run, consider Stage 1 Freezing:
- Freeze all layers except the input embeddings and the output head.
- Train for just 50–100 steps. This teaches the model only what the new tokens mean without touching the "reasoning" layers.
- Unfreeze everything.
- Continue fine tuning with synthetic data.
Stage 2: Training embedding model
Rather than explicitly labeling the “gold standard”, which may contain errors, we train the embedding layer to output the least tokens that still keep the final output from the generative model good.
We use “joint-training” to ensure the generative model gets the required information while minimizing prefill time.
We calculate the loss using “efficiency tax”, measuring how many tokens were outputted and the quality of the output.
We’ll use real world data for this, to effectively utilize the joint training to ensure both of the models can work with real world noise.
