t-tech/T-Search-NVFP4
T-Search-NVFP4
🚨 T-Search-NVFP4 is designed for use with the [official T-Search harness](https://github.com/turbo-llm/t-search-harness). The source T-Search checkpoint did not show noticeable differences relative to the Qwen3.6-35B-A3B base model on general-purpose benchmarks.
🚨 Users are advised to exercise caution and are responsible for any additional training and oversight required to ensure the model's responses meet acceptable ethical and safety standards. The responsibility for incorporating this model into industrial or commercial solutions lies entirely with those who choose to deploy it.
Highlights
We introduce T-Search-NVFP4, an agentic retriever built for difficult multi-step search in English and Russian. It plans and executes searches across multiple rounds with a configurable search budget.
- NVFP4 checkpoint: NVIDIA ModelOpt NVFP4 with 4-bit weights and input activations.
- Agentic retrieval: issues queries, inspects results, tracks coverage, carries compact evidence and search state across rounds, and returns a ranked chunk set.
- Retrieval quality: outperforms the base model and larger open models on average across the evaluated retrieval benchmarks.
- Retriever robustness: trained and evaluated with multiple retrievers, maintaining strong retrieval quality across configurations.
Description
T-Search-NVFP4 is built on Qwen3.6-35B-A3B and trained on fully synthetic search tasks generated for the same harness used at inference time.
The model is responsible for collecting evidence. Given a question and access to a corpus search backend, it formulates queries, reads retrieved snippets, decides which chunks are worth preserving, and returns a ranked evidence set. A downstream generator, reranker, or full-text fetcher can then consume that ranking.
⛷️ T-Search Harness
The official T-Search harness implements the inference protocol used for training and evaluation. The retriever operates in rounds. Within each round, it follows a ReAct loop with an approximately 32K-token context budget and three tools:
search_corpussearches the corpus and returns document chunks;save_and_advancepreserves important chunks and starts a new round;finalize_rankingends the search and returns a ranked evidence set.
The agent sees the original question, previously saved chunks, and the current coverage state: which parts of the question are supported by evidence and which remain open. It chooses search queries, inspects retrieved chunks, and decides whether to continue searching, preserve state for another round, or finalize.
At 75% context utilization, search_corpus is locked. The agent must either call finalize_ranking or use save_and_advance to build compact memory for the next round: preserved chunks with reasons, covered and open parts of the question, previous attempts, and a useful next step. Full tool history and intermediate noise do not cross the round boundary.
🧠 Training
Training proceeds in two stages: supervised fine-tuning on fully synthetic search trajectories, followed by reinforcement learning with GSPO and recall-based rewards. At each stage, separate English and Russian experts are trained on language-specific data and then merged into a single checkpoint.
Quantization
The checkpoint uses NVIDIA ModelOpt NVFP4 quantization. Linear weights and input activations use 4-bit floating-point quantization with a group size of 16, while the KV cache is configured for 8-bit floating point. The included MTP head and other modules listed under quantization_config.ignore remain in BF16.
Synthetic task factory
Training tasks contain a question, a fixed index, annotated evidence chunks, and a full tool-use trajectory. Candidate tasks pass adversarial checks for trivial query leakage, answerability from model weights, single-document shortcuts, missing evidence, and weak distractors.
Supervised fine-tuning
Long teacher trajectories are split into self-contained rounds. Invalid tool actions are masked from the loss, while useful recovery behavior after a tool error is retained. Productive rounds are selected by evidence gain rather than by final recall alone, preserving useful behavior from difficult tasks.
Training uses 11K SFT examples from 8K unique questions per language; one quarter targets robustness to different retrievers. An additional non-overlapping pool of 2K questions per language is reserved for RL.
Reinforcement learning
The policy is optimized with GSPO. RL optimizes the complete search policy on full tool-use trajectories. The main reward is recall over gold chunk_id values; precision and F-score are tracked for diagnostics but are not used as the primary training signal. This discourages the agent from finalizing early with a small high-precision but incomplete evidence set.
A detailed Russian-language training report will be available soon on Habr.
📊 Benchmarks
The results in this section were measured with the T-Search-FP8 checkpoint.
Evaluation datasets
The accompanying evaluation datasets are released as **TRuST** and **SynthComp**.
We use Recall@10 as the primary metric.
We report single-rollout results for all models and, for T-Search, an additional N=3 configuration that runs three independent rollouts in parallel and combines their rankings using reciprocal rank fusion (RRF).
We also study the latency–quality trade-off by varying the maximum number of search rounds and the number of parallel agent runs. This separates the effect of deeper sequential search from broader parallel exploration.
To evaluate retrieval robustness, we varied the search backend while keeping the model and agent configuration fixed.
👨💻 Usage
Recommended generation parameters
do_sample: true
temperature: 0.7
top_p: 1.0Serve t-tech/T-Search-NVFP4 through an OpenAI-compatible endpoint. See the harness README for installation and search-backend integration.
Reference SGLang serving setup
The ModelOpt FP4 backend requires NVIDIA Blackwell hardware. The following SGLang configuration follows the Qwen3.6 NVFP4 cookbook settings used to serve this checkpoint:
export SGLANG_ENABLE_SPEC_V2=1
python3 -m sglang.launch_server \
--model-path "/path/to/model" \
--served-model-name "t-tech/T-Search-NVFP4" \
--trust-remote-code \
--quantization modelopt_fp4 \
--host 0.0.0.0 \
--port 8000 \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--attention-backend trtllm_mha \
--speculative-algorithm EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--mamba-radix-cache-strategy extra_buffer \
--context-length 65536 \
--max-running-requests 16Example
from retriever_agent import (
AgentConfig,
HttpSearchClient,
OpenAILLMClient,
RetrieverAgent,
)
config = AgentConfig(model="t-tech/T-Search-NVFP4")
llm = OpenAILLMClient(["http://<sglang-host>:8000/v1"], config)
search = HttpSearchClient("http://<search-host>:8000")
agent = RetrieverAgent(config, llm, search)
result = agent.retrieve("your query")
for doc in result.documents:
print(doc.rank, doc.doc_id, doc.score, doc.text)